How can transformers improve video search beyond keyword matching?

How can AI transformer models (like BERT, RoBERTa, or other language models) make video search smarter than just matching exact words?