BERT, short for Bidirectional Encoder Representations from Transformers, is a language model introduced by Google researchers in 2018 that changed how machines process and understand text. Its defining innovation is reading words in both directions simultaneously, meaning it considers the full surrounding context of every word rather than processing text from left to right or right to left alone. That seemingly simple shift turned out to be enormously consequential, and understanding why requires looking at what BERT actually does, how people use it, and where its design both shines and falls short.
What Makes BERT Bidirectional
Before BERT, the dominant approach to training language models was essentially one-directional. A model would read a sentence from left to right, predicting each next word based only on the words that came before it. Some models read right to left, and a few tried combining both directions in separate passes, but none truly considered both sides at once during every step of training. BERT broke from that pattern by jointly conditioning on both left and right context in all layers of the model, meaning every word representation is informed by the entire sentence from the start.1ACL Anthology. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Why does this matter in practice? Consider the word “bank” in the sentences “She sat on the river bank” and “She went to the bank to deposit a check.” A left-to-right model processing “bank” has seen “She sat on the river” or “She went to the,” which might be enough in these clear-cut cases. But many real sentences are far more ambiguous, and the words that come after a tricky word can be just as important for figuring out its meaning as the words that come before. BERT gets both sides at once, which lets it build richer representations of what each word means in its specific context.
How BERT Learns During Pre-Training
BERT’s pre-training relies on a technique called masked language modeling. During training, the model is shown sentences with some words randomly hidden, and it has to predict what those hidden words are based on the surrounding context. This forces the model to develop a deep understanding of how language works, because it needs context from both directions to fill in the blanks accurately. The approach has been widely adopted as a standard framework for self-supervised language pre-training.2arXiv. Improving Self-supervised Pre-training via a Fully-Explored Masked Language Model
The original BERT also included a second training objective called next sentence prediction, where the model had to decide whether two sentences naturally followed each other. The idea was to help BERT understand relationships between sentences, not just individual words. This turned out to be one of the more controversial design choices. Subsequent research showed that next sentence prediction is actually detrimental to training, partly because of how it splits context and partly because the signal it provides is too shallow to be useful.3ACL Anthology. On Losses for Modern Language Models Later variants of BERT, most famously RoBERTa, dropped next sentence prediction entirely and saw improved results.4arXiv. RoBERTa: A Robustly Optimized BERT Pretraining Approach
Fine-Tuning for Specific Tasks
Pre-training gives BERT a general understanding of language, but the model becomes useful for real-world applications through a process called fine-tuning. You take the pre-trained model and train it a bit further on a smaller, task-specific dataset. Want BERT to classify customer reviews as positive or negative? Fine-tune it on labeled reviews. Need it to answer questions about a passage of text? Fine-tune it on question-answer pairs. This two-stage approach, pre-train broadly and then fine-tune narrowly, became the dominant pattern in natural language processing after BERT’s release.
There is a practical catch, though. Standard fine-tuning updates every single parameter in the model for each new task. That means if you have ten different tasks, you end up storing ten full copies of the model, each slightly different. Researchers have explored more efficient alternatives, such as adapter modules that insert small trainable layers into the frozen pre-trained model, dramatically reducing the number of parameters you need to update per task.5arXiv. Parameter-Efficient Transfer Learning for NLP These methods have become increasingly important as models have grown larger and storage costs have become a real concern for organizations running many specialized models.
Contextual Embeddings and Why They Matter
One of BERT’s lasting contributions is the concept of contextual word embeddings. Before BERT and similar models, the standard approach in language processing was to represent each word as a single fixed vector, a list of numbers capturing its meaning. The word “bank” would get one representation regardless of whether it appeared in a sentence about rivers or finance. BERT generates a different representation for the same word depending on the sentence it appears in, which is a fundamentally different and more expressive approach.
These contextual representations quickly made older static embeddings feel outdated for many applications. Interestingly, researchers have found that you can actually convert BERT’s contextual representations back into static ones by pooling across many different contexts, and the resulting static embeddings turn out to be higher quality than the originals, suggesting that BERT’s training process captures something richer about word meaning even when you strip away the context-sensitivity.6ACL Anthology. Interpreting Pretrained Contextualized Representations via Reductions to Static Embeddings
Probing studies have also revealed that BERT’s internal layers organize linguistic knowledge in structured ways. Different groups of hidden units appear to specialize in encoding different linguistic properties, from basic part-of-speech information to more complex syntactic relationships.7ACL Anthology. How Do BERT Embeddings Organize Linguistic Knowledge? The model develops this structure on its own, without ever being explicitly taught grammar rules.
How BERT Differs from GPT
BERT and GPT are probably the two most recognized names in language modeling, and people often wonder how they relate to each other. The core difference comes down to architecture and training direction. BERT is an autoencoder-style model: it receives input text and produces a contextual embedding for each token, making it strong at understanding and classifying text. GPT is an autoregressive model: it receives a sequence of tokens as input and predicts the next token, making it strong at generating text.8Artificial Intelligence: Foundations, Theory, and Algorithms. Pre-trained Language Models
This architectural distinction has practical consequences. BERT excels at tasks where you need to understand a piece of text as a whole: sentiment analysis, named entity recognition, question answering over a given passage. GPT excels at tasks where you need to produce text: writing summaries, holding conversations, generating code. BERT “reads” well; GPT “writes” well. In practice, the field has increasingly moved toward autoregressive models like GPT for general-purpose applications, partly because generation capabilities have proved more commercially versatile. But BERT-style models remain widely used in classification, search, and retrieval tasks where understanding existing text is the goal rather than generating new text.
Smaller and Faster Variants
The original BERT came in two sizes. BERT-base had 110 million parameters, and BERT-large had 340 million. Even BERT-base was too large and slow for many practical applications, especially on mobile devices or in systems that need to process thousands of queries per second. This spurred a wave of research into making BERT smaller without losing too much capability.
DistilBERT is one of the most successful compressed versions. Using a technique called knowledge distillation, where a smaller model learns to mimic the behavior of a larger one, DistilBERT reduces BERT’s size by about 40% while retaining roughly 97% of its language understanding capabilities and running about 60% faster.9arXiv. DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter For many applications where a small accuracy trade-off is acceptable, DistilBERT offers a much more practical option.
ALBERT took a different approach, using parameter-sharing techniques to reduce memory consumption and increase training speed. Rather than simply shrinking the model, ALBERT shares parameters across layers, meaning the same weights get reused multiple times during processing. It also replaced the next sentence prediction objective with a sentence-order prediction task. These changes allowed ALBERT to achieve state-of-the-art results on several benchmarks while using fewer parameters than BERT-large.10arXiv. ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
RoBERTa showed that you could also improve BERT dramatically just by training it better. The original BERT was, by the RoBERTa team’s assessment, significantly undertrained. By training longer, on more data, with larger batches, and without next sentence prediction, RoBERTa matched or exceeded every model published after BERT without any architectural changes.11arXiv. RoBERTa: A Robustly Optimized BERT Pretraining Approach The lesson was a bit humbling for the field: sometimes the architecture is fine and you just need to train it properly.
Multilingual BERT and Cross-Language Transfer
Google released a multilingual version of BERT, often called mBERT, trained simultaneously on text from 104 languages. What surprised researchers was how well it transferred knowledge across languages. You could fine-tune mBERT on an English task and then apply it to the same task in a language it had never been fine-tuned on, and it would perform competitively. This zero-shot cross-lingual transfer worked across tasks like named entity recognition, part-of-speech tagging, document classification, natural language inference, and dependency parsing, spanning dozens of languages from various language families.12ACL Anthology. Beto, Bentz, Becas: The Surprising Cross-Lingual Effectiveness of BERT
The effectiveness is genuine but uneven. Languages that are well represented in the training data and share structural similarities with English tend to benefit more from transfer. Languages with very different word orders, scripts, or morphological systems sometimes struggle more. Researchers have developed augmentation strategies, such as code-switching data augmentation, to help bridge these gaps.13arXiv. CoSDA-ML: Multi-Lingual Code-Switching Data Augmentation for Zero-Shot Cross-Lingual NLP Techniques like feature aggregation across layers have also been explored to squeeze more cross-lingual capability out of mBERT.14arXiv. Feature Aggregation in Zero-Shot Cross-Lingual Transfer Using Multilingual BERT
Domain-Specific Versions
General-purpose BERT, trained on books and Wikipedia articles, does not always perform well in specialized fields where the vocabulary and writing style differ substantially from general text. This realization has produced a long list of domain-adapted BERT models. BioBERT was trained on biomedical literature. SciBERT targeted scientific papers. ClinicalBERT focused on clinical notes from hospitals. FinBERT was trained on financial texts. The pattern is straightforward: take the BERT architecture and either pre-train from scratch on domain-specific text, or take a general BERT and continue pre-training on specialized data.
Some domain adaptations go further by building specialized vocabularies. One recent project created bilingual BERT models for Korean-English clinical text by building a custom vocabulary of 45,000 tokens tailored to medical terminology in both languages, starting from domain-specific BERT foundations like KM-BERT for Korean medical text and BioBERT for English biomedical text.15PubMed Central. Domain and Language adaptive pre-training of BERT models for Korean-English bilingual clinical text analysis This kind of vocabulary customization can make a real difference, because BERT’s default tokenizer often splits specialized terms into nonsensical fragments.
Vulnerability to Adversarial Attacks
BERT-based text classifiers can be surprisingly easy to fool. Adversarial examples, small changes to input text that are barely noticeable to a human reader, can cause the model to misclassify text entirely.16ACL Anthology. BAE: BERT-based Adversarial Examples for Text Classification These attacks typically work by swapping words for carefully chosen synonyms or making subtle character-level changes. Research has successfully attacked BERT along with other widely used model types on fundamental tasks like text classification and textual entailment.17Proceedings of the AAAI Conference on Artificial Intelligence. Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment
One finding that stands out is that fine-tuned BERT models can actually be more susceptible to synonym-based attacks than simpler neural network architectures. Work on Arabic text classification found that BERT was easier to fool with synonym substitutions than conventional convolutional and recurrent neural networks.18ACL Anthology. Arabic Synonym BERT-based Adversarial Examples for Text Classification This is somewhat counterintuitive: you might expect a larger, more sophisticated model to be harder to trick, but BERT’s sensitivity to context means that strategic word swaps can cascade through its attention layers in ways that simpler models’ shallower processing does not amplify.
Bias in BERT’s Representations
Because BERT learns from the text it is trained on, it absorbs whatever biases exist in that text. Studies have found that BERT’s word embeddings carry measurable gender bias that propagates into downstream tasks. When researchers trained simple predictors on top of BERT’s embeddings for tasks like emotion and sentiment intensity prediction, the outputs showed significant dependence on gender-specific words and phrases, even when the task design meant the model should have ignored gender entirely.19arXiv. Investigating Gender Bias in BERT
The picture gets more complicated when you consider how different demographic dimensions interact. Analysis of BERT’s representations of given names revealed that gender and race are interdependent in the model’s embeddings, meaning you cannot fully understand BERT’s gender bias without also considering racial bias and vice versa.20ACL Anthology. Interdependencies of Gender and Race in Contextualized Word Embeddings This matters for anyone deploying BERT in applications that affect people, like resume screening, content moderation, or customer service automation. Debiasing along a single axis might leave intersectional biases intact, a problem that does not have a simple fix.
The 512-Token Ceiling
BERT has a fixed maximum input length of 512 tokens, which translates to roughly 300 to 400 English words depending on the text. Any input longer than that gets truncated. For short tasks like classifying a tweet or extracting entities from a product description, this is fine. For longer documents like legal contracts, research papers, or book chapters, it becomes a genuine limitation.
Several models have tried to address this. Longformer modified BERT’s attention mechanism to handle much longer sequences. BigBird took a similar approach with sparse attention patterns. But research on long-text classification suggests that models specifically designed for longer inputs do not always outperform standard approaches. Experiments comparing models including Longformer across five languages on a policy topic classification task found no particular advantage for the long-context specialist.21arXiv. Beyond Token Limits: Assessing Language Model Performance on Long Text Classification In many real-world scenarios, the most important information in a document appears near the beginning, so truncation hurts less than you might expect. For cases where the critical information is spread throughout a long document, chunking strategies that split the text into overlapping windows and aggregate predictions remain a common workaround.
Where BERT Still Gets Used
Despite being overshadowed in public attention by large generative models, BERT-style architectures remain deeply embedded in production systems. Search engines use BERT and its descendants to understand queries and match them to relevant documents. E-commerce platforms use it to classify products, detect spam reviews, and power recommendation systems. Healthcare organizations use domain-adapted versions to extract information from clinical notes. Legal technology companies use it to classify and search contracts.
The reason is partly practical: BERT-style models are well-understood, relatively small, fast to run, and excellent at classification tasks. A fine-tuned BERT model running on a single GPU can classify thousands of text inputs per second, which matters when you are processing millions of customer queries daily. Generative models can do many of the same tasks, but they are orders of magnitude more expensive to run and often overkill when all you need is a label or a score rather than generated text. For organizations that need reliable, fast, and cost-effective text understanding, BERT and its variants remain a sensible choice rather than a legacy technology waiting to be replaced.

