Standard unidirectional language models limit the bidirectional context available during pre-training, especially for sentence- and token-level tasks. BERT addresses this with masked-language-model pre-training for deep bidirectional representations and advances state-of-the-art performance across eleven NLP tasks.
Problem
Standard unidirectional language models restrict pre-training architectures and limit context for sentence-level and token-level tasks.
↗Method
BERT pre-trains deep bidirectional representations using masked language modeling and next-sentence prediction on unlabeled text.
↗Results
80.5 GLUE score versus 72.8 for OpenAI GPT, with a 4.6% absolute accuracy improvement on MNLI.
↗Takeaways & Limitations
BERT reduces the need for heavily engineered task-specific architectures and successfully tackles a broad set of NLP tasks with one pre-trained model.
↗Masked-language-model pre-training creates a mismatch with fine-tuning because the [MASK] token does not appear during fine-tuning.
↗1 Introduction
BERT addresses the limits of unidirectional language-model pre-training by learning deep bidirectional representations from unlabeled text. Its pre-trained representations support strong performance across sentence-level and token-level tasks with minimal task-specific architecture.
BERT’s approach
Bidirectional pre-training lets representations fuse context from both directions, addressing a key limitation of standard unidirectional language models.The limitation is especially consequential for token-level tasks such as question answering, where both directions provide crucial context.
↗BERT uses masked language modeling to predict randomly masked tokens from both left and right context.This objective enables a deep bidirectional Transformer rather than a strictly left-to-right model.
↗BERT combines masked language modeling with next sentence prediction to jointly pre-train text-pair representations.
↗2 Related Work
Earlier NLP systems transferred information through word, sentence, or contextual representations, using either feature-based integration or fine-tuning. BERT builds on this transfer-learning tradition while replacing shallow or unidirectional context with deeply bidirectional representations.
Representation pre-training
Pre-training general language representations has a long history spanning non-neural and neural word-embedding methods.These methods established the value of learning representations from unlabeled text before downstream training.
↗Sentence representations
Prior sentence-representation methods trained models to rank candidate next sentences, generate sentence-conditioned text, or denoise corrupted inputs.
↗Feature-based approaches
ELMo creates contextual token features by concatenating representations from separate left-to-right and right-to-left language models.Its feature-based representations improve several benchmarks, including question answering, sentiment analysis, and named entity recognition.
↗3 BERT
BERT is a bidirectional Transformer encoder designed to transfer one broadly pretrained representation across downstream NLP tasks. It combines masked language modeling and next sentence prediction with a unified fine-tuning interface.
Model Architecture
BERT uses a multi-layer bidirectional Transformer encoder based on the original Transformer implementation.The reported BERTBASE and BERTLARGE configurations contain 110M and 340M total parameters, respectively.
↗Input/Output Representations
BERT represents single sentences and sentence pairs in one token sequence using [CLS], [SEP], and learned sentence-segment embeddings.The final [CLS] hidden vector serves as the aggregate sequence representation for classification tasks.
↗Each token representation sums its token, segment, and position embeddings.
↗Pre-training Tasks
During pre-training, BERT masks 15% of WordPiece tokens and predicts only the masked words from their contextual hidden vectors.
↗4 Experiments
BERT is evaluated on eleven NLP tasks spanning language understanding, question answering, and commonsense inference. Across these benchmarks, larger BERT models substantially improve on prior systems, though some comparisons use augmented data or ensembles.
GLUE
80.5 is the BERTLARGE GLUE leaderboard score, compared with 72.8 for OpenAI GPT.BERTBASE and BERTLARGE obtain 4.5% and 7.0% average accuracy improvements over the prior state of the art, respectively.
↗4.6% is BERT’s absolute accuracy improvement on MNLI over the prior state of the art.MNLI is described as the largest and most widely reported GLUE task.
↗SQuAD v1.1
+1.5 F1 is the ensemble improvement over the top SQuAD v1.1 leaderboard system, while the single system improves by +1.3 F1.
↗5 Ablation Studies
The ablations isolate why BERT works: deep bidirectionality, model scale, and flexible use of learned representations each improve downstream performance.
Pre-training tasks
Left-to-right pre-training performs worse than masked-language-model pre-training on every evaluated task, with large drops on MRPC and SQuAD.The left-context-only model also preserves the left-only constraint during fine-tuning to avoid a pre-train/ fine-tune mismatch.
↗Removing next sentence prediction hurts performance on QNLI, MNLI, and SQuAD 1.1.The comparison uses the same BERTBASE architecture and pre-training data.
↗A BiLSTM improves the left-to-right model on SQuAD but remains far worse than pre-trained bidirectional models and hurts GLUE performance.
↗6 Conclusion
BERT extends the benefits of unsupervised language-model pre-training from deep unidirectional systems to deep bidirectional representations. This lets one pre-trained model address a broad range of NLP tasks.
Conclusion
BERT’s major contribution is generalizing transfer learning to deep bidirectional architectures.The model uses rich unsupervised pre-training to support language understanding tasks.
↗Appendix for “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”
The appendix collects supplementary implementation details, experimental details, and additional ablation studies for BERT.
Appendix contents
Appendix A contains additional implementation details for BERT.
↗Appendix B provides additional details for the experiments.
↗A.1 Illustration of the Pre-training Tasks
BERT’s pre-training combines masked language modeling with next sentence prediction to learn bidirectional token and sentence-pair representations. The masking procedure supplies contextual targets while avoiding trivial self-prediction.
Masked LM
The masking procedure forces the Transformer encoder to maintain contextual representations for every input token.
↗A.2 Pre-training Procedure
BERT pre-training constructs paired token sequences, applies masking after WordPiece tokenization, and optimizes masked language modeling alongside next sentence prediction. Training uses long sequences, large batches, and scheduled optimization over a large corpus.
Input construction
Each input combines two sampled text spans with a combined length of at most 512 tokens.The spans are called “sentences” even though they are usually longer than ordinary sentences.
↗The second span is the actual successor of the first 50% of the time and a random sentence 50% of the time.This supplies positive and negative examples for next sentence prediction.
↗Masking
The masking procedure uses a uniform 15% rate after WordPiece tokenization, without special treatment for partial word pieces.
↗B Detailed Experimental Setup
The experiments evaluate BERT across GLUE tasks and compare its fine-tuning setup with established benchmarks. The setup also documents dataset definitions and important exclusions or reporting constraints.
Benchmarks
GLUE includes sentence-pair inference, question-sentence inference, sentiment, acceptability, similarity, paraphrase, and related classification tasks.MNLI predicts entailment, contradiction, or neutrality for sentence pairs.
↗MNLI is a large-scale crowdsourced entailment classification task over sentence pairs.The second sentence is classified as entailment, contradiction, or neutral relative to the first.
↗QNLI converts SQuAD into binary classification of whether a question-sentence pair contains the answer.
↗C.1 Effect of Number of Training Steps
Longer pre-training improves downstream accuracy, and masked language modeling reaches stronger absolute accuracy despite slightly slower convergence. The experiments directly compare training duration and objectives under matched conditions.
Training duration
Almost 1.0% additional MNLI accuracy is achieved by BERTBASE after 1M training steps compared with 500k steps.This supports using an extensive pre-training schedule for high fine-tuning accuracy.
↗Objective comparison
The MLM model converges slightly slower than the LTR model because it predicts only 15% of words per batch.
↗C.2 Ablation for Different Masking Procedures
This ablation tests how different masking procedures affect BERT during masked-language-model pre-training. The results show that fine-tuning is robust, while feature-based NER is more sensitive to the pre-training masking scheme.
Ablation setup
The ablation evaluates the effect of different masking strategies used for BERT’s masked language model objective.The study reports development results for MNLI and NER, using both fine-tuning and feature-based approaches for NER.
↗Masking strategies
MASK replaces the target with [MASK], SAME leaves it unchanged, and RND substitutes a random token.
↗BERT uses an 80%, 10%, 10% mixture of MASK, SAME, and RND strategies during MLM pre-training.
↗Generated with AI · May be inaccurate. Verify important details in the paper.