An AI reading layer for research papers

Read faster.Verify everything.

Summaries and margin notes sit on the original PDF. Click any point to open the exact sentence, equation, figure, or table behind it.

No account needed.Try an exampleAttention Is All You NeedBERTGPT-3
Paperlayerhigh

arXiv 1810.04805v2· 16 pp.

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Jacob Devlin, Ming-Wei Chang +2

Highlights69
Claims11Results21Methods13Problems1Caveats6Context17
Balanced
38%

TL;DR

Standard unidirectional language models limit the bidirectional context available during pre-training, especially for sentence- and token-level tasks. BERT addresses this with masked-language-model pre-training for deep bidirectional representations and advances state-of-the-art performance across eleven NLP tasks.

Problem

Standard unidirectional language models restrict pre-training architectures and limit context for sentence-level and token-level tasks.

Method

BERT pre-trains deep bidirectional representations using masked language modeling and next-sentence prediction on unlabeled text.

Results

80.5 GLUE score versus 72.8 for OpenAI GPT, with a 4.6% absolute accuracy improvement on MNLI.

Takeaways & Limitations

BERT reduces the need for heavily engineered task-specific architectures and successfully tackles a broad set of NLP tasks with one pre-trained model.
Masked-language-model pre-training creates a mismatch with fine-tuning because the [MASK] token does not appear during fine-tuning.

1 Introduction

BERT addresses the limits of unidirectional language-model pre-training by learning deep bidirectional representations from unlabeled text. Its pre-trained representations support strong performance across sentence-level and token-level tasks with minimal task-specific architecture.

BERT’s approach
Bidirectional pre-training lets representations fuse context from both directions, addressing a key limitation of standard unidirectional language models.
The limitation is especially consequential for token-level tasks such as question answering, where both directions provide crucial context.
BERT uses masked language modeling to predict randomly masked tokens from both left and right context.
This objective enables a deep bidirectional Transformer rather than a strictly left-to-right model.
BERT combines masked language modeling with next sentence prediction to jointly pre-train text-pair representations.

2 Related Work

Earlier NLP systems transferred information through word, sentence, or contextual representations, using either feature-based integration or fine-tuning. BERT builds on this transfer-learning tradition while replacing shallow or unidirectional context with deeply bidirectional representations.

Representation pre-training
Pre-training general language representations has a long history spanning non-neural and neural word-embedding methods.
These methods established the value of learning representations from unlabeled text before downstream training.
Sentence representations
Prior sentence-representation methods trained models to rank candidate next sentences, generate sentence-conditioned text, or denoise corrupted inputs.
Feature-based approaches
ELMo creates contextual token features by concatenating representations from separate left-to-right and right-to-left language models.
Its feature-based representations improve several benchmarks, including question answering, sentiment analysis, and named entity recognition.

3 BERT

BERT is a bidirectional Transformer encoder designed to transfer one broadly pretrained representation across downstream NLP tasks. It combines masked language modeling and next sentence prediction with a unified fine-tuning interface.

Model Architecture
BERT uses a multi-layer bidirectional Transformer encoder based on the original Transformer implementation.
The reported BERTBASE and BERTLARGE configurations contain 110M and 340M total parameters, respectively.
Input/Output Representations
BERT represents single sentences and sentence pairs in one token sequence using [CLS], [SEP], and learned sentence-segment embeddings.
The final [CLS] hidden vector serves as the aggregate sequence representation for classification tasks.
Each token representation sums its token, segment, and position embeddings.
Pre-training Tasks
During pre-training, BERT masks 15% of WordPiece tokens and predicts only the masked words from their contextual hidden vectors.

4 Experiments

BERT is evaluated on eleven NLP tasks spanning language understanding, question answering, and commonsense inference. Across these benchmarks, larger BERT models substantially improve on prior systems, though some comparisons use augmented data or ensembles.

GLUE
80.5 is the BERTLARGE GLUE leaderboard score, compared with 72.8 for OpenAI GPT.
BERTBASE and BERTLARGE obtain 4.5% and 7.0% average accuracy improvements over the prior state of the art, respectively.
4.6% is BERT’s absolute accuracy improvement on MNLI over the prior state of the art.
MNLI is described as the largest and most widely reported GLUE task.
SQuAD v1.1
+1.5 F1 is the ensemble improvement over the top SQuAD v1.1 leaderboard system, while the single system improves by +1.3 F1.

5 Ablation Studies

The ablations isolate why BERT works: deep bidirectionality, model scale, and flexible use of learned representations each improve downstream performance.

Pre-training tasks
Left-to-right pre-training performs worse than masked-language-model pre-training on every evaluated task, with large drops on MRPC and SQuAD.
The left-context-only model also preserves the left-only constraint during fine-tuning to avoid a pre-train/ fine-tune mismatch.
Removing next sentence prediction hurts performance on QNLI, MNLI, and SQuAD 1.1.
The comparison uses the same BERTBASE architecture and pre-training data.
A BiLSTM improves the left-to-right model on SQuAD but remains far worse than pre-trained bidirectional models and hurts GLUE performance.

6 Conclusion

BERT extends the benefits of unsupervised language-model pre-training from deep unidirectional systems to deep bidirectional representations. This lets one pre-trained model address a broad range of NLP tasks.

Conclusion
BERT’s major contribution is generalizing transfer learning to deep bidirectional architectures.
The model uses rich unsupervised pre-training to support language understanding tasks.

Appendix for “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”

The appendix collects supplementary implementation details, experimental details, and additional ablation studies for BERT.

Appendix contents
Appendix A contains additional implementation details for BERT.
Appendix B provides additional details for the experiments.

A.1 Illustration of the Pre-training Tasks

BERT’s pre-training combines masked language modeling with next sentence prediction to learn bidirectional token and sentence-pair representations. The masking procedure supplies contextual targets while avoiding trivial self-prediction.

Masked LM
The masking procedure forces the Transformer encoder to maintain contextual representations for every input token.

A.2 Pre-training Procedure

BERT pre-training constructs paired token sequences, applies masking after WordPiece tokenization, and optimizes masked language modeling alongside next sentence prediction. Training uses long sequences, large batches, and scheduled optimization over a large corpus.

Input construction
Each input combines two sampled text spans with a combined length of at most 512 tokens.
The spans are called “sentences” even though they are usually longer than ordinary sentences.
The second span is the actual successor of the first 50% of the time and a random sentence 50% of the time.
This supplies positive and negative examples for next sentence prediction.
Masking
The masking procedure uses a uniform 15% rate after WordPiece tokenization, without special treatment for partial word pieces.

B Detailed Experimental Setup

The experiments evaluate BERT across GLUE tasks and compare its fine-tuning setup with established benchmarks. The setup also documents dataset definitions and important exclusions or reporting constraints.

Benchmarks
GLUE includes sentence-pair inference, question-sentence inference, sentiment, acceptability, similarity, paraphrase, and related classification tasks.
MNLI predicts entailment, contradiction, or neutrality for sentence pairs.
MNLI is a large-scale crowdsourced entailment classification task over sentence pairs.
The second sentence is classified as entailment, contradiction, or neutral relative to the first.
QNLI converts SQuAD into binary classification of whether a question-sentence pair contains the answer.

C.1 Effect of Number of Training Steps

Longer pre-training improves downstream accuracy, and masked language modeling reaches stronger absolute accuracy despite slightly slower convergence. The experiments directly compare training duration and objectives under matched conditions.

Training duration
Almost 1.0% additional MNLI accuracy is achieved by BERTBASE after 1M training steps compared with 500k steps.
This supports using an extensive pre-training schedule for high fine-tuning accuracy.
Objective comparison
The MLM model converges slightly slower than the LTR model because it predicts only 15% of words per batch.

C.2 Ablation for Different Masking Procedures

This ablation tests how different masking procedures affect BERT during masked-language-model pre-training. The results show that fine-tuning is robust, while feature-based NER is more sensitive to the pre-training masking scheme.

Ablation setup
The ablation evaluates the effect of different masking strategies used for BERT’s masked language model objective.
The study reports development results for MNLI and NER, using both fine-tuning and feature-based approaches for NER.
Masking strategies
MASK replaces the target with [MASK], SAME leaves it unchanged, and RND substitutes a random token.
BERT uses an 80%, 10%, 10% mixture of MASK, SAME, and RND strategies during MLM pre-training.

Generated with AI · May be inaccurate. Verify important details in the paper.

Trace shared pre-trained parameters into separate fully fine-tuned task models; compare unchanged core architecture with task-specific output layers.
Figure/table guide
Model architecture: a multi-layer bidirectional Transformer encoder.
Definition
[MASK] creates a pre-training/ fine-tuning mismatch; replacement variants only mitigate it.
Caveat
WordPiece embeddings: subword embeddings from a 30,000-token vocabulary.
Definition
Masked LM: predict randomly masked input tokens from their final hidden vectors.
Definition
Next Sentence Prediction: classify sentence B as the actual successor of A or a random sentence.
Definition
Compare token, segment, and position embedding components; their sum forms each input embedding.
Figure/table guide
Compare BERT with OpenAI GPT across task rows; check each row’s metric, training-set size, and single-model status.
Figure/table guide
80.5 GLUE versus GPT’s 72.8 shows bidirectional pre-training transfers beyond matched architecture size.
Why this matters
+1.3 F1 as a single system shows the gain is not dependent on ensembling.
Why this matters

The Paperlayer reader opened on “BERT” in Summary view: the source-linked summary on the left and the paper on the right with highlights and margin notes, joined by a line from a summary point to its supporting passage.

The reading layer

Start with the main argument. Go deeper whenever you need.

Highlights come in five kinds: claims, methods, results, problems, and caveats. Drag the density slider to set how many of them you see.

4 of 9 sentences highlighted

Balanced

1 Introduction

Recent work has achieved significant improvements in computational efficiency through factorization tricks [21] and conditional computation [32], while also improving model performance in case of the latter. The fundamental constraint of sequential computation, however, remains.

Attention mechanisms have become an integral part of compelling sequence modeling and transduction models in various tasks, allowing modeling of dependencies without regard to their distance in the input or output sequences [2, 19]. In all but a few cases [27], however, such attention mechanisms are used in conjunction with a recurrent network.

In this work we propose the Transformer, a model architecture eschewing recurrence and instead relying entirely on an attention mechanism to draw global dependencies between input and output. The Transformer allows for significantly more parallelization and can reach a new state of the art in translation quality after being trained for as little as twelve hours on eight P100 GPUs.

7 Conclusion

For translation tasks, the Transformer can be trained significantly faster than architectures based on recurrent or convolutional layers. On both WMT 2014 English-to-German and WMT 2014 English-to-French translation tasks, we achieve a new state of the art. In the former task our best model outperforms even all previously reported ensembles.

In the reader

Three questions every paper raises.

Each one is answered without leaving the PDF.

Is this paper worth my time?

Summary with source links. A concise guide to the paper: its problem, idea, results, and caveats. Every point links to the evidence that supports it.

Decide in minutes whether it deserves a deep read.

What does this part actually mean?

Margin notes. Short notes sit in the margin where you need them. They define terms, unpack equations, and explain figures in context.

Keep reading instead of leaving the page to search.

Where's the catch?

Problems and caveats, marked. Limitations and open problems get their own highlight color, so the weak points are as easy to find as the results.

See the limitations before you build on the results.

Also in the reader

Highlight densityFocused

Choose how much of the paper stands out.

Reading languageEnglish

Read the summary and notes in Korean, Japanese, Chinese, Spanish, French, or German.

Keyboard navigationJK

Move through the paper without a mouse.

Paper chatSoon

Ask the paper a question and get answers with numbered citations.

Open your next paper.

Find a paper by title or author, or paste an arXiv link. Start reading the original PDF while Paperlayer prepares its summary and notes.

Or change the domain: arxiv.org/abs/… becomes paperlayer.ai/abs/…