Source-linked AI summary
FrameBench:A Language Understanding Benchmark Based on Frame Semantics
Chihiro Yano, Ryohei Sasano
TL;DR
It remains unclear whether strong LLM performance reflects the implicit, context-dependent frame-semantic enrichment humans use when interpreting language. FrameBench addresses this gap with FrameNet-grounded four-choice benchmarks in English and Japanese, generated and validated with native-speaker judgments. Across evaluations, performance varies with model scale, reasoning mode, and language, and several frontier models reach or exceed the human reference score.
Problem
The central gap is whether LLMs can reliably perform implicit, context-dependent frame-semantic enrichment, which broad downstream-task benchmarks do not directly test.
Method
FrameBench uses FrameNet resources and LLM-based generation to build English and Japanese four-choice tasks, followed by native-speaker validation.
Results
Several frontier models reached or exceeded the human reference score, while performance varied substantially with model scale, reasoning mode, and language.
Takeaways & Limitations
FrameBench provides a resource for evaluating frame-semantic interpretation across two typologically distant languages.
Takeaways & Limitations
FrameBench evaluates context-dependent frame-semantic discrimination in a four-choice setting, while its human reference scores may overestimate performance because they reuse validation judgments.
Abstract
from arXiv · showhide
In frame semantics, sentence comprehension is assumed to proceed by relating lexical meaning to background knowledge called semantic frames, thereby enabling readers to implicitly enrich the text with unstated information. Recent large language models (LLMs) have achieved strong performance across a wide range of downstream tasks. However, it remains unclear whether they can reproduce the kinds of implicit enrichment that humans naturally make during comprehension. To address this question, we introduce FrameBench, a benchmark grounded in frame semantics. FrameBench consists of multiple-choice questions that test whether models distinguish the frames evoked by the same verb across contexts. We construct the benchmark for English and Japanese using FrameNet-style resources and a generation-and-verification pipeline with native-speaker judgments. Our experiments on a diverse set of models reveal challenges for small models, while several large models surpass the human reference scores. We release the constructed FrameBench dataset and the code for dataset construction and evaluation at https://github.com/SasanoLab/FrameBench.
1 Introduction
FrameBench addresses whether LLMs can perform the implicit, context-dependent frame-semantic enrichment that humans use to interpret the same verb differently across contexts. It introduces English and Japanese benchmarks generated and validated with FrameNet-style resources and native-speaker judgments.
- The same verb can evoke different situations across contexts, causing readers to infer unstated information such as an employment relation.“Left” can describe physical departure from a location or quitting, depending on context.
- Existing broad benchmark suites do not directly test whether LLMs reliably perform context-dependent frame-semantic enrichment.
- FrameBench evaluates these distinctions indirectly through multiple-choice questions about situations implied by sentences rather than explicit frame-label prediction.
- The benchmark covers English and Japanese, using FrameNet resources, LLM-based generation, and native-speaker validation.
- FrameBench requires context-sensitive interpretation because its example item cannot be solved by matching the shared predicate alone.
- The study reports that performance varies with model scale, reasoning mode, and language, while high-performing models can still face fine-grained frame-semantic challenges.
2 Related Work
Related work includes broad LLM evaluations, context-based word-sense tasks, FrameNet resources, and LLM-based frame-semantic analysis. FrameBench differs by testing context-dependent frame discrimination grounded in broader conceptual situations rather than dictionary senses or explicit frame induction.
- Broad LLM evaluations commonly bundle diverse downstream tasks, while WSD and WiC more directly test whether words have the same meaning across contexts.
- Unlike standard WSD, FrameBench is grounded in frame semantics and captures broader conceptual situations.
- FrameNet is a manually annotated lexical knowledge base containing evoked frames, participant roles, and structured relations between frames.
- Frame-semantic resources extend beyond English, and recent work applies LLMs to frame-semantic analysis and FrameNet annotation.
- NutFrame evaluates explicit frame-semantic structure induction, whereas FrameBench evaluates discrimination of context-dependent interpretations.
3 FrameBench: Task Design and Dataset Construction
FrameBench is a four-choice task that contrasts sentences evoking different frames for the same polysemous verb. Its construction combines FrameNet-guided generation, sentence-pair expansion, native-speaker validation, and filtering into English and Japanese evaluation sets.
- Task Design: Each entry uses a question targeting one frame and four candidate sentences: two designed for that frame and two for a contrasting frame.
- Task Design: Models choose Sentence A, Sentence B, Both Sentences, or Neither Sentence, although each evaluation pair has exactly one correct sentence.
- Construction Pipeline: The pipeline extracts polysemous verbs and frame pairs from FrameNet-style resources, then uses an LLM to generate a question and base sentence pair conditioned on frame information.
- Construction Pipeline: Additional extended-pair sentences preserve each original sentence’s frame and ground-truth label while increasing evaluation diversity.
- Human Validation: Native speakers judged item correctness and description acceptability, and the evaluation retained entries meeting minimum thresholds on both judgments.
- Dataset Statistics: The retained main evaluation set contains 599 English entries and 500 Japanese entries.
4 Evaluation on FrameBench
FrameBench evaluates LLMs on English and Japanese context-dependent frame-semantic interpretation, using controlled evaluation protocols and comparisons with human and benchmark scores. Performance generally improves with model scale and reasoning, with several large models exceeding human references, while cross-lingual and dataset-construction effects remain important caveats.
- Evaluation setup: FrameBench evaluates closed-source, open-weight, Japanese-oriented, and multimodal models on English and Japanese benchmark versions, with comparisons to established LLM benchmarks.Each item is tested in base and extended forms, with both sentence orders and five prompt templates per language.
- Evaluation setup: Human reference scores may be upwardly biased because the evaluation subset retains entries judged correct by at least two annotators.An independently administered human evaluation could produce a lower score.
- Overall results: Across both languages, FrameBench performance increases with model scale and reasoning, and several high-performing models surpass the human reference scores.The scale and reasoning trends broadly match comparison benchmarks, while reasoning gains are often larger on FrameBench.
- English results: 99.3 is the highest English score, achieved by Gemini 3.1 Pro, while Qwen3.5-2B rises from 48.5 to 80.3 with reasoning.GPT-5 reaches 99.1, and the strongest listed English scores exceed the human reference of 96.3.
- Robustness: Using Gemini 3.1 Pro rather than GPT-5 for constructing 100 additional English entries produced a similar overall performance pattern.This robustness check addresses possible effects of the dataset-construction model.
- Japanese results: 99.1 is the highest Japanese score, achieved by Gemma 4 31B, while Qwen3.5-2B rises from 30.0 to 60.7 with reasoning.The strongest listed Japanese scores exceed the human reference of 94.9.
- Japanese-oriented models: Qwen3 Swallow models perform comparably to or slightly below corresponding Qwen3 models on FrameBench despite outperforming them on JamC-QA.Qwen3 Swallow 8B scores 84.0 versus 85.0 for Qwen3-8B, and Swallow 32B scores 93.2 versus 93.9 for Qwen3-32B.
- Cross-lingual comparison: Japanese results separate models more clearly than English results, with weaker models showing larger cross-lingual drops.GPT-5 nano falls from 93.5 to 83.9, whereas Gemma 4 31B changes from 98.6 to 99.1.
5 Analysis
The analysis examines pair type, sentence position, error type, multimodality, and challenging examples to identify why FrameBench remains difficult. Results point to lexical overlap, order sensitivity, borderline frame distinctions, and fine-grained semantic cues as important factors.
- Pair type: Lexical similarity makes base pairs harder for English LLMs, whereas this pair-type effect is weaker in Japanese.The Japanese human score gap between pair types is 0.93 points, and LLM advantages for extended pairs are less consistent.
- Sentence position bias: Most models show an absolute performance difference when the correct sentence appears first versus second, supporting evaluation in both orders.Swapped-order testing controls for item difficulty while mitigating sentence-order sensitivity.
- Error tendencies: For models above 85% accuracy, opposite-frame errors become rare, leaving over-selection and under-selection as the dominant error patterns.This indicates that high-performing models mainly struggle with borderline frame-semantic distinctions rather than choosing the opposite frame.
- Impact of multimodality: Multimodal models outperform corresponding language models in most matched FrameBench comparisons, but the effect of visual grounding cannot be isolated from other training differences.Some FrameBench gains occur alongside lower MMLU-Pro scores, while many multimodal models also score higher on MMLU-Pro.
- Case studies: Even the largest Qwen3.5 and Gemma 4 models fail examples requiring fine-grained distinctions despite strong lexical or contextual overlap.The examples include shifts between contact and injury frames and a ceremonial distinction associated with christen.
- Case studies: Question 3 remains challenging for smaller models because distinguishing related senses within a close semantic domain requires moderate semantic sensitivity.Question 4 is solved by all models when salient contextual cues place the two meanings in clearly different domains.
6 Conclusion
FrameBench provides a FrameNet-grounded resource for evaluating whether LLMs distinguish frames evoked by the same verb across contexts in English and Japanese. Experiments show that frontier-model performance can reach or exceed human reference scores, while results vary with model scale, reasoning mode, and language.
- FrameBench evaluates context-dependent frame-semantic interpretation across English and Japanese using the same verb in different contexts.The benchmark is grounded in FrameNet resources and uses generation, verification, and native-speaker judgments.
- Several frontier models reached or exceeded the human reference score, while performance varied substantially with model scale, reasoning mode, and language.The experiments covered a wide range of LLMs and included behavioral analyses of model performance.
- The analyses characterized model behavior by sentence-pair type, sentence position, and error type, and examined whether multimodal training may support frame-semantic interpretation.
Limitations
FrameBench targets a specific form of semantic competence rather than broad language understanding, and its human reference scores were derived from the same annotators who validated the items.
- FrameBench focuses on frame-semantic discrimination for context-dependent verb interpretations in a four-choice setting.It does not directly evaluate broader language understanding, open-ended generation, or explicit frame prediction.
- The human reference scores were calculated from judgments by the same annotators who validated and filtered the benchmark items.
- Because only items answered correctly by at least two validation annotators were used, the reported human scores may overestimate performance relative to independent annotation.
Ethical considerations
The benchmark construction and evaluation involved human annotation in English and Japanese, with explicit acceptability criteria, agreement reporting, and controlled model-evaluation procedures.
- English and Japanese benchmark items underwent human annotation for correctness and acceptability.English annotations used three expert native English annotators, while Japanese evaluation used three native Japanese-speaking university students.
- English acceptability used a three-way Unacceptable, Acceptable, Natural scale, whereas Japanese acceptability used a binary naturalness label.
- For cross-lingual analysis, English Natural judgments were mapped to positive and Acceptable or Unacceptable judgments to negative.The mapping was intended to use a stricter notion of linguistic naturalness and produce a more informative label distribution.
- The study reports positive-label rates, Fleiss’ κ, and 3/3 agreement rates, while using a strict human-validation criterion for the main experiments.The additional statistics support interpretation because κ is sensitive to skewed label distributions.
- Evaluation prompts randomized the mapping between answer options and numbers to reduce the influence of specific choice-number output probabilities.
B.5 Details of the Frame Identification Evaluation
The frame-identification evaluation tests whether models can select an evoked frame when candidate frame names and definitions are explicitly provided. It uses two candidate frames plus a negative option and evaluates 400 instances derived from 200 verb–frame pairs.
- Unlike FrameBench, this evaluation explicitly provides candidate frame names and definitions from the relevant English or Japanese FrameNet resource.It tests frame identification when the relevant candidate frames are given.
- Each instance contains a target frame-evoking verb, two candidate frames that the verb can evoke, and the negative option “Neither frame is evoked.”Answer-option order is randomized.
- The evaluation reuses verb–frame pairs from FrameBench and samples one sentence for each candidate frame for every selected pair.When a verb appeared in multiple FrameBench pairs, one pair was randomly selected to avoid redundancy.
- 400 test instances were evaluated in a 3-shot setting from 200 verb–frame pairs, with two independently evaluated target sentences per pair.
C Effect of the LLM Used for Dataset Construction
Using a different LLM for dataset construction largely preserves FrameBench’s overall model-performance pattern, although same-family models tend to perform slightly better. The comparison reports a very high rank correlation between alternative and original entries.
- The construction comparison used 100 additional English candidate entries generated with Gemini 3.1 Pro from verb–frame pairs also used in the original GPT-5-based construction.
- Using Gemini 3.1 Pro instead of GPT-5 for dataset construction preserved the overall FrameBench performance pattern, with ρ = 0.986.The additional entries were evaluated with the same protocol as the main experiment, but were not human-validated or filtered.
- Models from the same family as the LLM used to construct the benchmark tended to perform slightly better.
D Analysis on the Japanese Subset
The Japanese subset generally matches the English analysis for correct-sentence position and error type but differs in sentence-pair type. Extended pairs provide only a small advantage over base pairs in Japanese.
- Japanese results are generally consistent with English for correct sentence position and error type, but differ in sentence-pair type.
- In Japanese, the extended-pair advantage over base pairs is weaker than in English.The corresponding Japanese gap is 0.93 points, compared with a 2.57-point human-score gap in English.
- The human-score gap between extended and base sentence pairs is 2.57 points in English and 0.93 points in Japanese.