Source-linked AI summary

LongBench Pro: A More Realistic and Comprehensive Bilingual Long-Context Evaluation Benchmark

Ziyang Chen, Xing Wu, Junlong Jia, Chaochen Gao, Qi Fu, Debing Zhang, Songlin Hu

arXiv:2601.02872v1cs.CLcs.AI

TL;DR

Existing long-context benchmarks lack a scalable combination of realism, breadth, and extreme-length coverage. LongBench Pro addresses this gap with a bilingual benchmark and human–model construction pipeline, and its evaluation of 46 models shows that context optimization, effective-length alignment, and native reasoning training are central to performance.

  • Problem

    Existing long-context benchmarks trade off realistic complexity against scalable construction as context lengths and model capabilities expand.

  • Method

    LongBench Pro combines 1,500 bilingual samples from authentic documents with model drafting, expert verification, and multidimensional evaluation across tasks, lengths, context requirements, and difficulty.

  • Results

    Evaluating 46 long-context LLMs shows that context optimization outperforms parameter scaling, effective context lengths often fall below claimed lengths, and native reasoning training is important for thinking gains.

  • Takeaways & Limitations

    LongBench Pro provides a realistic, comprehensive testbed for analyzing long-context understanding and comparing scaling, language, length, and reasoning effects.

  • Takeaways & Limitations

    As task length and complexity grow, human–model construction faces a tension between verification accuracy and production efficiency.

Abstract

from arXiv · show

The rapid expansion of context length in large language models (LLMs) has outpaced existing evaluation benchmarks. Current long-context benchmarks often trade off scalability and realism: synthetic tasks underrepresent real-world complexity, while fully manual annotation is costly to scale to extreme lengths and diverse scenarios. We present LongBench Pro, a more realistic and comprehensive bilingual benchmark of 1,500 naturally occurring long-context samples in English and Chinese spanning 11 primary tasks and 25 secondary tasks, with input lengths from 8k to 256k tokens. LongBench Pro supports fine-grained analysis with task-specific metrics and a multi-dimensional taxonomy of context requirement (full vs. partial dependency), length (six levels), and difficulty (four levels calibrated by model performance). To balance quality with scalability, we propose a Human-Model Collaborative Construction pipeline: frontier LLMs draft challenging questions and reference answers, along with design rationales and solution processes, to reduce the cost of expert verification. Experts then rigorously validate correctness and refine problematic cases. Evaluating 46 widely used long-context LLMs on LongBench Pro yields three findings: (1) long-context optimization contributes more to long-context comprehension than parameter scaling; (2) effective context length is typically shorter than the claimed context length, with pronounced cross-lingual misalignment; and (3) the "thinking" paradigm helps primarily models trained with native reasoning, while mixed-thinking designs offer a promising Pareto trade-off. In summary, LongBench Pro provides a robust testbed for advancing long-context understanding.

1. Introduction

LongBench Pro addresses limitations in existing long-context benchmarks by combining realistic bilingual data, broad task coverage, and scalable expert-validated construction. Evaluations of 46 models reveal that context optimization, effective-length gaps, language alignment, and native reasoning shape long-context performance.

  • Motivation: Existing synthetic and manually annotated benchmarks increasingly underserve next-generation models as context lengths and capabilities expand.Their limitations include insufficient realism, scalability, task coverage, and difficulty at extreme lengths.
  • Benchmark: LongBench Pro contains 1,500 bilingual samples from authentic long documents, spanning 11 primary tasks, 25 secondary tasks, and 8k–256k-token inputs.Its taxonomy covers full versus partial context dependence, six length levels, and four model-calibrated difficulty levels.
  • Construction: Human-Model Collaborative Construction combines model-generated questions and answers with expert verification to improve realism, quality, and scalability.Models also provide design rationales and solution processes, while experts verify correctness and refine problematic cases.
  • Findings: Long-context optimization improves comprehension more than parameter scaling.The reported findings favor a context-optimization-first over a scale-first paradigm.
  • Findings: Models’ effective context lengths often trail their claimed lengths, while cross-lingual performance remains misaligned.These gaps persist across most long-context models, although stronger models reduce them.
  • Findings: Thinking substantially helps primarily models with native reasoning training, whereas mixed-thinking models offer a Pareto trade-off between fast response and deep reasoning.Conventional instruct models show limited or degraded gains when prompted to think.

2. Task Framework of LongBench Pro

LongBench Pro organizes long-context evaluation through 11 primary task categories and an orthogonal context-requirement dimension. Crossing task type with global versus localized dependence produces 25 secondary categories.

  • Task taxonomy: The benchmark consolidates prior formulations into 11 primary task categories covering core long-context capability dimensions.Its task taxonomy is mapped against existing benchmark tasks.
  • Context Requirement: Context Requirement measures how globally a solution must depend on the input document.It is independent of task type and distinguishes integration from localization and retrieval.
  • Context Requirement: Full-context tasks require integrating evidence dispersed across multiple distant document spans.These tasks emphasize global integration and reasoning.
  • Context Requirement: Partial-context tasks primarily rely on localized spans and emphasize localization and retrieval.They require less global dependence on the document than full-context tasks.
  • Secondary taxonomy: Crossing 11 primary categories with context requirements yields 25 secondary categories.The secondary taxonomy combines task formulation with the document-dependence dimension.

3. Construction Process of LongBench Pro

LongBench Pro constructs its benchmark from compliant natural documents, scalable model drafts, and iterative human verification. Difficulty labels and prompt standardization support consistent evaluation across tasks, languages, and lengths.

  • Document Collection: Documents come from diverse public domains and formats, balanced across languages, document settings, and six target lengths from 8k to 256k tokens.Human review excludes privacy-sensitive, copyrighted, or otherwise non-compliant content.
  • Sample Generation: Frontier LLMs draft three candidate samples per document, including questions, reference answers, design rationales, and solution processes.Candidates are aligned with target task definitions and context requirements.
  • Human Verification: Human annotators verify task alignment, context requirements, answer correctness, and difficulty, selecting or minimally editing qualifying samples.Failed cases are revised until they satisfy the criteria, or the process moves to another document.
  • Human Verification: The workflow combines model scalability with human reliability to mitigate cognitive limitations at extreme lengths and model hallucinations.Human expertise is used for verification while models provide diverse candidate components.
  • Prompting: The benchmark uses standardized non-thinking and thinking prompts, each specifying task instructions, output requirements, and examples.Both prompt types require answer elements to be presented line by line using the specified identifier.
  • Quality Control: Two annotators independently verify samples, with an additional long-context expert reviewing potential problems.Annotators inspect original answer components and model predictions to improve precision and recall.
  • Difficulty Classification: Difficulty is defined from model performance using high-, mid-, and low-performing tiers to assign Extreme, Hard, Moderate, and Easy levels.Extreme samples are answerable by at most one high-performing model; the remaining levels are assigned progressively by tier performance.
  • Difficulty Classification: The progressive difficulty scheme aligns labels with model capabilities and supports scalable cross-difficulty analysis.It is designed as a model-centric alternative to subjective human difficulty ratings.

4. Data Statistics and Validation of LongBench Pro

LongBench Pro uses a balanced factorial design across tasks, languages, and lengths, then audits sampled items for attribute and answer correctness.

  • Data Statistics: 1,500 samples result from 5 samples for each combination of 25 secondary tasks, 2 languages, and 6 length buckets.The design balances task, language, and target-length coverage.
  • Validation: A uniform audit of 300 samples checks attribute correctness and answer correctness across secondary tasks, languages, and lengths.The audit evaluates whether metadata and answers satisfy the benchmark criteria.

5. Evaluation

Evaluation across model tiers, languages, difficulty levels, lengths, and task requirements shows that long-context performance depends more on effective comprehension than nominal context capacity. Native reasoning and mixed-thinking designs help, but extreme, full-context, and aggregation tasks remain difficult.

  • 73.42, 72.61, and 69.87 are the overall scores of Gemini-2.5-Pro, GPT-5, and Claude-4-Sonnet, respectively.
  • Long-context optimization outperforms parameter scaling, with optimized Qwen3-4B scoring 45.68 versus 44.34 for Qwen3-8B, and optimized Qwen3-30B scoring 54.52 versus 51.12 for Qwen3-32B.
  • Claimed context length often exceeds effective performance: MiniMax-Text-01 claims 4M tokens but scores only 45.00, behind many 128k models.
  • Cross-lingual performance is uneven: GPT, Claude, Mistral, and Llama generally favor English, while GLM, Kimi, and MiniMax generally favor Chinese.The language gap narrows as model scale and overall capability increase.
  • Thinking gains are larger on Easy than Extreme samples; Claude-4-Sonnet improves by 15.36 points on Easy but only 4.13 on Extreme.Extreme samples combine long-context memory, integration, and reasoning, leaving substantial room for improvement.
  • Native reasoning training is crucial: conventional instruct models gain little or decline with thinking, including a 1.03-point drop for Llama-3.1-8B-Instruct.Mixed-thinking models combine fast responses with deep reasoning and can approach or surpass dedicated thinking models.
  • Performance generally declines with sample length, but Gemini-2.5-Pro remains nearly length-insensitive, scoring 71.77 at 256k versus 74.50 at 8k.The reported bottleneck shifts from reading more tokens toward handling long-range dependencies and complex logical relationships.

6. Related Works

Prior long-context benchmarks range from scalable synthetic probes to realistic human-verified tasks, but often provide narrower realism, language, or task coverage. LongBench Pro responds with fully natural bilingual documents and diverse tasks and metrics.

  • Existing benchmarks include synthetic probes, multilingual evaluations, and methodology-focused protocol suites for testing long-context retrieval, integration, and reasoning.
  • LongBench Pro differs by using fully natural long documents, bilingual English-Chinese coverage, and diverse tasks and metrics for fine-grained evaluation.

7. Conclusion and Future Work

LongBench Pro is presented as a realistic, comprehensive bilingual benchmark for long-context evaluation, with analyses spanning multiple task and model dimensions. The authors identify a remaining tension between verification accuracy and production efficiency as tasks become longer and more complex.

  • LongBench Pro evaluates 46 representative long-context LLMs across task, length, context requirement, difficulty, and language settings.
  • Human-model collaborative construction remains subject to a tension between verification accuracy and production efficiency as task length and complexity grow.The authors are exploring recursive critique to decompose verification into easier subproblems for human annotators.

A. Task Definitions

The benchmark defines tasks that test retrieval, ordering, evidence-based question answering, summarization, citation alignment, clustering, and global frequency analysis under full or partial context requirements.

  • Retrieval: Global and local retrieval tasks distinguish full-document reorganization from locating a target fragment in a specified paragraph, using NDCG@k.
  • Ordering: Timeline reconstruction and paragraph reordering test global versus local ordering with Pairwise Accuracy.
  • Question Answering: Evidence-based question answering includes multi-document integration and local-paragraph questions, both evaluated with Accuracy.
  • Summarization: Summary generation distinguishes full-text coverage from subtopic summarization and combines semantic similarity with ROUGE-L.
  • Alignment and Statistics: Other tasks align generated sentences with source locations, cluster documents by methodology, retrieve category instances, and sort terms by global frequency.

T9 Version & Code Diff Analysis

The supplied task definitions cover version-difference analysis, rule induction from document-wide or targeted examples, and long-range dialogue-state tracking, with explicit context requirements and metrics.

  • Version Differences: Dependency-aware multi-version impact analysis identifies methods whose deprecation status changes between API versions using F1.
  • Version Differences: Local version-difference analysis identifies renamed, removed, or structurally refactored fields between resource-model versions.
  • Rule Induction: Large-scale in-context rule induction infers formatting rules from a full document, while targeted induction derives rules from selected examples.
  • Dialogue Tracking: Long-range entity and commitment tracking evaluates global dialogue-state tracking, whereas short-range reference resolution queries local states and references.

B.2. Sample Verification Criteria

Sample verification systematically checks whether model-generated questions satisfy task requirements, whether answers are correct and document-grounded, and whether the questions provide sufficient challenge.

  • Verification begins by assessing compliance with the task type, context requirement, document grounding, topic relevance, and absence of unsupported facts.
  • Annotators evaluate answer correctness through grounded reasoning, internal consistency, document verifiability, and absence of hallucinated information.
  • Samples are discarded when questions diverge from the task or document, while minor issues may be corrected before further verification.
  • The process prohibits fabricated information, subjective questions, unverifiable answers, and overly simple questions while encouraging deep reasoning and cross-paragraph inference.

B.3. Sample Rewriting Criteria

LongBench Pro standardizes each sample into parallel non-thinking and thinking prompts, then verifies answers through independent human checks supported by model predictions and expert adjudication.

  • Prompt standardization: Each question uses non-thinking and thinking prompts with the same structure, differing mainly in output requirements and examples.Both formats include task description, optional supplementary content, and a mandatory output example.
  • Prompt standardization: Thinking prompts require step-by-step reasoning before the “[Answer]” identifier and specified answer elements.The answer format remains constrained to the annotator-defined elements.
  • Prompt standardization: Annotators rewrite each sample by extracting the task description, output requirement, supplementary content, and output example.The procedure then constructs fixed templates for both prompt types.
  • Answer verification: Verification combines human precision checks with model-assisted recall checks for missing reasonable answer components.Two annotators independently inspect answer correctness and consult predictions from five state-of-the-art models.
  • Answer verification: Reviewers identify document, question, and answer issues, while long-context experts adjudicate disputes and decide whether reconstruction is needed.Reconstruction may modify the document, rewrite the question, or revise the answer.

B.5. Sample Quality Evaluation Criteria

Sample quality is evaluated across five dimensions using three-expert scoring on a discrete 0, 0.5, or 1 scale, with final scores averaged across experts.

  • Evaluation dimensions: Each long-context question is scored on five evaluation dimensions, and every dimension allows scores of 0, 0.5, or 1.The dimensions assess task alignment, context dependency, difficulty, naturalness, and answer correctness.
  • Evaluation dimensions: Task alignment measures whether the question matches the predefined objective and intended task type.The benchmark contains 25 secondary tasks.
  • Evaluation dimensions: Context dependency measures whether document reliance matches the intended requirement level, such as full-document versus local-paragraph dependence.A Partial question requiring the entire document is treated as incorrect dependency alignment.
  • Evaluation dimensions: Difficulty is estimated from five-model accuracy, assigning the highest score to questions answered correctly by at most one model.Scores of 1, 0.5, and 0 correspond to 0%–20%, 40%–60%, and 80%–100% accuracy, respectively.
  • Scoring procedure: Three experts independently score each sample, and the final score for every dimension is the average of their scores.Scoring follows the criteria without subjective assumptions and requires explanations for review.

E. Truncation Length Setting

Inference truncation lengths are adjusted to preserve sufficient thinking space and avoid instability when models operate near their claimed context limits.

  • Truncation settings: Some thinking models produce outputs far beyond the default 8k limit, so selected models use 120k truncation lengths to permit 32k thinking outputs.This setting is applied to DeepSeek-V3.2, GLM-4.6, and MiniMax-M2.
  • Truncation settings: GLM-4.6 becomes unstable on near-190k samples when truncated at 190k, causing a sharp metric drop despite its claimed 198k context length.The discrepancy motivates setting its truncation length to 120k.
  • Truncation settings: Table 4 compares GLM-4.6 non-thinking scores across sample lengths and truncation lengths.The comparison is used to examine effective versus claimed context length.
  • Truncation settings: Inference parameter settings are documented for different models, with lengths measured uniformly in tokens.Table 5 distinguishes non-thinking and thinking settings and records the standardized truncation choice.
Loading 2601.02872v1…