Source-linked AI summary
Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA
Minzheng Wang, Longze Chen, Cheng Fu, Shengyi Liao, Xinghua Zhang, Bingli Wu, Haiyang Yu, Nan Xu, Lei Zhang, Run Luo, Yunshui Li, Min Yang, Fei Huang, Yongbin Li
TL;DR
Existing long-context benchmarks do not adequately evaluate realistic multi-document question answering, where evidence is distributed across relevant documents. Loong addresses this gap with a multi-domain benchmark spanning four task types and varied context lengths, and finds that powerful long-context LLMs still struggle. Its scope is limited to financial, legal, and academic domains, with high annotation costs.
Problem
Existing benchmarks lack sufficient alignment with realistic multi-document question answering, often centralizing evidence or using distracting noise texts.
Method
Loong evaluates multi-document long-context comprehension with four task categories, varied context lengths, and newly annotated question-answering instances.
Results
Even the most powerful long-context LLMs fail to achieve satisfactory performance on Loong, leaving substantial room for improvement.
Takeaways & Limitations
Loong provides a realistic benchmark for assessing long-context modeling across multi-document evidence and diverse evaluation tasks.
Takeaways & Limitations
Loong covers only financial, legal, and academic domains, and its expert proofreading requires substantial time and effort.
Abstract
from arXiv · showhide
Long-context modeling capabilities have garnered widespread attention, leading to the emergence of Large Language Models (LLMs) with ultra-context windows. Meanwhile, benchmarks for evaluating long-context LLMs are gradually catching up. However, existing benchmarks employ irrelevant noise texts to artificially extend the length of test cases, diverging from the real-world scenarios of long-context applications. To bridge this gap, we propose a novel long-context benchmark, Loong, aligning with realistic scenarios through extended multi-document question answering (QA). Unlike typical document QA, in Loong's test cases, each document is relevant to the final answer, ignoring any document will lead to the failure of the answer. Furthermore, Loong introduces four types of tasks with a range of context lengths: Spotlight Locating, Comparison, Clustering, and Chain of Reasoning, to facilitate a more realistic and comprehensive evaluation of long-context understanding. Extensive experiments indicate that existing long-context language models still exhibit considerable potential for enhancement. Retrieval augmented generation (RAG) achieves poor performance, demonstrating that Loong can reliably assess the model's long-context modeling capabilities.
1 Introduction
Loong addresses the lack of realistic multi-document benchmarks by scattering answer evidence across documents so that every document matters. It introduces varied tasks and context lengths, and experiments show that even powerful long-context LLMs still struggle.
- Existing benchmarks often centralize evidence or add distracting texts, allowing models to overlook documents and use shortcuts.
- Loong scatters answer evidence across multiple documents, so bypassing any document leads to an erroneous answer.
- Loong introduces Spotlight Locating, Comparison, Clustering, and Chain of Reasoning for multi-document long-context evaluation.
- Even current powerful LLMs struggle with Loong, indicating substantial room for improving long-context modeling.
- Loong evaluates long-context modeling with varying input lengths, task difficulties, and newly annotated, quality-checked instances.
2 Related Work
Long-context models and benchmarks are advancing, but existing evaluations remain insufficiently aligned with realistic, sufficiently long, multi-document question answering. Prior approaches include context extension, sliding windows, synthetic tasks, and retrieval-based methods.
- Long-Context Language Models: Context-extension methods adapt positional embeddings during fine-tuning, whereas sliding-window methods improve efficiency but fail to exploit the entire context.
- Long-Context Benchmarks: Synthetic tasks such as Needle-in-a-Haystack and Counting Stars mainly indicate surface-form long-context understanding.
- Long-Context Benchmarks: Earlier comprehensive benchmarks often use 5k–25k contexts, while other benchmarks provide longer data but still lack full real-world multi-document alignment.
- Long-Context Benchmarks: Existing benchmarks still lack sufficient length, contamination control, and alignment with real-world multi-document question answering.
- Retrieval-Augmented Language Models: Retrieval-augmented language models use retrieved long documents as external knowledge and can match or exceed task-specific long-context models.
3 Loong: A Long-Context Benchmark
Loong is a bilingual, multi-domain benchmark with four task categories and context sizes up to 250K tokens. Its tasks require locating, comparing, clustering, or reasoning over evidence distributed across documents, supported by structured annotation workflows.
- Benchmark Overview: Loong contains 1,600 Chinese and English test instances across financial reports, academic papers, and legal cases, with context sets from 10K to 250K tokens.
- Evaluation Tasks: The benchmark covers Spotlight Locating, Comparison, Clustering, and Chain of Reasoning to model diverse multi-document semantic relationships.
- Spotlight Locating: Spotlight Locating tests finding evidence in one relevant document among semantically similar, unrelated documents.
- Comparison: Comparison requires locating dispersed evidence and correlating values across documents through enumeration, extrema, or range-based selection.
- Clustering: Clustering extracts and integrates evidence across documents into groups based on textual, numerical, citation, or legal criteria.
- Chain of Reasoning: Chain of Reasoning requires locating evidence across documents and modeling logical relationships for temporal, citation, matching, and sequential inference tasks.
- Data Construction: Annotation combines information compression, template-based and free annotation, GPT-4o generation, evidence recall, self-checking, and manual review.
4 Experiments
Loong evaluates seven advanced LLMs on realistic multi-document tasks using Avg Scores and Perfect Rate, finding substantial weaknesses in long-context modeling. Performance generally declines with context length, while RAG fails to improve overall results because it may omit distributed evidence.
- 4.2 Main Results: Gemini-1.5-pro achieved the best overall performance, with a comprehensive score of 55.37 and a perfect rate of 27%.It particularly excelled on ultra-long contexts in Set3 and Set4, followed by GPT-4o.
- 4.3 Task Analysis: Models generally performed better on Spotlight Locating than on Comparison and Clustering, while Chain of Reasoning performance declined sharply as context length increased.Comparison and Clustering require cross-document matching, contrasting, and classification, whereas shorter Chain of Reasoning settings remained comparatively strong.
- 4.4 Scaling Law of Context Window: Performance notably declined as context length increased, revealing an effective capability boundary below some models’ claimed window sizes.GPT-4o and Qwen2-72B-Instruct began degrading within 50–100K despite being trained on 128K data.
- 4.5 RAG or Not: Adding RAG did not improve overall Loong performance and noticeably harmed tasks requiring comprehensive evidence integration.RAG was more useful for sparse-evidence tasks such as Spotlight Locating, while evenly distributed evidence across documents made retrieval incomplete.
- 4.5 RAG or Not: RAG can help on ultra-long contexts by compressing information and recalling evidence lost through length truncation, but it does not cover all evidence.The paper therefore argues that stronger training on longer texts is needed rather than relying only on RAG.
5 Conclusion
The paper proposes Loong, a benchmark for long-context comprehension in realistic multi-document QA scenarios. Experiments show that even powerful long-context LLMs perform unsatisfactorily, while analyses examine RAG and context-window scaling.
- Loong evaluates long-context comprehension in real-world multi-document question answering and analyzes model parameter sizes, context windows, RAG, and context-length scaling.
Limitations
Loong’s main limitations are restricted domain coverage and substantial expert annotation demands. Its evaluation also relies on GPT-4 judging outputs against gold answers using accuracy, hallucination, and completeness criteria.
- Scope: Loong covers only financial, legal, and academic multi-document domains, leaving many real-world domains outside its scope.The authors attribute this restriction to annotation costs and model evaluation efficiency.
- Annotation: Expert proofreading requires substantial time and effort because annotators evaluate evidence across documents averaging up to 110k in length.Experts must understand questions, search multiple documents, and judge answer consistency.
B.3 Extremum Acquisition
The financial-statement question asks which company has the highest Total Non-current Assets. BLUE DOLPHIN ENERGY CO is identified as the highest, with $56,787,000.
- BLUE DOLPHIN ENERGY CO has the highest Total Non-current Assets, at $56,787,000.
B.4 Range Awareness
Range Awareness evaluates whether models can organize multi-document entities according to a threshold or conceptual range. Its examples include counting qualifying companies and grouping them by total shares outstanding.
- B.4 Range Awareness: Four companies have Total Shares Outstanding exceeding 10,000,000 shares in the example.
- B.4 Range Awareness: The example also groups companies into below 10,000,000 shares and 10,000,000 shares or more.
- B.4 Range Awareness: Citation and reference prompts require identifying directed relationships among the provided papers and returning titles in JSON lists.The examples distinguish papers cited by a given paper from papers that cite it.
B.8 Temporal Analysis
Temporal Analysis asks models to infer attribute changes across ordered financial periods, while related examples test citation chains and document classification over multiple inputs.
- B.8 Temporal Analysis: ARVANA INC’s share capital increased from $4,611 in 2021 to $107,847 in 2024.The reported intermediate values are $34,149 in 2022 and $35,949 in 2023.
- B.8 Temporal Analysis: Temporal Analysis requires analyzing changes or trends in an attribute using financial reports from consecutive years or quarters.The task uses temporal relationships to connect values across documents.
- B.8 Temporal Analysis: The examples also ask models to construct the longest citation chain among provided papers.
- B.8 Temporal Analysis: Loong’s length-distribution data is concentrated around 30-150k while also covering shorter and longer intervals.This supports evaluation across multiple context-length ranges.
D RAG Detailed Results
The reported RAG experiments show no improvement on Loong, while Loong distributes answer evidence across every document rather than concentrating it in one passage.
- The RAG evaluation was conducted on GPT-4o and Qwen2-72B-Instruct, with detailed results reported across multiple tables.
- RAG achieved subpar results on Loong, indicating that the benchmark requires genuine long-context understanding.
- Loong distributes answer-related evidence across every document, unlike LongBench examples where evidence can be confined to Passage 1.
F Comparison of Results with Other Benchmarks
Loong is compared with RULER and NOCHA to examine whether benchmark results align across synthetic, novel-domain, and more realistic long-context tasks.
- Model performance across Loong, RULER, and NOCHA is described as essentially consistent despite their different task and domain emphases.
- Gemini-1.5-pro achieved the best results on Loong and RULER, whereas GPT4o had a significant lead on NOCHA.
- Qwen2-72B-Instruct underperformed smaller models on RULER but not on Loong, where the reported pattern favored larger parameter counts.
G Results of Recall Rate by RAG
The recall analysis finds that RAG often fails to retrieve all documents needed for Loong questions, and its document-level recall remains limited even at larger top-k values.
- The analysis evaluates whether retrieved top-k passages cover all documents because Loong answers distribute evidence across documents.
- Recall@n equals 1 only when the retrieved passages contain all documents and otherwise equals 0, making it a document-coverage measure.
- At topK 50, RAG’s highest document recall on Loong was only 0.64, while actual evidence recall would be lower.
- The recall experiment is presented in Table 6, while related RAG results are reported in Tables 8–10.