Source-linked AI summary

BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese

Peilin Zhou, Bruce Leon, Xiang Ying, Can Zhang, Yifan Shao, Qichen Ye, Dading Chong, Zhiling Jin, Chenxuan Xie, Meng Cao, Yuxin Gu, Sixin Hong, Jing Ren, Jian Chen, Chao Liu, Yining Hua

arXiv:2504.19314v2cs.CL

TL;DR

Existing browsing benchmarks provide limited direct evaluation of web retrieval and reasoning in the Chinese information environment, whose fragmented structure and linguistic characteristics differ from English web settings. BrowseComp-ZH fills this gap with 289 reverse-designed, multi-hop Chinese questions and two-stage quality control, then benchmarks more than 20 systems. Most standalone models perform poorly, while the best reported system, DeepResearch, reaches 42.9%, and retrieval helps inconsistently across systems.

  • Problem

    Existing benchmarks lack direct evaluation of browsing competence in the Chinese web environment, where fragmented content, inconsistent naming, and linguistic complexity challenge English-oriented or translated evaluations.

  • Method

    BrowseComp-ZH reverse-designs native Chinese multi-constraint questions from factual answers and applies three-engine validation plus human-in-the-loop checks for difficulty and answer uniqueness.

  • Results

    Most standalone models achieve limited accuracy, while AI search products with iterative retrieval perform better; DeepResearch reaches 42.9% and Doubao (Deep Search) reaches 26.0%.

  • Takeaways & Limitations

    BrowseComp-ZH demonstrates that Chinese web agents need both effective retrieval and strong reasoning and information reconciliation to solve difficult multi-hop questions.

  • Takeaways & Limitations

    The dataset is relatively small, answer uniqueness cannot be fully guaranteed, and changing web facts create stability and reproducibility challenges.

Abstract

from arXiv · show

As large language models (LLMs) evolve into tool-using agents, the ability to browse the web in real-time has become a critical yardstick for measuring their reasoning and retrieval competence. Existing benchmarks such as BrowseComp concentrate on English and overlook the linguistic, infrastructural, and censorship-related complexities of other major information ecosystems -- most notably Chinese. To address this gap, we introduce BrowseComp-ZH, a high-difficulty benchmark purpose-built to comprehensively evaluate LLM agents on the Chinese web. BrowseComp-ZH consists of 289 multi-hop questions spanning 11 diverse domains. Each question is reverse-engineered from a short, objective, and easily verifiable answer (e.g., a date, number, or proper noun). A two-stage quality control protocol is applied to strive for high question difficulty and answer uniqueness. We benchmark over 20 state-of-the-art language models and agentic search systems on our proposed BrowseComp-ZH. Despite their strong conversational and retrieval capabilities, most models struggle severely: a large number achieve accuracy rates below 10%, and only a handful exceed 20%. Even the best-performing system, OpenAI's DeepResearch, reaches just 42.9%. These results demonstrate the considerable difficulty of BrowseComp-ZH, where success demands not only effective retrieval strategies, but also sophisticated reasoning and information reconciliation -- capabilities that current models still struggle to master. Our dataset, construction guidelines, and benchmark results have been publicly released at https://github.com/PALIN2018/BrowseComp-ZH.

1 Introduction

BrowseComp-ZH addresses the lack of direct evaluation for Chinese web browsing by introducing a native, high-difficulty benchmark. Results show that success requires coordinated retrieval and reasoning, while retrieval integration can help or hurt depending on the system.

  • Motivation: Existing browsing evaluations largely target English and do not directly measure agents’ ability to retrieve, filter, and reason over web information.This gap matters for assessing real-world information-seeking with dynamic external evidence.
  • Motivation: Chinese web browsing adds fragmented platforms, inconsistent naming, and linguistic and cultural features that complicate straightforward search.Direct translation of English benchmarks can produce unnatural queries, mismatched information pathways, or trivial keyword matches.
  • Benchmark: BrowseComp-ZH is a native Chinese benchmark whose reverse-designed questions begin with factual answers and use multi-constraint queries that are difficult to retrieve but easy to verify.Queries, evidence chains, and browsing steps are localized, with validation across Baidu, Bing, and Google and a two-stage quality-control process.
  • Benchmark: 289 complex questions span 11 domains and evaluate open-source models, closed-source APIs, and AI search products on multi-hop retrieval and reasoning.The dataset is designed around concise factual answers while requiring integration of online information.
  • Findings: 42.9% and 26.0% are the reported accuracies for DeepResearch and Doubao (Deep Search), respectively, among systems using orchestrated retrieval and reasoning pipelines.Naive models often remain below 10%, while reasoning-oriented models improve over comparable systems.
  • Findings: DeepSeek-R1 accuracy drops from 23.2% without web access to 7.6% with web search enabled, showing that retrieval integration is not uniformly beneficial.The benchmark exposes difficulties in processing, evaluating, and aligning retrieved information with internal representations.

2 Related Work

Prior work studies tool use, retrieval augmentation, and English information-seeking benchmarks, but Chinese web evaluations remain limited in difficulty control, multi-hop reasoning, and native coverage. BrowseComp-ZH addresses these gaps with reverse-designed Chinese tasks, rigorous validation, and broad system evaluation.

  • Existing evaluations: WebGPT, Toolformer, ReAct, and retrieval-augmented generation frameworks examine how language models use tools and external knowledge for complex question answering.Related work also covers summarization and fact verification applications of injected external information.
  • Existing evaluations: English benchmarks such as TriviaQA, HotpotQA, FEVER, KILT, and GAIA cover multi-hop reasoning and fact checking but often permit simple keyword retrieval.Their reliance on structured sources can under-test complex search planning and information synthesis.
  • Chinese benchmarks: Chinese retrieval datasets have addressed web search, dynamic question answering, or RAG, but individually lack rigorous difficulty control, multi-hop emphasis, or comparable scope.These limitations leave native Chinese browsing competence incompletely evaluated.
  • BrowseComp-ZH: BrowseComp-ZH contributes native Chinese reverse-designed tasks, multi-step validation for retrieval difficulty and answer verifiability, and broad coverage of open and proprietary agents.It targets fragmented, unstructured, and linguistically diverse Chinese information sources.

3 The BrowseComp-ZH Dataset

BrowseComp-ZH is constructed by reverse-designing difficult, multi-constraint Chinese queries from objective answers and validating them for retrieval difficulty and uniqueness. The resulting 289-question dataset spans 11 domains with dense questions and short answers.

  • Dataset construction: Ten expert contributors select objective, specific, independently verifiable answers across 11 domains before designing multi-constraint questions.Topics include Film & TV, Technology, Art, History, Sports, Music, Geography, Policy & Law, Medicine, Video Games, and Academic Research.
  • Dataset construction: Annotators reverse-design queries by combining contextual constraints so each answer is unique and direct retrieval is non-trivial.Each sample includes at least one authoritative source URL linking the constraints to the target answer.
  • Difficulty control: Queries are tested across Baidu, Bing, and Google, revised when the answer appears on a first page, and further hardened if GPT-4o and DeepSeek solve them with minimal effort.The process initially produces 480 preliminary samples.
  • Quality control: A two-stage quality-control protocol screens question difficulty and answer uniqueness using timed human search checks and human-in-the-loop verification of model-generated answers.Questions solved within 10 minutes are labeled low difficulty; alternative valid answers cause rejection.
  • Quality control: 76 simple samples are removed during difficulty validation, leaving 404 high-difficulty candidates.The cross-check uses search engines without LLM assistance and applies a strict 10-minute limit per question.
  • Quality control: 115 ambiguous samples are eliminated, yielding 289 validated questions with high difficulty and answer verifiability.Ambiguity is identified when an alternative answer satisfies all task constraints and differs from the original.
  • Data statistics: Film & TV accounts for 15.6% of samples, followed by Art at 13.8% and Geography at 12.8%, while Policy & Law and Academic Papers have 3.5% and 2.4%.The distribution covers a broad range of Chinese web knowledge domains.
  • Data statistics: Most questions contain 60 to 90 characters, whereas answers typically contain 5 to 10 characters.This pairing makes questions information-dense while keeping target answers concise and verifiable.

4 Benchmarks

The benchmark evaluates diverse language models and AI search products using accuracy and calibration error, with answer extraction and grading procedures tailored to system type. Results show strong variation across systems, with iterative retrieval and reasoning generally associated with better performance, while some search configurations degrade accuracy.

  • Evaluation setup: Over 20 models and AI search products are evaluated, covering open-source models, closed-source APIs, and retrieval systems.AI search products are evaluated through human annotator GUI interactions, while model outputs are extracted and graded automatically.
  • Metrics: Accuracy and calibration error are reported using five probability bins and a weighted average of bin-level confidence–accuracy differences.ECE compares each bin’s average predicted confidence with its empirical accuracy.
  • Performance: 42.9% accuracy makes OpenAI’s DeepResearch the strongest evaluated system, followed by O1 at 29.1% and Gemini-2.5-Pro at 27.3%.Qwen2.5-72B-Instruct, GPT-4o, and Claude-3.5-Sonnet achieve 6.6%, 6.2%, and 5.5% accuracy, respectively.
  • Reasoning and retrieval: Reasoning-enhanced systems outperform their non-reasoning counterparts across the reported model comparisons.Doubao Deep Search reaches 26.0% versus 18.7% for Doubao Standard, a nearly 8% absolute improvement.
  • Reasoning and retrieval: Single-retrieval systems can underuse reasoning, whereas multi-round retrieval iteratively refines searches for multifaceted questions.The benchmark distinguishes single-retrieval systems from systems such as DeepResearch, Perplexity, and Doubao that perform multiple retrieval rounds.
  • Failure cases: DeepSeek-R1 accuracy falls from 23.2% without web access to 7.6% with search enabled, while its search variant shows no significant advantage over the standard version.The authors hypothesize that unreliable retrieved content can override more accurate internal knowledge; search also raises DeepSeek-R1’s calibration error from 59% to 65%.

5 Conclusion and Discussion

BrowseComp-ZH is presented as a Chinese-web benchmark built from difficult, verifiable multi-hop questions and evaluated across diverse systems. Results indicate that reasoning and iterative retrieval improve performance, while the dataset remains limited by size and the web’s changing factual content.

  • Conclusion: BrowseComp-ZH evaluates Chinese-web browsing and reasoning through questions requiring multi-hop retrieval, information filtering, and logical reasoning.Its two-stage quality control combines three-engine keyword validation with human verification to make answers difficult to retrieve and unambiguous.
  • Conclusion: 289 questions across 11 topics support evaluations of more than 20 open-source models, closed-source APIs, and AI search products.Topics include Film & TV, History, Technology, and Medicine.
  • Conclusion: 42.9% for DeepResearch and 26.0% for Doubao Deep Search exceed the performance of weaker standalone models and illustrate the value of iterative retrieval.O1 and Gemini-2.5-Pro reach 29.1% and 27.3%, respectively.
  • Limitations: The dataset’s small size limits representativeness, and answer uniqueness cannot be fully guaranteed as web facts change or become inconsistent.The authors identify stability and reproducibility as ongoing challenges.
  • Future work: Future work will expand question coverage, analyze reasoning and search strategies, study failure cases, and explore post-training methods.These directions are intended to support broader evaluation and improved browsing and reasoning capabilities.

A.1 Instructions for model prediction

The prediction instructions require models to answer using intrinsic knowledge when external resources would otherwise be needed, rather than refusing. Responses must follow a fixed explanation, answer, and confidence format.

  • Prediction instruction: Models are instructed to rely on intrinsic knowledge instead of refusing when a question would require external resources.The instruction is applied to open-source and closed-source models that may state they lack search capabilities.
  • Prediction instruction: Each response must contain an explanation, an exact answer, and a confidence score from 0% to 100%.The required labels are Explanation, Exact Answer, and Confidence.

A.2 Instruction for grading

The grading procedure adopts the BrowseComp grading prompt and uses GPT-4o to assess model responses.

  • Grading instruction: GPT-4o performs grading with the same prompt used in BrowseComp.The grading prompt is adopted without a separately described modification.
Loading 2504.19314v2…