Source-linked AI summary
Interactive Benchmarks
Baoqing Yue, Zihan Zhu, Yutong Han, Brian Fan, Qian Sun, Jichen Feng, Hufei Yang, Yifan Zhang, Mengdi Wang
TL;DR
Existing reasoning evaluations are limited by saturation, contamination, and subjective judgment, motivating assessment of how models acquire and use information. The paper introduces budgeted, multi-turn Interactive Benchmarks spanning proof and game settings, and finds that interactive evaluation reveals substantial room for improvement in current models.
Problem
Fixed benchmarks can be saturated and contaminated, while preference-based evaluations rely on subjective judgments and leave active information acquisition under-measured.
Method
Interactive Benchmarks evaluate models through budgeted multi-turn interaction in proof tasks with judge feedback and games requiring strategic action for long-horizon utility.
Results
Across Logic, UI2Html, Math, Poker, and the Trust Game, interactive evaluation captures information-acquiring ability and shows current models still have substantial room to improve.
Takeaways & Limitations
Interactive Benchmarks provide a unified way to assess reasoning through objective feedback and strategic interaction.
Takeaways & Limitations
Measured scores depend on judge behavior and, for UI2Html, the summarizer, so results reflect a particular interaction protocol rather than a completely judge-invariant capability estimate.
Abstract
from arXiv · showhide
Existing reasoning evaluation paradigms suffer from different limitations: fixed benchmarks are increasingly saturated and vulnerable to contamination, while preference-based evaluations rely on subjective judgments. We argue that a core aspect of intelligence is the ability to decide what information to acquire and how to use it effectively. We propose Interactive Benchmarks, a unified evaluation paradigm that assesses a model's reasoning ability through budgeted multi-turn interaction. We evaluate models under this framework in two settings: Interactive Proofs, where models interact with a judge to solve Logic, UI2Html, and Mathematics tasks under objective feedback; and Interactive Games, where models reason strategically to maximize long-horizon utilities. Our results show that interactive benchmarks provide a more robust assessment of this dimension of model intelligence, revealing substantial room for improvement in interactive scenarios.
1 Introduction
Interactive Benchmarks target a neglected component of intelligence: deciding what information to acquire and how to use it under uncertainty. They organize evaluation as sequential interaction for objective target recovery or strategic utility maximization.
- Existing fixed datasets are increasingly saturated and contamination-prone, while preference-based arenas rely on subjective judgments.
- Interactive Benchmarks assess how models decide what evidence to seek and use, complementing static tests that often miss active information acquisition.
- Interactive Games: Interactive Games place Players in stochastic or adversarial environments where they act strategically to maximize long-term utility.
- Interactive Proofs: Interactive Proofs use judge feedback to help Players converge on verifiable targets in Logic, UI2Html, and Mathematics tasks.
2 Interactive Benchmarks
Interactive Benchmarks model evaluation as budgeted interaction between a model and an environment, spanning proof tasks that recover hidden answers and games that optimize rewards. The concrete protocols require models to query selectively, integrate sparse feedback, and revise reasoning or actions.
- Each benchmark instance is a horizon-T interaction in which a model selects actions from its history and receives subsequent environment observations until termination or budget exhaustion.
- Interactive Proofs: In interactive proofs, a verifier provides restricted feedback on queries, and the model must submit the correct hidden solution within a cost budget.
- Interactive Proofs: Logic: Situation Puzzle requires recovering a hidden causal explanation through yes/no-style queries, while no-interaction evaluation yields 0% accuracy for all evaluated models.
- Interactive Proofs: UI2Html: UI2Html has Players submit complete HTML and clarification questions over 20 rounds, with a Judge scoring layout, components, style, text, and polish.
- Interactive Proofs: Math: Interactive mathematics lets Players query intermediate claims or submit solutions, enabling early pruning of incorrect branches and explicit hypothesis testing under sparse feedback.
- Interactive Games: Texas Hold’em evaluates agents under partial observations and strategic uncertainty, using repeated-play cumulative bankroll as the outcome.
3 Experiments
The experiments evaluate Interactive Benchmarks across interactive proofs and games, measuring performance under fixed interaction or token budgets. Results show that interaction can improve task performance, but gains vary by model and strategic success remains uneven.
- Interactive Proofs: Interactive Proofs evaluate six frontier models on Situation Puzzle, UI2Html, and Mathematics tasks under controlled judges and interaction budgets.Situation Puzzle uses 46 puzzles and a 20-turn budget; UI2Html fixes the summarizer and judge to qwen-vl-max with 20 interaction rounds plus finalization.
- Interactive Proofs: Logic: 30.4% accuracy is Gemini-3-flash’s best Situation Puzzle result, while Qwen3-max reaches 4.3% and Deepseek-v3.2 averages 18.0 turns among solved puzzles.Kimi-k2 is fastest among solved cases at 12.3 turns, followed by Gemini-3-flash at 13.3.
- Interactive Proofs: UI2Html: All six UI2Html models score higher with 20-round interaction than with a single-round baseline, with GPT-5-mini highest at 57.62.Grok-4.1-fast gains most, from 53.19 to 57.12, whereas Kimi-k2 changes only from 48.88 to 49.03.
- Interactive Proofs: Mathematics: Interactive evaluation exceeds budget-matched pass@k by roughly 20%-50% across models, while Grok-4.1-fast achieves 76.9% interactive accuracy.Qwen3-max uses the fewest turns among correct trials at 5.2, whereas DeepSeek-v3.2 requires 12.0 turns and reaches 48.1% accuracy.
- Interactive Games: Texas Hold’em: Gemini-3-flash leads Texas Hold’em winnings at 31.8 ± 42.4 per hand, while agents display distinct participation and folding styles.GPT-5-mini has the highest VPIP at 23.7% ± 1.1%; DeepSeek-v3.2 is tightest at 9.0% ± 2.0% VPIP and 90.5% ± 1.4% fold rate.
- Interactive Games: Trust Game: 1.867 is Qwen3-max’s highest Trust Game payoff per round, and only Qwen3-max and GPT-5-mini outperform both heuristic baselines.The Grim Trigger and TFT baselines score 1.811 and 1.782, respectively, leaving substantial room for improvement in adaptive game playing.
4 Related Work
Prior interactive benchmarks test multi-turn capability but often target specialized settings and do not isolate interaction's contribution. Interactive Benchmarks address these gaps with a unified, principled framework for objective comparison across tasks.
- Benchmarks that require interaction: Existing interactive benchmarks include hypothesis-driven puzzles, question-selection games, abstract-task refinement, and specialized multi-turn dialogue evaluations.Examples include TurtleBench, Entity-deduction Arena, ARC-AGI, MT-Eval, TurnBench-MS, and medical consultation evaluation.
- Research gap: These benchmarks do not explicitly isolate interaction from task-specific priors, environment design, or reward shaping.Their protocols also lack a broadly applicable mathematical principle for objective comparison across tasks and settings.
- Research gap: Interactive Benchmarks formalize interaction theoretically and provide a general framework for principled, reproducible evaluation.The framework is designed to address both the isolation problem and limited cross-task generalization.
5 Conclusion
The paper introduces Interactive Benchmarks as a unified, budgeted multi-turn evaluation framework spanning objective target recovery and strategic utility maximization. Across five testbeds, experiments find that interactive evaluation exposes information-acquisition ability and substantial room for improvement.
- Conclusion: Interactive Benchmarks measure reasoning through budgeted, multi-turn interaction in Interactive Proofs and Interactive Games.Proofs use judge feedback for Logic, UI2Html, and Math; Games require strategic interaction in Poker and the Trust Game.
- Conclusion: Across Logic, UI2Html, Math, Poker, and the Trust Game, interactive evaluation captures information-acquiring ability that previous benchmarks struggle to assess.The conclusion covers both objective-feedback and long-horizon strategic settings.
- Conclusion: Current models still have substantial room to improve in interactive scenarios.The authors propose broader task coverage and training methods for interactive performance as future work.
A.1.1 Judge Sensitivity
Judge choice changes absolute Situation Puzzle scores, but has limited impact on player rankings. Increasing the interaction budget benefits stronger logic players more clearly, while weaker players show smaller and non-monotonic changes.
- Judge sensitivity: Judge choice affects absolute Situation Puzzle scores, while stronger players remain consistently ahead across judges.The relative ranking is broadly stable, although judge sensitivity helps calibrate absolute performance.
- Interaction budget: The ablation varies the maximum interaction budget over 0, 5, 10, 15, and 20 rounds while reporting success rate.Dataset, judge, prompt format, and decoding settings remain fixed.
- Interaction budget: 30.4%: Gemini-3-flash rises from 17.4% at 5 rounds to 30.4% at 20 rounds.DeepSeek-v3.2 also increases from 4.3% to 15.2%, whereas Kimi-k2 and Qwen3-max remain low with small, non-monotonic changes.
A.2 UI2Html
UI2Html judge choice materially changes measured reconstruction quality, although player rankings remain broadly stable. The strongest judge configuration produces higher scores and clearer separation among players, making judge sensitivity an important evaluation caveat.
- Judge ablation: UI2Html judge ablation varies the shared summarization, comparison, and scoring stack while comparing player models by average reconstruction score.The x-axis varies the judge model, and each bar represents a player model.
- Judge ablation: 57.22, 58.06, 54.50, and 52.64: qwen-vl-max scores for grok-4.1-fast, gpt-5-mini, deepseek-v3.2, and qwen3-max, respectively.Replacing qwen-vl-max with gpt-4.1-mini causes a moderate consistent drop, while gemini-2.5-flash is weakest in nearly every case.
- Judge ablation: 31.28: grok-4.1-fast falls from 57.22 under qwen-vl-max to 31.28 under gemini-2.5-flash.DeepSeek-v3.2 similarly drops from 54.50 to 30.84, and qwen3-max from 52.64 to 30.40.
- Limitation: Measured reconstruction quality is not judge-invariant, although relative player rankings remain broadly stable across configurations.The authors suggest reporting judge sensitivity explicitly or averaging over multiple judge configurations.
A.3 Math
The math judge ablation varies only the judge model while holding the dataset, prompt, temperature, and 20-turn budget fixed. Player-level differences remain visible across judges, although judge identity affects absolute accuracy and average turns.
- Accuracy: Gemini-3-flash records 59.6%–69.2% accuracy across judges, compared with 34.6%–48.1% for DeepSeek-v3.2 and 32.7%–38.5% for Kimi-k2.The heatmap’s row-level separation remains visually intact despite changes in judge identity.
- Interpretation: Judge identity changes absolute outcomes without eliminating the player-level ordering visible in the accuracy heatmap.The experiment keeps all other evaluation conditions fixed and represents player–judge pairings as heatmaps.
- Average turns: DeepSeek-v3.2 averages 11.3–13.9 questions among solved instances, versus 7.3–8.4 for Gemini-3-flash and 7.8–10.8 for Kimi-k2.The average-turn heatmap is also primarily row-wise, but some judge columns exert a global effect.
A.4 Texas Hold’em
The Texas Hold’em sanity check compares six LLM agents with deterministic all-in and fold baselines across 1000 hands and randomized eight-player tables. All LLM agents remain profitable, while extreme all-in play is severely exploitable.
- Experimental setup: The sanity check uses two deterministic baselines—AllIn-BL always goes all-in when able, whereas Fold-BL folds whenever folding is available.Each table contains the same six LLM agents plus both baselines.
- Results: 401.5–699.5 chips per hand are the average winnings of the six LLM agents, while AllIn-BL loses −3320.8 chips per hand and Fold-BL loses −19.8.The auxiliary experiment uses 1000 hands across 10 independent eight-player tables with randomized seat order.
- Results: 99.3% of the aggregate gains obtained by the six LLM agents correspond to the loss of AllIn-BL, indicating that naive extreme aggression is strongly exploitable.LLM fold rates rise to 80.1%–89.6% in this setting, consistent with more selective play against the all-in opponent.
A.5 Trust Game
The Trust Game ablation varies the continuation probability and measures payoff alongside cooperation and betrayal behavior. Intermediate horizons perform best for most models, while longer horizons produce model-dependent, non-monotonic effects and ranking changes.
- Continuation probability: Most models peak in payoff around an expected horizon of 13.3 rounds, where cooperation is also near its maximum.This corresponds to δ = 0.925 in the sweep over continuation probabilities.
- Continuation probability: Beyond an expected horizon of 13.3 rounds, increasing δ has model-dependent and non-monotonic effects, with rankings changing noticeably.The benchmark describes δ ≈0.925 as the most balanced choice; shorter horizons compress interaction, while longer horizons tend to introduce instability.
- Behavioral statistics: Cooperation rate and betrayal rate complement average payoff by summarizing how models balance reciprocity against opportunistic defection.Betrayal rate measures how often a model defects when the opponent cooperates in the last round.
- Behavioral statistics: Behavioral statistics reveal substantial diversity: some models sustain high cooperation with low betrayal, while others behave more opportunistically.These statistics make strategic styles explicit beyond average payoff per round.
C.2 UI2Html
UI2Html evaluates iterative webpage reconstruction by comparing rendered outputs with a hidden reference and using judge feedback to guide revisions. The interaction trace illustrates successive structural decisions, ending after 20 rounds with a final rendered webpage.
- Interaction trace: The Reddit reconstruction trace uses judge feedback to accept or reject proposed additions to navigation, sidebars, buttons, and content sections.Accepted changes include left-sidebar navigation, related posts, resources, topics, a See more button, and a User Settings icon.
- Example: The example contrasts the hidden reference, the first-round rendering, and the final webpage after 20 interaction rounds.These stages show the progression from the initial textual prompt through iterative refinement.
D Limitations
The benchmark’s results are constrained by judge and protocol design, limited task coverage, domain-specific skills, and incomplete realism. These boundaries limit how broadly performance should be interpreted as general interactive reasoning.
- Scores depend on the judge and, for UI2Html, the summarizer, so the benchmark measures performance under a particular interaction protocol.Absolute scores can shift across judge choices even when broad player rankings remain relatively stable.
- The benchmark spans two interaction regimes and five tasks, but its datasets and game configurations remain narrow relative to real interactive intelligence.The proof evaluations use 46 logic puzzles, 50 UI2Html screenshots, and 52 math problems, while game results use one poker engine and one Trust Game parameterization.
- Fixed budgets and restricted interfaces improve reproducibility but may favor models especially adept at adapting to the benchmark’s structured interaction format.Examples include a fixed 20-round proof budget, restricted Logic and Math feedback, and full HTML revisions plus binary questions in UI2Html.
- Task success remains entangled with domain-specific ability, such as HTML/CSS competence in UI2Html and game knowledge in poker.The benchmark therefore measures interactive reasoning together with non-interactive domain priors.
- The evaluation only partially captures deployment cost and realism because it excludes some system overheads and uses controlled or stylized environments.The math comparison matches player-side tokens but excludes judge-side cost and latency; proof and game settings also simplify real-world conditions.
E.7 Benchmark Interface
The benchmark interface treats the model as an agent that receives structured poker-game state and must emit parser-recognized actions. The broader contribution is an evaluation framework rather than a deployable decision-making system.
- At each poker decision, the agent receives the round, private cards, community cards, pot and call amount, stack sizes, and recent action history.
- The agent must output one parser-recognized action: FOLD, CHECK, CALL, RAISE, or ALL_IN.
- The paper contributes an evaluation framework and benchmark protocol rather than a new deployable model or decision-making system.The authors position robust interactive evaluation as a way to assess reasoning more accurately and reduce reliance on static benchmarks.