Source-linked AI summary

PolyBench: Benchmarking LLM Forecasting and Trading Capabilities on Live Prediction Market Data

Pu Cheng, Juncheng Liu, Yunshen Long

arXiv:2604.14199v1q-fin.CPcs.AIcs.LG

TL;DR

Existing benchmarks do not capture open-domain forecasting under live, multimodal market conditions with point-in-time order-book dynamics. PolyBench addresses this gap with a contamination-proof benchmark and evaluates seven LLMs, finding that only MiMo-V2-Flash and Gemini-3-Flash achieve positive returns.

  • Problem

    Open-domain forecasting across unseen topics remains largely unsolved, while existing benchmarks overlook prediction-market order-book mechanics, capital risk, and point-in-time grounding.

  • Method

    PolyBench couples 38,666 binary market snapshots across 4,997 events with exact CLOB states, contemporaneous news, and official resolution criteria to evaluate LLM forecasting and trading.

  • Results

    Only two of seven evaluated models achieve positive returns: MiMo-V2-Flash reaches 17.6% CWR and Gemini-3-Flash reaches 6.2% CWR, while the remaining five incur losses.

  • Takeaways & Limitations

    PolyBench provides a financially grounded, contamination-proof evaluation standard showing that strong language modeling does not guarantee profitable trading under live market conditions.

Abstract

from arXiv · show

Predicting real-world events from live market signals demands systems that fuse qualitative news with quantitative order-book dynamics under strict temporal discipline -- a challenge existing benchmarks fail to capture. We present \textbf{PolyBench}, a multimodal benchmark derived from Polymarket that records point-in-time cross-sections of 38,666 binary prediction markets spanning 4,997 events, synchronously coupling each snapshot with a Central Limit Order Book (CLOB) state and a real-time news stream. Using PolyBench, we evaluate seven state-of-the-art Large Language Models -- spanning open- and closed-source families -- generating 36,165 predictions under identical, timestamp-locked market states collected between February 6 and 12, 2026. Our multidimensional framework assesses directional accuracy, our proposed Confidence-Weighted Return (CWR), Annualized Percentage Yield (APY), and Sharpe ratio via realistic order-book execution simulation. The results reveal a pronounced performance divergence: only two of seven models achieve positive financial returns -- MiMo-V2-Flash at \textbf{17.6%} CWR and Gemini-3-Flash at 6.2% CWR -- while the remaining five incur losses despite uniformly high stated confidence. These findings highlight the gap between surface-level language fluency and genuine probabilistic reasoning under live market uncertainty, and establish PolyBench as a contamination-proof, financially-grounded evaluation standard for future LLM research. Our dataset and code available at \underline{\href{https://github.com/PolyBench/PolyBench}{https://github.com/PolyBench/PolyBench}}.

1 Introduction

PolyBench addresses the open problem of forecasting diverse real-world events on live prediction markets by combining multimodal, timestamp-locked market data with realistic trading evaluation. Its results show that language-modeling strength alone does not guarantee profitable trading.

  • Open-domain forecasting of reliable probabilistic predictions across diverse unseen topics without domain-specific fine-tuning remains largely unsolved.
  • Live decentralized prediction markets require jointly assessing resolution criteria, breaking news, CLOB states, spreads, and liquidity under point-in-time constraints.
  • PolyBench comprises 38,666 binary prediction markets across 4,997 real-world events, coupling each snapshot with an exact CLOB state, contemporaneous news, and official resolution criteria.
  • Its framework evaluates seven state-of-the-art LLMs on forecasting accuracy, CWR, APY, Sharpe ratio, and instruction adherence using realistic order-book execution simulation.
  • Only two of seven evaluated models achieve positive financial returns, while the remaining five incur losses despite strong language modeling and high stated confidence.

2 Related Work

Existing forecasting and agent benchmarks largely omit live quantitative market context and remain vulnerable to pre-training contamination. PolyBench fills this gap by combining future-event evaluation with synchronized order-book, news, and return-based assessment.

  • Financial forecasting research mainly targets continuous asset variables, while text-based event benchmarks lack real-time quantitative market grounding.
  • Static benchmarks can suffer data leakage that inflates reported accuracy, whereas future prediction-market outcomes cannot be pre-trained from ground-truth results.
  • Prior LLM agent evaluations on Polymarket and live events assess textual prediction but omit order-book mechanics and trade-execution costs.
  • PolyBench combines Polymarket data, live order books, aligned news, and return-based evaluation at scale.

3 Polymarket Definitions and Notations

PolyBench formalizes Polymarket’s event, market, resolution, and contract concepts for decentralized, order-book-based, snapshot evaluation. Events contain tradable binary markets that resolve under official rules, with share costs determined by live liquidity.

  • PolyBench adapts prediction-market definitions to decentralized mechanisms, order-book structures, and snapshot-based evaluation.
  • An event is a high-level container for a future occurrence that establishes context, scope, and official resolution criteria for its markets.
  • Each event contains one or more actively tradable binary markets resolving definitively to Yes or No and accompanied by volume metrics and CLOB states.
  • An event resolves when all constituent markets are definitively settled under Polymarket’s official rules, including conditional constraints such as forced 50-50 splits.
  • A Yes share pays 1 when its market resolves Yes and 0 otherwise, while its executed purchase price is determined dynamically by available market liquidity.

4 PolyBench

PolyBench is a return-focused, automated benchmark that constructs timestamped multimodal market snapshots and evaluates LLM forecasting and trading under realistic execution constraints. Its framework combines strict temporal grounding, structured decisions, selective abstention, and metrics designed to capture accuracy, confidence calibration, returns, and risk.

  • PolyBench aims to provide a return-focused, extendable, automated evaluation of LLM forecasting and market-trading capabilities on Polymarket.
  • Construction: The construction pipeline collects active markets, retrieves contemporaneous news, captures precise CLOB states, dispatches snapshots to LLM endpoints, and matches predictions to resolved outcomes.
  • Construction: Timestamp-locked prompts provide historical event descriptions, market options, order-book spreads, and midpoint prices while enforcing resolution-rule primacy, value identification, microstructure analysis, evidence-backed confidence, and temporal awareness.
  • Construction: Models must return structured JSON decisions, while trades with no positive expected value or confidence below 0.6 are discarded and explicit SKIP choices allow abstention from ambiguous markets.
  • Dataset: The dataset contains actively traded, yet-to-be-resolved markets collected from February 6–12, 2026, supporting future longitudinal assessments after resolution.
  • Evaluation: Directional accuracy evaluates correct settlements, but CWR additionally scales capital by stated confidence and simulates CLOB liquidity absorption, rewarding calibrated successful bets and penalizing overconfident losses.
  • Evaluation: APY annualizes returns by holding duration, Sharpe ratio measures mean return relative to return volatility, and temporal resolution error supplies an additional timing metric.

5 Experiment

PolyBench evaluates seven LLMs on forecasting and trading using market snapshots, execution-aware financial metrics, and domain- and strategy-level analyses. Results show positive returns are rare, position sizing exposes liquidity limits, and confidence or instruction adherence alone does not ensure profitability.

  • 5.1 Overall Performance: 17.6% CWR for MiMo-V2-Flash and 6.2% CWR for Gemini-3-Flash were the only positive returns among seven evaluated models.The other five models incurred losses, while CWR was used as the principal return metric.
  • 5.1 Overall Performance: Exceeding 1,000% returns on some successful low-probability predictions produced annotated spikes in the settlement trajectory.The trajectory runs from the initial February 6 snapshot batch through progressive settlement on February 21, 2026.
  • 5.1 Overall Performance: At L ≤$100, agents preserved theoretical alpha, but larger positions exhausted top order-book levels and worsened execution prices.MiMo-V2-Flash remained relatively resilient up to $500, whereas Gemini-3-Flash’s alpha eroded before both met liquidity limits.
  • 5.2 Domain Expertise and Miscalibrated Conviction: Models maintained high confidence across domains, producing robust Politics alpha but deep negative Crypto returns under severe miscalibration.The domain comparison covers CWR and Average Declared Confidence across eight event categories.
  • 5.3 Strategy Utilization and Instruction Governance: MiMo-V2-Flash generated 50.0% CWR with news_catalyst and 19.2% CWR with value_bet, while some compliant models still had negative returns.Gemini-3-Flash and DeepSeek-V3.2 used stable_yield in over 20% of trades; instruction adherence alone did not guarantee profitability.

6 Conclusion

PolyBench introduces a contamination-proof benchmark that couples live prediction-market snapshots with contemporaneous order-book and news data. Across seven LLMs, only two achieved positive returns, while domain miscalibration, instruction adherence, and order-book slippage shaped financial outcomes.

  • 6 Conclusion: PolyBench evaluates LLMs as trading agents on live decentralized prediction markets using 38,666 market snapshots with aligned CLOB states and news.The benchmark is presented as a financially grounded standard beyond static benchmarks.
  • 6 Conclusion: 17.6% CWR for MiMo-V2-Flash and 6.2% CWR for Gemini-3-Flash were the only positive returns; the remaining five models lost money despite high confidence.The conclusion also identifies uneven domain competence, insufficient instruction adherence, and larger-lot slippage as key findings.
Loading 2604.14199v1…