Source-linked AI summary
CEO-Bench: Can Agents Play the Long Game?
Haozhe Chen, Karthik Narasimhan, Zhuang Liu
TL;DR
CEO-Bench addresses whether agents can sustain strategic control over long-horizon, uncertain, changing real-world tasks. It simulates operating a startup for 500 days, finding that most models struggle, none surpass the rule-based baseline, and only three finish above $1M.
Problem
Existing evaluations largely test short-horizon execution, leaving agents’ sustained strategic control under uncertainty, noisy information, and changing conditions insufficiently tested.
Method
CEO-Bench simulates an agent operating a startup for 500 days while coordinating pricing, growth, product, operations, communication, and sales through programmable tools.
Results
Most state-of-the-art models struggle to complete the simulation, only three finish above the $1M starting balance, and the rule-based baseline achieves $15.76M.
Takeaways & Limitations
CEO-Bench exposes a gap between agents’ local tool competence and their ability to sustain coherent strategy as delayed, uncertain decisions compound over time.
Takeaways & Limitations
The simulation approximates startup operations and omits qualitative product changes, compliance, security, and fundraising.
Abstract
from arXiv · showhide
Language model agents are becoming proficient executors at isolated, short-horizon tasks such as software engineering and customer service. Yet real-world challenges require a combination of sophisticated skills that remain largely untested in agents: (1) navigating long horizons amid uncertainty; (2) acquiring information in noisy environments; (3) adapting to a changing world; (4) orchestrating multiple moving parts toward a coherent goal. We introduce CEO-Bench, which evaluates these capabilities together by simulating a representative real-world task: operating a startup for 500 days. An agent manages pricing, marketing, budgeting, and many other aspects of a fictional company through a programmable Python interface, operating in the same environment and facing the same challenges as a human CEO. Success demands analyzing noisy, interconnected business databases, translating signals into sound strategy, and coordinating many decisions with programming. The strongest agents write sophisticated code that forecasts churn regimes, billing timing, customer losses, and future cash under different scenarios. Even so, most state-of-the-art models struggle in this environment. Only Claude Fable 5, GPT-5.6 Sol, and Claude Opus 4.8 finish above the $1M starting balance, and all evaluated models remain below the rule-based baseline. CEO-Bench takes a first step toward measuring the intelligence required to drive sustained, adaptive progress over time.
1 Introduction
CEO-Bench tests whether agents can sustain strategic control over 500-day startup operations under uncertainty, noisy information, changing conditions, and many coordinated decisions. Current state-of-the-art agents remain challenged: only three finish above the $1M starting balance, and none reaches the non-LLM rule-based baseline.
- Motivation: The benchmark targets sustained strategic control: prioritizing, allocating limited resources, interpreting noisy signals, and adapting as conditions change.These capabilities extend beyond short-horizon tasks with clear goals and rapid feedback.
- Benchmark design: CEO-Bench simulates operating a startup for 500 days through a programmable Python interface with 34 tools and a 19-table business database.Agents can execute code, query the database with SQL, and compose tools into custom workflows.
- Results: Only Claude Fable 5, GPT-5.6 Sol, and Claude Opus 4.8 finish above the $1M starting balance.Most agents produce valid tool calls and analytics queries but struggle to sustain coherent strategy and often go bankrupt before completion.
- Results: No evaluated model reaches the non-LLM rule-based baseline.This result shows that the simulated long-horizon challenge remains difficult for current state-of-the-art agents.
- Analysis: Performance correlates with inferring hidden structure from noisy data, forecasting delayed consequences, and adapting to competitive pressure.Trajectory analysis reveals distinct strategies, including late adjustment, early exploration followed by passivity, and other model-specific behavior patterns.
2 Designing CEO-BENCH
CEO-BENCH simulates operating a subscription-software startup for 500 days in a partially observable, changing world, with granular mechanics and delayed, interconnected consequences. Its programmable interface lets agents organize open-ended actions into workflows while managing uncertainty across customers, markets, products, and competitors.
- Simulator setup: 500 simulated days begin with zero customers and $1M in cash, and bankruptcy occurs if cash falls strictly below zero.Agents are graded on cash on hand at the end of the simulation.
- Action interface: Agents can take unlimited-turn actions each simulated week across 34 tools, using a programmable interface to organize granular decisions into custom workflows.The interface is designed to make the action space open-ended while remaining manageable through programming.
- Granular world mechanics: The simulator models 26 customer groups and individual customer trajectories with heterogeneous preferences, acquisition responses, budgets, support expectations, and behavioral patterns.Subscription decisions follow a product-value-versus-price participation rule, while acquisition depends on channel responses and broader market conditions.
- Partial observability: Agents observe only partial and delayed evidence, inferring hidden customer, market, competitor, and demand conditions from dashboards, records, social media, research, and negotiation history.True satisfaction, willingness to pay, churn propensity, competitor schedules, and demand parameters remain unobserved.
- Dynamic challenges: Interconnected dynamics, delayed and stochastic effects, and a non-stationary environment force agents to coordinate long-horizon decisions and continually adapt.Reputation can spill across groups, costs may precede benefits by weeks, R&D timelines are stochastic, and customer, competitor, and macroeconomic conditions change over time.
3 Experiments and Results
Experiments evaluate 16 agents over three 500-day simulations with $1M starting cash, showing that most models struggle to avoid bankruptcy and coordinate long-horizon growth, quality, and cash flow. The strongest agents succeed through distinct customer strategies, adaptive exploration, and increasingly sophisticated analysis of churn, payment timing, costs, and cash.
- Models: 16 models are evaluated on the full 500-day simulation, with $1M starting cash, three simulations per model, random seed 42, and maximum reasoning effort.The evaluated set includes Claude Fable 5, GPT-5.6 Sol, Claude Opus 4.8, Claude Opus 4.7, Claude Sonnet 5, Claude Sonnet 4.6, Claude Haiku 4.5, GPT-5.5, Qwen 3.7 Max, Gemini 3.5 Flash, Gemini 3 Flash, GLM 5.2, GLM 5.1, Kimi K2.6, DeepSeek V4 Pro, and Grok 4.20.
- Overall results: Most state-of-the-art models struggle to complete the simulation without bankruptcy, while the strongest models show high-upside strategic behavior but often fail to coordinate growth, quality, and cash flow across the full horizon.DeepSeek V4 Pro, Gemini 3 Flash, and Grok 4.20 bankrupt on all runs.
- Overall results: $11.31M is GPT-5.6 Sol’s best-run finishing cash, but its other two runs bankrupt around day 190; Claude Fable 5, GPT-5.6 Sol, and Claude Opus 4.8 are the only evaluated models whose best runs finish above starting cash.The provided passage identifies GPT-5.6 Sol as second only to Claude Fable 5 in best-run cash.
- Agent behaviors: Claude Fable 5 keeps adjusting its plan as conditions change, Claude Opus 4.8 explores broadly before becoming more passive, and Claude Opus 4.7 narrows earlier into waiting and protecting cash.Quantitatively, Claude Fable 5 and Claude Opus 4.8 spread their actions more evenly across tools than Claude Opus 4.7.
- Agent behaviors: Claude Fable 5, Claude Opus 4.8, and GPT-5.6 Sol finish above starting cash through distinct customer-base strategies, including early scaling, mid-simulation collapse, gradual decline, and sustaining a smaller base.Claude Opus 4.8 drops to zero customers mid-simulation, GPT-5.6 Sol reaches zero by day 420, and Claude Fable 5 sustains a smaller base centered on Individual Group 2 through the end.
- Sophisticated analytics: Top-performing agents write analytics that estimate customer retention and cancellations, track payment timing and accumulating costs, and forecast final-week cash under customer-loss scenarios.Claude Opus 4.8 uses customer payment and cancellation behavior, while Claude Fable 5 models payment timing, costs, and alternative customer-loss assumptions.
4 Ablating Simulator and Agent Configurations
Ablations show that competitor difficulty and simulation horizon materially affect CEO-Bench outcomes, while changing the agent harness can significantly alter behavior and results. Even with a 50-day horizon, no evaluated model exceeds its starting cash balance.
- Overall findings: Existing models remain challenging to evaluate even over a short horizon, and harness choice impacts outcomes significantly.These findings indicate that both simulator and agent-configuration choices meaningfully shape measured performance.
- Simulation horizon: 50 days: no evaluated model finishes above its starting cash balance, despite the horizon being one-tenth of the original.The shortened horizon reduces long-term planning challenges but still exposes weaknesses in orchestrating decisions toward a short-term goal.
- Simulator difficulty: Competitor configuration provides an effective difficulty control, with weaker or absent competitors making CEO-Bench easier.The competitor raises expectations through both a preset stationary sequence and adaptive responses to agent actions.
- Agent harness: Changing the agent harness significantly changes results and agent behavior even when the underlying model remains fixed.The study compares Claude Fable 5 with Claude Code and GPT-5.6 Sol with Codex against a custom minimal terminal-using harness.
5 Related Work
Related work has shifted language-model evaluation toward realistic agentic execution and longer-horizon processes. CEO-Bench further challenges long-horizon planning by requiring delayed investments whose returns arrive much later.
- Language model evaluations: Language-model evaluation has shifted from static knowledge and reasoning toward realistic agentic task execution.Recent benchmarks have also broadened evaluation toward economically valuable tasks.
- Long-horizon agent evaluation: Memory and continual-learning benchmarks test information retention, but often reduce evaluation to retrieval or a single static task.Long-horizon benchmarks instead evaluate processes unfolding over time.
- Long-horizon agent evaluation: Vending-Bench permits relatively steady accumulation of successes, whereas this simulator requires significant investments whose payoffs arrive much later.The delayed payoff structure creates a stronger challenge for long-horizon planning.
6 Limitations and Conclusion · Appendix
CEO-BENCH approximates startup operations but omits some real-world complexities and cannot evaluate qualitative product changes reliably. Its results expose a gap between local tool competence and sustained strategic performance under delayed feedback, hidden state, and non-stationarity.
- 6.1 Limitations: CEO-BENCH approximates real-world startup operations and challenges, but discrepancies may remain between the simulation and reality.The authors explicitly frame the benchmark as an approximation rather than a complete reproduction of startup operations.
- 6.1 Limitations: Because qualitative product changes are difficult to evaluate reliably, the simulation represents products using only a quality measure.This design choice narrows how agents can affect products within the environment.
- 6.1 Limitations: Economic feasibility requires limiting the scope of possible actions in each simulation run.The passage presents this restriction as a consequence of making simulations economically feasible.
- 6.1 Limitations: The benchmark leaves out aspects of startup operation such as compliance.The passage identifies compliance as one omitted aspect of the simulated environment.
- 6.2 Conclusion: Existing-model agents can take plausible actions locally but fail when those actions must compound under delayed feedback, hidden state, and non-stationarity.CEO-BENCH therefore distinguishes isolated action quality from sustained strategic competence.
- 6.2 Conclusion: CEO-BENCH reveals a gap between existing models’ local tool competence and crucial sustained strategic skills.The conclusion positions this gap as central to what the benchmark measures.
- 6.2 Conclusion: Future evaluations should test whether agents can organize evolving systems toward distant goals rather than only execute isolated tasks.The conclusion calls for developing agents and training models beyond isolated task execution.
- 6.2 Conclusion: CEO-BENCH is presented as one step toward developing agents and training models for these sustained strategic capabilities.The benchmark is framed as an initial contribution toward that future.
A Simulator Mechanics
This section provides the full details of the simulator mechanics.
- A Simulator Mechanics: The section describes the simulator mechanics in full detail.
A.1 Agent Commands and Observable State
CEO-Bench gives agents a Python action surface for managing pricing, marketing, analytics, research, market intelligence, infrastructure, and enterprise decisions. Agents advance time and query partial observables while important preferences, competitor schedules, and macro conditions remain latent.
- Python action surface: Agents operate the company by importing novamind_api functions and executing Python commands.The action surface is explicitly operational rather than omniscient.
- Operational controls: Pricing, marketing, and analytics commands control offers, promotion, spending, advertising, and targeted operations or development investment.Pricing includes set_prices, set_model_tiers, set_usage_quotas, and set_promotion; marketing includes advertising and social-media actions.
- Information gathering: Research and market commands let agents start projects, inspect research, investigate markets and groups, and retrieve market insights.The interface includes start_research_project, list_research_projects, research_market, research_group, get_market_overview, and get_group_insights.
- Time and observability: Time advances through next_week, while query reads the company database and get_vars reports runtime variables.Observable inputs include dashboards, database tables, social posts, and inbox messages, but true preferences, satisfaction, competitor schedules, and hidden macro state remain latent.
A.2 Customers, Plans, and Participation
The simulator models heterogeneous customers through segment-based parameter sampling with individual variation, market drift, and plausible bounds. Participation and plan choice depend on each customer’s effective price–quality requirements, budget ceiling, and surplus, with cancellation when no plan is acceptable.
- Customer parameter sampling: Customers receive private parameters sampled from group-level distributions, covering willingness to pay, quality thresholds, usage demand, sensitivities, and enterprise traits.This gives each customer both segment-typical characteristics and individual preferences.
- Customer parameter sampling: Market drift changes customer preferences over time, while clipping bounds keep sampled values plausible within realistic ranges.Customers in the same segment therefore resemble one another without being identical.
- Customer participation curve: Customers evaluate plans using differentiated price–quality participation curves: higher prices require higher perceived quality, especially near budget ceilings.The model incorporates willingness to pay, quality floors and ceilings, and heterogeneous price sensitivity.
- Plan choice and billing-period switching: Each customer chooses the acceptable plan with the largest surplus, and cancels on a billing decision day when no acceptable plan remains.Surplus captures perceived quality above the customer’s minimum requirement at the effective price.
A.3 Product Quality, Usage, and Monetization … A.8 Costs and Cash Flow
The simulator couples product quality, customer experience, reputation, discovery, sales, and cash flow into a delayed, noisy operating environment. Agents must balance immediate monetization and operating costs against retention, acquisition, competitive adaptation, and liquidity.
- A.3 Product Quality, Usage, and Monetization: Operations spending buys reliability and capacity tiers provide headroom, but overload still raises outage risk and can degrade delivered quality.The model separates product capability from delivery reliability: a strong product can still feel poor when overloaded.
- A.3 Product Quality, Usage, and Monetization: Diminishing-return development improves quality continuously, whereas R&D delivers larger uncertain gains only after delayed completion.R&D projects move the quality frontier but do not instantly address current churn risk.
- A.4 Satisfaction, Retention, and Support: Satisfaction is an exponential moving average of experienced quality relative to price expectations, so outages and support recovery affect behavior across multiple periods.Voluntary churn occurs when no plan clears a customer’s participation curve, while separate involuntary churn captures background cancellations.
- A.5 Reputation, Social Media, and Acquisition: Social media is a noisy public signal, while reputation aggregates satisfaction and cancellation damage across groups and spills over through related segments.The simulator samples only up to Kpost posts per day, making social media informative but incomplete.
- A.5 Reputation, Social Media, and Acquisition: Acquisition combines reputation, seasonality, macro conditions, social effects, referrals, saturation, and Poisson noise, with temporary demand surges requiring sufficient pricing, quality, and capacity to retain customers.Market availability prevents growth from continuing indefinitely as subscriber counts approach group capacity.
- A.6 Market Discovery and Non-Stationarity: Research improves market information only after delay, while hidden macro cycles and adaptive competitor catch-up make conditions non-stationary and partially observed.Competitors can consume unreleased global or targeted quality gains, with broad improvements attracting stronger catch-up.
- A.7 Enterprise Sales and Negotiation: Enterprise sales evaluates limited plan-price menus and unfolds through counter-offers, response delays, procurement latency, and macro-sensitive deal velocity.Later counter-offers approach the customer’s maximum acceptable price, but negotiations can terminate after the configured turn limit.
- A.8 Costs and Cash Flow: Cash flow subtracts infrastructure, compute, operations, development, marketing, acquisition, and research costs from subscription and ad revenue, coupling growth investment to runway risk.The agent must choose when to burn cash for future growth and when to preserve liquidity to avoid bankruptcy.
B Rule-Based Baseline Strategy and Configuration Search
The rule-based baseline provides a deliberately simple, non-LLM comparison for agent results, using fixed operating policies rather than information-gathering or language-model capabilities. Its configuration search spans 24 combinations of price books, targeting rules, spend packages, and cash floors.
- Baseline strategy: The baseline is a non-LLM point of comparison for agent results in Section 3.2.It is intentionally simple and does not use market research, enterprise negotiation, social media analysis, promotions, or language-model calls.
- Baseline strategy: The policy fixes pricing, product-quality and advertising spend, and customer targeting throughout a run.At run start, it sets prices, model tiers, and usage quotas from one price book.
- Configuration search: 24 total configurations are searched across two price books, two target rules, three spend packages, and two cash floors.The search space varies each of these rule-based baseline components.
C Comparison to Vending-Bench 2 Curves
CEO-BENCH exhibits a delayed-payoff cash trajectory that contrasts with Vending-Bench 2’s steadier accumulation after early volatility. Claude Fable 5’s best CEO-BENCH run falls sharply to fund growth before compounding to $12.6M, whereas Vending-Bench 2 averages growth from $500 to roughly $5.7K.
- C Comparison to Vending-Bench 2 Curves: $626K is the approximate day-30 balance of the best Claude Fable 5 CEO-BENCH run, down from $1M as it funds acquisition and development.The early drawdown reflects spending on growth and product investment.
- C Comparison to Vending-Bench 2 Curves: $5.7K is the Vending-Bench 2 leaderboard endpoint, averaged across five 365-day Claude Fable 5 High runs from a $500 starting balance.Its curve grows steadily after early operating volatility once the agent finds workable suppliers and products.
- C Comparison to Vending-Bench 2 Curves: $12.6M is the later balance reached by the best Claude Fable 5 CEO-BENCH run after its early drawdown, illustrating the simulator’s delayed-payoff structure.The trajectory later compounds after funding growth and product investment.
D Upper Bound Final-Cash Estimate
CEO-Bench estimates an optimistic final-cash upper bound by combining full-market revenue and simulator costs, then discounting the subtotal for execution frictions. The resulting estimate is approximately $2.2B, but it is intended as headroom calibration rather than a proof of optimality or an executable strategy.
- Method: 0.49 is the friction factor applied to the pre-friction subtotal to reflect issue-driven churn, enterprise negotiation friction, and acquisition delay.The factor discounts the clean accounting upper bound toward a more conservative execution path.
- Revenue and costs: $6.69B in pre-friction subscription revenue comprises $1.93B from individual groups and $4.76B from enterprise groups.The calculation covers all 26 customer groups under maximum supportable pricing and modeled retention.
- Interpretation: The accounting assumes full conversion, full enterprise close rates, rapid acquisition, maximum supportable prices, and no additional losses beyond modeled cost categories.These assumptions make the estimate an approximate headroom calculation rather than a demonstrated executable strategy.
- Revenue and costs: $2.2B is the estimated final cash upper bound after adjustments for execution frictions.Table 5 reports this estimate after applying the friction adjustment.