Source-linked AI summary
Can LLMs Take the Pulse of the Economy? A Real-Time Evaluation of LLM Nowcasts on Macroeconomic Indicators
Xinyue Zhao, Ruiyi Zhang, Liqin Ye, Rui Cao, Pengtao Xie, Sudheer Chava
TL;DR
Macroeconomic nowcasting is important but difficult to evaluate for LLMs because historical GDP and CPI releases may be memorized. The paper introduces a live, contamination-resistant benchmark with market-based scoring and finds that LLM agents can match or compete with professional and institutional benchmarks, while remaining limited by sparse U.S. release events and U.S.-only coverage.
Problem
LLM macroeconomic nowcasting lacks reliable historical evaluation because widely reported indicators may be memorized, creating pretraining-contamination risk.
Method
LiveMacroEval runs hourly web-search-enabled LLM nowcasts for sixteen U.S. indicators in strictly pre-release windows and scores them against equity returns and prediction-market trading returns.
Results
Top LLM agents tie Bloomberg ECOS on aggregate LiveMacro Score, and several outperform corresponding Federal Reserve nowcasts on LiveBetting Score for selected indicators.
Takeaways & Limitations
The results indicate strong potential for LLM agents as real-time macroeconomic estimators and provide a template for contamination-resistant evaluation of economically consequential tasks.
Takeaways & Limitations
The live evaluation contains few release events because high-value indicators arrive monthly or quarterly, and all indicators and reference nowcasts are U.S.-denominated.
Abstract
from arXiv · showhide
Nowcasting headline macroeconomic indicators, i.e., estimating an indicator's value for the current reference period before its official release, is critical for monetary policy and financial markets, and central banks devote dedicated teams of expert economists to producing such estimates. Large language model (LLM) agents are a promising candidate for this task, combining broad world knowledge with real-time web search and supporting queries at higher frequency than institutional nowcasts. Evaluating their nowcasting capability is, however, challenging: headline indicators such as GDP and CPI are widely reported and likely memorized during pretraining, so any evaluation on historical releases is vulnerable to data contamination. To address this, we introduce LiveMacroEval, a live, contamination-resistant benchmark in which LLM agents produce hourly nowcasts for sixteen major U.S. macroeconomic indicators over a pre-release window closing at each official release. Nowcast quality is assessed through a LiveMacro Score against announcement-window equity returns and a LiveBetting Score from simulated Polymarket-style trading, with Federal Reserve regional-bank nowcasts, the Bloomberg ECOS professional consensus, and an auto-ARIMA baseline as comparators. Over six months with four state-of-the-art LLM agents configured with web search, aggregate nowcast accuracy is broadly comparable to the institutional and professional benchmarks, with performance varying widely across individual indicators. This highlights LLM agents' potential as real-time estimators of macroeconomic conditions.
1 Introduction
LiveMacroEval addresses contamination risks in LLM macroeconomic nowcasting by evaluating web-search-enabled agents live before official releases. Across six months, results suggest strong potential, with GPT-5 tying Bloomberg ECOS on aggregate LiveMacro Score and several agents competing with Federal Reserve nowcasts for selected indicators.
- Headline macroeconomic indicators matter for monetary policy and asset pricing, but official releases arrive weeks to months after the reference period.
- Historical evaluation is vulnerable to pretraining contamination because widely reported GDP and CPI values may be memorized by LLMs.
- LiveMacroEval introduces a live, contamination-resistant benchmark covering sixteen major U.S. indicators across four thematic blocks.
- The benchmark evaluates hourly pre-release nowcasts using LiveMacro Score and LiveBetting Score, with econometric, Federal Reserve, and Bloomberg ECOS comparators.
- GPT-5 ties Bloomberg ECOS on aggregate LiveMacro Score, while several LLM agents compete with Federal Reserve nowcasts on LiveBetting Score for GDP and unemployment.
2 Related Work
The paper connects macroeconomic nowcasting, LLM forecasting, memorization, and live future-prediction benchmarks. Its focus is the underexplored use of contamination-resistant live evaluation for macroeconomic indicator nowcasting.
- Nowcasting estimates headline indicators before official release and is a standard task in empirical macroeconomics.
- Prior LLM economic and financial forecasting work studies prompted frontier models producing point estimates of macroeconomic and financial variables.
- LLM memorization can reproduce widely reported facts and benchmark instances encountered during pretraining.
- Live future-prediction benchmarks evaluate unresolved events at query time to obtain pretraining-contamination-free assessments.
- LiveMacroEval applies this live-evaluation approach specifically to macroeconomic indicator nowcasting.
3 LiveMacroEval
LiveMacroEval combines live LLM nowcasts, expert and econometric comparators, and market-based scoring for sixteen U.S. macroeconomic indicators. Its strictly pre-release workflow prevents target values from entering training data or retrieved web content, while its metrics translate nowcast quality into equity-return and prediction-market outcomes.
- 3.1 Overview: LiveMacroEval evaluates real-time pre-release estimation across sixteen U.S. indicators organized into four thematic blocks.
- 3.1 Overview: The framework compares LLM agents with Bloomberg ECOS, Federal Reserve regional-bank nowcasts, and an auto-ARIMA baseline.
- 3.4 LLM Agent Nowcasts: Four web-search-enabled agents synthesize heterogeneous mixed-frequency information and decide when and what to search.
- 3.4 LLM Agent Nowcasts: Each prediction window spans weeks to months, closes at the scheduled release, and contains continuous hourly nowcasts without post-release information.
- 3.6 Evaluation Metrics: The LiveMacro Score uses calibrated equity-return effects to compare model-implied returns from predicted surprises with realized announcement-window returns.
- 3.6 Evaluation Metrics: The LiveBetting Score computes cumulative returns from hourly $1 bets on the prediction-market bucket containing each latest nowcast.
4 Results
Across six months, LLM nowcasts were broadly competitive with professional and institutional benchmarks, but performance varied sharply by indicator, theme, and agent design. GPT-5 tied Bloomberg’s aggregate consensus, while inflation-related indicators remained a major weakness and multi-agent design improved results relative to simpler configurations.
- Aggregate LiveMacro Score: +0.004: GPT-5 tied the Bloomberg consensus on the aggregate LiveMacro Score and exceeded the auto-ARIMA baseline.The score aggregates implied surprises across all sixteen indicators, weighting each by its historical causal effect on equity returns.
- Indicator heterogeneity: 87% of the average agent’s LiveMacro shortfall came from CPI, retail sales, and the PCE price index.Value was concentrated on activity and housing-supply series rather than distributed evenly across indicators.
- LiveBetting Score: Four LLM agents achieved positive cumulative LiveBetting returns on real GDP, with three beating the Atlanta Fed, while the leading agent exceeded the Chicago Fed on unemployment.On CPI, the Cleveland Fed and Bloomberg consensus outperformed every LLM agent.
- Theme-level results: LLM agents materially exceeded Bloomberg’s consensus on Supply and Production and Housing, led by GPT-5 and Qwen3-235B, respectively.The theme decomposition shows that LLM performance is specialized rather than uniform across macroeconomic sectors.
- Information-event response: GPT-5’s hourly March 2026 CPI and PCE nowcasts shifted upward within a tight intraday window on April 8, 2026 and then tracked the realized print.Hourly cadence enables ex post tracing of how macroeconomic information events enter the agent’s information set.
- Tool and agent design: The multi-agent team was the only Claude-based configuration with a positive LiveMacro Score, while the tool-augmented agent did not significantly differ from the plain-prompt control.The team predicted smaller surprises and directional outcomes correctly on 54% of releases versus 38% for the plain prompt.
5 Conclusion
The paper introduces LiveMacroEval, a live benchmark for evaluating LLM macroeconomic nowcasting without historical-release contamination. Across six months, top agents matched Bloomberg’s aggregate LiveMacro performance and beat institutional Federal Reserve nowcasts on several LiveBetting indicators, while hourly revisions enabled interpretable analysis of information updating.
- Contribution: LiveMacroEval combines live, contamination-resistant evaluation with LiveMacro Score and LiveBetting Score for macroeconomic nowcasting.The benchmark evaluates LLM agents before official releases rather than relying on potentially contaminated historical releases.
- Main conclusion: Over six months, top LLM agents tied Bloomberg’s professional consensus on aggregate LiveMacro Score and outperformed corresponding Federal Reserve nowcasts on several indicators’ LiveBetting Score.The reported results indicate potential for LLM agents as real-time macroeconomic nowcasting technology.
- Interpretability: Hourly nowcasting supports event-study-style attribution of revisions, enabling interpretable analysis of belief updating and surfacing candidate signals for economic scientific discovery.The framework’s high-frequency cadence connects revisions to information events after the fact.
Limitations
The evaluation is constrained by sparse release events and a U.S.-only scope. These boundaries reflect the requirements of contamination-resistant, temporally valid macroeconomic nowcasting.
- Limitations: The live evaluation contains few release events because high-value indicators arrive only monthly or quarterly.The authors mitigate this through hourly nowcasts, sixteen indicators, multiple expert baselines, and continued operation, but cannot remove the constraint without sacrificing temporal validity.
- Limitations: The benchmark covers only U.S.-denominated indicators and reference nowcasts.The authors identify non-U.S. economies, especially the euro area, as a natural extension because comparable release calendars, market reactions, consensus surveys, and institutional nowcasts exist.
- Motivation and scope: Historical-release evaluation is vulnerable to training-data contamination because headline macroeconomic values may be memorized during pretraining.A strictly live setting avoids this problem because the target has not yet been officially released at prediction time.
- Motivation and scope: Nowcasting matters because policymakers must act amid multi-week publication lags and repeated revisions in headline economic data.Major central banks therefore maintain dedicated nowcast products, while existing institutional nowcasts update weekly or when major source data arrive.
E Post-Release Recovery
Post-release, GPT-5’s nowcasts rapidly align with the official values for CPI, unemployment, and GDP. This recovery indicates that the web-search pipeline can read and report released data soon after publication.
- Recovery pattern: Within the first day after release, the post-release nowcast clusters snap onto the released level for all three evaluated indicators.The residual dispersion remains small relative to both the pre-release cluster spread and the release-versus-prior-cluster gap.
- Indicator examples: The March CPI cluster shifts from roughly 2.6% before release to the 3.3% official value.This is one of three high-attention indicators continued for five days after release.
- Indicator examples: The March unemployment-rate cluster moves from a 4.40–4.45% band to the 4.30% release.The figure marks the release timestamp and released value alongside raw hourly nowcasts and a smoothed trajectory.
- Indicator examples: The 2026Q1 GDP nowcast contracts from a 2.1–2.3% band to the 2.0% official value.The continued series covers the advance estimate released April 30, 2026.
- Interpretation: The web-search pipeline reliably recovers the official value once published and reports it in the requested numeric format.This post-release behavior supports interpreting pre-release dispersion primarily as nowcasting uncertainty rather than retrieval failure.
F Are LLM Agents Copying Institutional Nowcasts?
Three direct tests find no evidence that LLM agents simply copy Bloomberg or Federal Reserve nowcasts. Their levels, revisions, and revision timing diverge from institutional comparators.
- Copying concern: LLM agents cannot directly retrieve the Bloomberg ECOS consensus through web search because it is distributed through the paid Bloomberg Terminal.The open web does not republish the real-time cumulative-median consensus used as the comparator.
- Copying concern: The agents’ hourly revision paths are denser than Federal Reserve nowcasts, which update weekly or on event-triggered schedules.Producing this intra-window path without an LLM agent would require continuous human analyst labor.
- Test 1: level convergence: The median final gap is 0.42σ from Bloomberg, with gaps at least 0.5σ in 46% and at least 1σ in 25% of series.Against Bloomberg, ρS = −0.001, providing no evidence that the gap closes as release approaches.
- Test 2: revision co-movement: The median absolute revision correlation is 0.11, and signs split ten to ten across twenty cells.These results concern co-movement of changes rather than similarity in levels.
- Test 3: revision timing: None of the three event-time tests finds the positive revision concentration expected from copying.The two significant cells are both anti-clustering, with ∆share < 0; GPT-5’s Atlanta Fed GDPNow result is also anti-clustering at ∆share = −0.04.
G What the Nowcast Rationales Contain
The rationale analysis shows that agents usually cite official sources, often exercise judgment, and frequently adjust computed values into forecasts. Arithmetic checks largely validate the reported calculations, but forecasts commonly depart from them.
- Reported behavior: 56.2% of rationale blocks name a statistical agency, while 47.0% contain explicit judgment over the agent’s figure.The analyzed corpus contains 7,078 self-reported rationale blocks from 305 nowcast runs.
- Citation composition: 98.6% of rationale blocks name sources, and official statistical sources account for 62.1% of citations across 167 domains.Census, BEA, BLS, Federal Reserve, and ISM sites are the leading official domains; consensus aggregators form a smaller share.
- Arithmetic checking: 93.4% of 652 explicit calculations match the reported figure within a 2% tolerance.A second independently written checker agrees on 94.7% of calculations, including many multi-step derivations.
- Forecast adjustment: The forecast differs from the computed value in 69.5% of 604 blocks containing both.Adjustment language appears in 40.0% of departures, and forward-period language appears in 40.5%.
- Baseline context: The benchmark uses a per-indicator auto-ARIMA point forecast fit on the latest pre-target FRED vintage and held fixed throughout each prediction window.The model selects ARIMA specifications per indicator and does not condition on intra-window information.
J Nowcast Error in Real Time
Real-time error comparisons show that model performance varies sharply by indicator and theme: consensus dominates several series, while continuous LLM updating helps on others near release. The section also defines the scale-normalized error measure and the market-oriented LiveMacro evaluation framework.
- Error measurement: Squared relative error normalizes each indicator’s forecast miss by its mean absolute release value, making errors comparable across units.The metric is reported separately by indicator rather than summed into a cross-indicator ranking.
- Supply and Production: Consensus remains strong for real GDP and durable goods orders, while Claude-sonnet and Qwen3 narrowly meet or cross consensus for industrial production and ISM manufacturing near release.Model errors decline sharply toward release for the latter two indicators.
- Demand and Inflation: Consensus errors are near zero for CPI, PCE price index, real PCE, and retail sales, but every model beats consensus on PPI and all models beat it for ISM services at resolution.For ISM services, consensus error rises near release while model errors decline.
- Labor Market: Every model beats consensus on nonfarm payrolls, whereas only Claude-sonnet and GPT-5 cross below consensus for the unemployment rate.Consensus remains large for payrolls but tightly tracks the unemployment rate.
- Housing: Claude-sonnet performs consistently well across housing series, while GPT-5 and Qwen3 are weakest in this theme.Model nowcasts continue updating and typically fall below flat consensus error near housing releases.
- LiveMacro evaluation: The LiveMacro Score uses predicted macro surprises to explain announcement-window equity returns, with field-specific sensitivities estimated from historical events and a bounded range of [−1, +1].A score of 0 matches the consensus reference, while +1 represents perfect anticipation of realized announcement-window returns.
K.3 Inference and robustness
Inference and robustness analyses show that headline LiveMacro rankings are broadly stable under alternative historical fits, while conventional aggregate error sums can be dominated by scale or near-zero denominators. Confidence intervals also expose an information-set mismatch that disadvantages LLMs under one sampling construction.
- Inference: The sampling-based confidence construction lowers point estimates relative to the headline parametric-bootstrap results, including GPT-5 moving from +0.004 to −0.010.The construction uses five recent pre-release nowcasts as samples.
- Inference: The sampling interval systematically disadvantages LLM agents because their latest five hourly samples reach roughly six hours before release, while Bloomberg consensus incorporates submissions until announcement.This breaks the terminal-snapshot alignment of information sets.
- Inference: Qwen3-235B is closest to consensus at −0.007 with a 90% interval of [−0.033, +0.009], narrowly ahead of GPT-5 at −0.017.The interval for Qwen3-235B straddles zero.
- Robustness: Robustness movements are uniformly below 0.025 in absolute value, with essentially no change at the top of the ranking.Qwen3-235B has the largest reported shift, +0.023, comparable to its bootstrap uncertainty.
- Error aggregation: Summed raw squared errors are unusable across indicators because more than 99.99% comes from five level series, with existing home sales alone contributing 45%.The resulting ranking primarily reflects housing and payroll levels rather than inflation or activity performance.
- Error aggregation: Relative-error aggregation shifts the distortion to near-zero denominators: PPI contributes 42% of the relative-error sum while CPI contributes 0.1%.The effective number of indicators under economic LiveMacro weights is 4.7, versus 16 under equal weighting.
- Evaluation design: The LiveMacro framework weights indicators by historical surprise-to-return sensitivities, whereas simulated betting prices each bucket using the crowd’s pool-implied probability.This links evaluation either to announcement-window market impact or to prediction-market returns.
M Detailed Analysis of LiveMacro Score Results
GPT-5 is the strongest aggregate LiveMacro performer, tying Bloomberg ECOS, while model performance varies substantially across indicators and directional accuracy.
- Aggregate results: GPT-5 scores +0.004 against Bloomberg ECOS, implying an approximately 0.8% reduction in predicted announcement-window equity-return-shock mean-squared error.Its 90% bootstrap interval narrowly straddles zero.
- Aggregate results: Claude-sonnet-4.5 and Qwen3-235B score approximately −0.028, while auto-ARIMA and Qwen3-80B perform worst at −0.118 and −0.122.The latter models have overlapping 90% bootstrap intervals.
- Per-indicator decomposition: The average agent contributes positively on 4 of 16 indicators, led by ISM manufacturing (+0.0012), existing home sales (+0.0011), building permits (+0.0007), and industrial production (+0.0004).These are activity or housing-supply series.
- Per-indicator decomposition: CPI, retail sales, and PCE price index contribute −0.0380 together, accounting for 87% of the average agent’s aggregate score of −0.0437.The average contributions are −0.0228, −0.0077, and −0.0075, respectively.
- Per-indicator decomposition: The per-indicator best model is positive on 12 of 16 indicators, with GPT-5 leading eight and each other agent leading at least two.Best-model identity changes across indicators.
- Directional performance: GPT-5 leads the directional score at 0.572, and it, Qwen3-80B, and auto-ARIMA are the only models at or above the 0.5 coin-flip reference.Claude-sonnet-4.5 and Qwen3-235B score below coin-flip, with Qwen3-235B last at 0.388.
O Detailed Analysis of LiveBetting Score Results
Simulated hourly betting reveals substantial indicator-specific differences: LLMs can outperform institutional references on GDP and unemployment, but inflation remains led by Cleveland Fed and Bloomberg consensus.
- Scoring interpretation: Cumulative betting returns combine hourly $1 stakes with inverse-price payoffs, so long windows can generate large returns through many separate bets.Accuracy and early timing can both increase returns when agents identify the eventual bucket before market repricing.
- Real GDP: The New York Fed Staff Nowcast is strongest on real GDP, while Qwen3-235B, Qwen3-80B, and Claude-sonnet-4.5 rise above the Bloomberg consensus before release.Three LLM agents beat Atlanta Fed GDPNow.
- Headline CPI: The Cleveland Fed Inflation Nowcasting model achieves the highest cumulative CPI betting return, followed by Bloomberg consensus and the two Qwen variants.GPT-5 ends the two-month aggregation negative, while Claude-sonnet-4.5 performs worst.
- Headline CPI: Cleveland Fed and Bloomberg returns decline through the CPI window partly because early profits are diluted across a growing number of break-even bets.The Cleveland Fed falls from +138% to +91.5%, while Bloomberg falls from +66% to +44%.
- Headline CPI: Around the March CPI shock, Qwen3-80B under-reacts, Qwen3-235B over-reacts in the wrong month, and GPT-5 suffers many noisy misses.These distinct timing patterns produce different betting outcomes despite the same event window.
- Unemployment rate: Claude-sonnet-4.5 is strongest on unemployment, materially exceeding every other LLM agent and the institutional reference.GPT-5 follows, while Qwen3-80B, Qwen3-235B, and Bloomberg consensus reach the −100% floor.
P Detailed Analysis of LiveMacro Score by Theme
GPT-5 is the most consistent cross-theme performer, while gains concentrate in supply, production, and housing; demand, inflation, and labor remain more difficult.
- Per-model breakdown: GPT-5 scores +0.057 on Supply and Production, approximately −0.003 on Demand and Inflation, approximately 0 on Labor Market, and −0.021 on Housing.It matches or ties Bloomberg consensus in three of four themes.
- Cross-theme numerical detail: On Demand and Inflation, only GPT-5 sits at consensus, while the other LLM agents and auto-ARIMA show the largest negative gaps.This theme-level pattern complements the aggregate indicator decomposition.
- Cross-theme numerical detail: Labor Market scores cluster within ±0.03 across every model, consistent with a signal-to-noise ceiling for unemployment-rate information.The labor theme shows little separation among models.
Q Cross-Model Error Analysis
The error analysis identifies two shared weaknesses—anchored inflation and sharp reversals—alongside model-specific calibration and updating patterns; multi-agent design is the only configuration clearing consensus.
- Shared failures: Across five CPI months, consensus misses by at most 0.1 percentage point in four, while it is exactly right on four of five PCE prints.This leaves little room for agents to add signal, especially against Cleveland Fed’s purpose-built model.
- Shared failures: All four agents miss several sharp reversals in the same direction, including March industrial production and March retail sales.They predict growth into an industrial-production contraction and understate a +1.7 retail-sales print with forecasts between +0.4 and +0.6.
- Per-model patterns: GPT-5 leads at +0.004 and has the highest directional score at 0.572, but systematically under-predicts housing levels.Its January housing-starts nowcast was 1.26M versus a 1.49M release, with limited market-score cost because housing sensitivities are small.
- Per-model patterns: Claude-sonnet-4.5 is a high-variance overshooter: its advance Q4 GDP nowcast was 4.24% against a 1.4% release.Its LiveMacro standing therefore relies on cautious magnitude calibration rather than directional skill.
- Agent configurations: The tool-augmented agent leads the plain-prompt control on both scores, but paired intervals contain zero, so the comparison is treated as a tie.The paired bootstrap accounts for correlated errors on identical releases.
- Agent configurations: The multi-agent team is the only configuration that clears the consensus baseline, with a significant advantage over both other configurations on the consensus-relative score.Its advantage on LiveMacro Score points the same way but is not significant.
S.4 Why the multi-agent team clears the baseline
The multi-agent team clears the baseline by combining restrained position sizing with directional accuracy. Its design also reconciles independent estimates, while the comparison remains limited by a short live evaluation and configuration costs.
- Sizing discipline: The multi-agent team is the only configuration with negative magnitude–error correlation, indicating that its largest predicted surprises are not its largest misses.For the plain prompt, plug-in, search agent, and ARIMA baseline, the correlation is strongly positive.
- Design decomposition: The verifier reduces mean predicted surprise magnitude by a factor of 3.7, whereas adding the plug-in leaves it essentially unchanged.The verifier also lowers correlation with the plain prompt from 0.60 to 0.33, but these comparisons are descriptive associations on small samples.
- Design decomposition: The orchestrator reconciles two independent sub-agent passes into one estimate, so disagreements produce a value between the two estimates.Completed runs delegate to both sub-agents, and the verifier accounts for 49% of retrieval.
- Market-impact heterogeneity: The tool helps where coverage is deepest: relative error is 0.206 versus 0.397 on high-market-impact releases, but 0.635 versus 0.525 on the other half.The releases are split at the median market-impact weight |β_i|.
- Scope and cost: The comparison spans about one month, uses one live configuration per design, and the multi-agent team costs roughly fifteen minutes per run.These constraints motivate paired intervals rather than relying only on a bare point estimate.