Source-linked AI summary
Conv-FinRe: A Conversational and Longitudinal Benchmark for Utility-Grounded Financial Recommendation
Yan Wang, Yi Han, Lingfei Qian, Yueru He, Xueqing Peng, Dongji Feng, Zhuohan Xie, Vincent Jim Zhang, Rosie Guo, Fengran Mo, Jimin Huang, Yankai Chen, Xue Liu, Jian-Yun Nie
TL;DR
Financial recommendation benchmarks can mistake noisy behavioral imitation for decision quality when investor actions conflict with long-term goals. Conv-FinRe introduces a conversational, longitudinal benchmark with multi-view, investor-specific references for ranking stocks over a fixed horizon. Its evaluations reveal a persistent tension: utility-oriented models may diverge from user choices, while behaviorally aligned models may overfit short-term noise.
Problem
Existing recommendation benchmarks commonly use user behavior as ground truth, but financial actions can be noisy or short-sighted and may not reflect long-term risk preferences or objectives.
Method
Conv-FinRe evaluates stock rankings from onboarding interviews, step-wise market contexts, and advisory dialogues against user-choice, rational-utility, market-momentum, and risk-sensitivity references.
Results
Evaluations reveal a persistent tension between utility-based ranking and behavioral alignment: general-purpose models often optimize utility more effectively, whereas domain-specific models tend to overfit transient user actions.
Takeaways & Limitations
The benchmark supports evaluation that separates rational decision quality from observed behavior and exposes the limits of behavior-only assessment.
Takeaways & Limitations
The latent preference signal assumes user-specific volatility and downside-risk parameters are time-invariant and is used only as a reference representation, not exposed to the model.
Abstract
from arXiv · showhide
Most recommendation benchmarks evaluate how well a model imitates user behavior. In financial advisory, however, observed actions can be noisy or short-sighted under market volatility and may conflict with a user's long-term goals. Treating what users chose as the sole ground truth, therefore, conflates behavioral imitation with decision quality. We introduce Conv-FinRe, a conversational and longitudinal benchmark for stock recommendation that evaluates LLMs beyond behavior matching. Given an onboarding interview, step-wise market context, and advisory dialogues, models must generate rankings over a fixed investment horizon. Crucially, Conv-FinRe provides multi-view references that distinguish descriptive behavior from normative utility grounded in investor-specific risk preferences, enabling diagnosis of whether an LLM follows rational analysis, mimics user noise, or is driven by market momentum. We build the benchmark from real market data and human decision trajectories, instantiate controlled advisory conversations, and evaluate a suite of state-of-the-art LLMs. Results reveal a persistent tension between rational decision quality and behavioral alignment: models that perform well on utility-based ranking often fail to match user choices, whereas behaviorally aligned models can overfit short-term noise. The dataset is publicly released on Hugging Face, and the codebase is available on GitHub.
1 Introduction
Financial recommendation cannot be evaluated reliably through behavioral imitation alone because market noise and investor preferences can make observed choices diverge from long-term utility. Conv-FinRe addresses this gap with conversational, longitudinal, multi-view evaluation and reveals tension between utility-based quality and behavioral alignment.
- Financial actions may reflect short-term noise, emotions, and shifting constraints rather than stable risk tolerance or long-term objectives.
- Existing benchmarks often treat clicks, ratings, or choices as truth, leaving utility grounding and diagnosis of risk-sensitive reasoning unresolved.
- Conv-FinRe evaluates stock rankings against user choice, rational utility, market momentum, and risk sensitivity over conversational, longitudinal interactions.
- User-specific risk preferences are inferred from longitudinal decision trajectories through inverse optimization to construct utility- and risk-based reference rankings without exposing latent utility to models.
- Models show a persistent tension between rational decision quality and behavioral alignment: utility-strong models can conflate long-term risk with momentum, while specialized models overfit noisy actions.
- The benchmark contributes a diagnostic framework for separating behavioral alignment from rational decision quality and identifying advisory patterns under competing signals and user noise.
2 Related Works
Prior recommendation research largely studies consumer personalization from interaction histories or coarse user traits, using single relevance signals and increasingly structured or sequential evaluation protocols.
- Consumer recommendation benchmarks commonly model personalization from interaction histories or coarse user traits, supervised by ratings or clicks.
- Recent work evaluates structured explanations and LLM recommenders through point-wise, pair-wise, and list-wise protocols.
- Related analyses also address fairness, bias, and sequential alignment in recommendation systems.
3 Conv-FinRe
Conv-FinRe models financial recommendation as a conversational, longitudinal, multi-view alignment task grounded in user-specific risk preferences. It combines market data, user trajectories, simulated advisory dialogues, and latent utility estimation to evaluate LLM rankings.
- 3.1 Task Formulation: Conv-FinRe evaluates personalized financial recommendations across User Choice, Rational Utility, Market Momentum, and Risk Sensitivity reference views.The task simulates iterative advisor–user interactions over a fixed investment horizon and compares model rankings with complementary behavioral and normative signals.
- 3.1 Task Formulation: The benchmark provides each decision step with current market context, prior interaction history, and recommendations from three specialized advisors.The longitudinal history includes user decision patterns and conflicting Rational Utility, Market Momentum, and Risk Sensitivity signals.
- 3.2 Latent Preference Grounding via Inverse Optimization: Latent user preferences are represented by a utility function balancing expected return against volatility and downside risk.User-specific volatility and downside-risk sensitivities are assumed time-invariant and estimated from longitudinal behavior using inverse optimization.
- 3.3.1 Data Collection: The benchmark constructs a controlled stock universe from S&P 500 constituents, grouping candidates by market beta across low, moderate, and high-risk regimes.The resulting universe contains ten representative stocks, while market states use daily and intraday price data over a 30-day horizon.
- 3.3.1 Data Collection: User data combine questionnaire-based profiles with longitudinal buy decisions and portfolio feedback collected through an asset simulation tool.Ten participants provide demographic, financial, experience, and risk-attitude information, while actions and realized returns and volatility form temporally ordered traces.
- 3.3.2 Conversation Simulation: Observed profiles and trajectories are transformed into reproducible conversations with onboarding interviews followed by longitudinal advisory dialogues.The onboarding dialogue captures financial background, constraints, goals, and emotional reactions to risk; later turns condition on prior history and current market state.
4 Experiments and Results
Experiments expose distinct ways models balance rational utility, user behavior, market momentum, and risk sensitivity. Conversational history improves some models’ utility alignment, but gains vary: some integrate preferences, some rely on contemporaneous signals, and others overfit noisy actions.
- Overall Performance: Most models achieve high uNDCG scores (0.92–0.97) for Rational Utility, but strong utility rankings do not consistently recover User Choice.Llama-3.3-70B-Instruct leads in uNDCG (0.97) while showing lower Hit Rates, whereas Qwen2.5-72B-Instruct and Llama3-XuanYuan3-70B-Chat excel in MRR and HR@K.
- Expert Alignment Analysis: Llama-3.3-70B-Instruct couples high Utility and Market Momentum alignment with sharply lower Risk alignment, while DeepSeek-V3.2 maintains the most balanced profile.The results attribute this contrast to difficulty decoupling downside protection from growth-oriented signals during trending markets.
- Preference Discovery Dynamics: GPT-5.2 and DeepSeek-V3.2 show significant positive utility-alignment gains during early-to-middle conversational steps, especially steps 1–10.Across most models, later fluctuations and plateauing indicate that coarse preference representations form early while consistent long-term tracking remains challenging.
- Preference Discovery Dynamics: Once a stable investor persona is established, additional historical context has diminishing marginal utility and ranking depends more on immediate market context.This connects conversational preference discovery with the observed plateau in utility-based alignment.
- Preference Discovery Dynamics: Figure 3 identifies Adaptive Advisors, Transaction-driven Analysts, and Behavioral Overfitters according to how utility alignment changes when conversational history is available.Adaptive Advisors improve with history, Transaction-driven Analysts remain near the diagonal, and Behavioral Overfitters degrade, with XuanYuan’s drop highlighting non-uniform preference discovery.
5 Conclusion
Conv-FinRe shifts financial recommendation from behavioral matching toward utility-grounded decision alignment. Using inverse optimization and multi-view evaluation, it separates rational decision quality from observed behavior and finds persistent tension between utility alignment and behavioral alignment.
- 5 Conclusion: Conv-FinRe shifts stock recommendation evaluation from surface behavioral matching to utility-grounded decision alignment.The benchmark is conversational and longitudinal, rather than limited to static behavior imitation.
- 5 Conclusion: Inverse optimization of latent risk preferences supports multi-view evaluation that separates rational decision quality from observed user behavior.The framework distinguishes utility-oriented alignment from alignment with transient user actions.
- 5 Conclusion: General-purpose LLMs often optimize utility more effectively, whereas domain-specific models tend to overfit transient user actions.These findings expose limits of behavior-only evaluation and motivate benchmarks that disentangle long-term investor preferences from observed behavior.