Source-linked AI summary
Same Day, Same Story; One Day Ahead, a Different Signal: The Dual Validity of Financial Sentiment
AS Aravinthkakshan, Laven Srivastava, Harsh Nandwani
TL;DR
Financial NLP commonly treats agreement with human sentiment labels as evidence that an instrument will extract market signal, although the two evaluations are usually conducted on different corpora. This paper links 70,500 X messages from securities class actions to abnormal returns, evaluates five instruments through one pipeline, and finds that the relationship between construct and predictive validity depends on sampling and score representation. Agreement establishes semantic validity but does not by itself determine predictive rankings, while content informs where message volume does not.
Problem
Financial sentiment tools are usually validated against human labels separately from market outcomes, leaving unclear whether construct validity implies predictive validity.
Method
The paper links securities-class-action messages to abnormal returns and evaluates five sentiment instruments, human labels, and shared versus method-specific samples through one pipeline.
Results
The relationship between human agreement and return association varies with sampling, horizon, score representation, and zero-score inclusion, while content informs where volume does not.
Takeaways & Limitations
Benchmark agreement is evidence of semantic validity but does not by itself justify selecting an instrument for predictive use.
Takeaways & Limitations
The rank comparisons cover only five instruments and are therefore descriptive.
Abstract
from arXiv · showhide
Financial NLP has a standard workflow: validate a sentiment tool against human labels, then trust it to extract market signal. This assumes the two evaluations measure the same thing. We test that assumption in a setting where both can be measured at once: a corpus of securities class actions (2002-2025) linking 70,500 X messages to abnormal stock returns, with a single-annotator human labelled gold sample. Running five instruments (VADER, Loughran-McDonald, FinBERT, Twitter-RoBERTa, and an LLM annotator) through one identical pipeline, we find that the relationship between construct and predictive validity depends on the sampling convention and score representation. Under conventional method-specific sampling, human agreement aligns more closely with graded same-day associations than with one-day leads. On a fixed-n panel, however, agreement has similar graded rank correlations at both horizons, while the coarse ordering remains weak. Benchmark agreement therefore establishes semantic validity but does not by itself determine predictive rankings. In a conversation that is 17.6% spam, message volume predicts neither market damage nor settlement size.
1 Introduction
The paper tests whether agreement with human sentiment labels implies usefulness for predicting market outcomes by measuring both validities on the same financial-social-media messages. It finds that their relationship varies with sampling, horizon, and score representation, while content carries information where message volume does not.
- Motivation: The study tests the assumption that human-label agreement identifies the best instrument for extracting market signals.It addresses the usual separation between labeled sentiment benchmarks and market datasets by evaluating both criteria together.
- Motivation: The domain combines genuine investor discussion with solicitation, repeated headlines, automated promotion, and 17.6% spam, making raw counting a poor measure by construction.Solicitation and spam together account for 25,166 messages.
- Contributions: 66,890 event-anchored messages from 845 securities class actions are released with five instrument labels, multi-axis LLM annotation, and a 400-message human gold standard.The corpus links the same messages to human annotation and market outcomes through one pipeline.
- Findings: Under conventional method-specific sampling, human agreement aligns more closely with graded same-day associations than with one-day leads, while fixed-n graded rankings are similar across horizons but coarse rankings remain weak.The relationship depends on horizon, representation, and treatment of zero-score days.
- Findings: The largest one-day point estimate comes from an instrument pretrained on financial news, while news accounts carry the strongest attributed signal (ρ = −0.261).A local-projection placebo fails at negative horizons.
- Findings: Content is informative where volume is not: message volume predicts neither the depth of the price decline nor dollar settlement size.This is the paper’s domain-level informative null.
2 Related Work
Prior work distinguishes agreement with human labels from association with market outcomes, but these evaluations are rarely performed on the same financial messages. This paper situates a linked litigation corpus and standardized comparison within that gap.
- Content and volume in financial discourse: The study separates message content from volume, motivated by prior evidence that content carries return information while posting activity mainly captures attention or volatility.A count-based attention measure can combine informative reporting, investor reaction, solicitation, and automation, motivating a separate volume comparison.
- Sentiment instruments and their benchmark evaluation: Conventional sentiment benchmarking emphasizes agreement with human annotation, using Financial PhraseBank and newer financial-language-model suites as reference paradigms.The paper contrasts these standard benchmark practices with evaluation against market outcomes.
- Intrinsic versus extrinsic evaluation: Agreement with human labels and association with market outcomes represent distinct intrinsic and extrinsic evaluations, whose relationship is rarely tested on the same financial messages.The paper frames this distinction as construct versus predictive validity and connects it to broader measurement critiques in NLP.
- Social media, returns, and securities litigation: The corpus combines securities class-action filings, event windows, market returns, litigation metadata, and social-media messages to study event-anchored financial discourse.Court filings define company identities, class-period bounds, corrective disclosures, allegations, defendants, and retrieval windows.
- Sentiment instruments and their benchmark evaluation: Five instruments span a social-media lexicon, finance dictionary, financial-news and tweet-pretrained transformers, and an LLM annotator.The comparison includes VADER, Loughran–McDonald, FinBERT, Twitter-RoBERTa, and Claude Haiku, with common aggregation and testing procedures.
- Intrinsic versus extrinsic evaluation: The human gold evaluation covers polarity and relevance but not intensity, emotion, severity, or topic, and uses one annotator without a measured human-agreement ceiling.The paper therefore treats the evaluation as agreement with intended linguistic constructs rather than a complete validation of every instrument output.
4 Empirical Setup
The study compares sentiment with abnormal returns using rank correlations, lead–lag and Granger tests, while distinguishing predictive precedence from causation. Analyses use a panel of cases and case days, with conventional method-specific samples and a fixed-n intersection.
- Spearman ρ measures monotonic association, with daily analyses pooling case days and cross-sectional analyses collapsing each case to one observation.Because the daily panel is large, instruments are compared primarily on effect size rather than p-values.
- Lead–lag correlations compare sentiment at t − ℓ with abnormal returns at t, while second-order Granger tests assess incremental predictive content beyond return history.The broader timing design also includes distributed-lag and local-projection specifications.
- These tests establish predictive precedence rather than structural causation, and between-instrument ρ differences remain descriptive point-estimate comparisons.Significance for one instrument but not another does not establish a significant difference between them.
- The panel contains 179 cases and 103,542 case days, while conventional correlations exclude method-specific zero-score days.A common-day intersection holds n fixed for comparison.
5 Predictive Validity: Content versus Counting
The daily panel tests whether message content carries stronger same-day market associations than message volume. Sentiment representations generally exceed tweet volume in association strength, although the result varies by scoring choice.
- Message volume is uncorrelated with worst single-day abnormal return (ρ = −0.083, p = 0.60, n = 43) and CAR (ρ = +0.003, p = 0.99, n = 36).These case-level results come from the 2025 cohort.
- Every sentiment family contains a representation with a larger association than tweet volume, whose ρ is −0.0311; the strongest instruments exceed volume by two to three times.Loughran–McDonald coarse is the exception at 0.80×, so the result applies to method families rather than every threshold choice.
- Impressions, a reach-based count, are not significant in the daily comparison.
- Twitter-RoBERTa, rather than the LLM, is the strongest same-day instrument, so the content-over-counting result does not depend on LLM annotation.The comparison uses the ρ column because p-values differ partly with row-specific sample sizes.
6 Temporal Structure
Temporal associations differ across instruments under conventional method-specific sampling, while the visual comparisons emphasize same-day and lead–lag correlations across score representations. The reported comparisons are point estimates and do not establish causal effects or significant differences between instruments.
- 6 Temporal Structure: Every instrument shows a negative and significant same-day association, so the contemporaneous result is not instrument specific.Table 4 reports the lead–lag correlations for all messages, with coarse and graded representations shown separately.
- 6 Temporal Structure: The refreshed price panel reproduces published values within |∆ρ| ≤ 0.002, with identical signs and significance ordering.All cross-method comparisons use the refreshed same data panel.
- 6 Temporal Structure: FinBERT coarse has the largest one-day point estimate (−0.0466, pFDR = 0.003), while FinBERT graded and Twitter-RoBERTa graded also survive FDR.VADER and Loughran–McDonald do not survive FDR at one day; at two days, Claude graded and Twitter-RoBERTa graded are significant at the unadjusted 5% level.
- 6 Temporal Structure: Four baseline specifications survive FDR on forward Granger tests, while Twitter-RoBERTa never survives despite its strongest same-day association.The passing specifications include Loughran–McDonald, FinBERT, and the LLM reference.
- 6 Temporal Structure: Figure 1 separates conventional method-specific sampling from a fixed 134-case comparison of event-minus-baseline sentiment shifts versus CAR.Panel (a) uses strongest same-day representations with method-dependent n; panel (b) uses identical cases and reports exact ρ values.
7 Where the Two Validities Diverge
The relationship between human-agreement rankings and predictive associations changes with sampling convention, horizon, and score representation. Conventional samples favor graded same-day alignment, while fixed-n panels show similar graded correlations across horizons and weak coarse relationships.
- 7 Where the Two Validities Diverge: Human-agreement rankings align most closely with graded same-day associations under conventional method-specific sampling, while one-day alignment is weaker and representation-dependent.Table 5 reports Spearman rank correlations across five instruments; these comparisons are descriptive and do not survive adjustment for 12 comparisons.
- 7.1 Instrument rankings are unstable: The conventional analysis excludes zero-score case days separately by instrument, whereas the fixed-n panel retains only days with nonzero scores for every instrument.Because the retained days differ, method rankings can change with sample construction; on common days, FinBERT is strongest and VADER is indistinguishable from zero.
- 7 Where the Two Validities Diverge: None of the 12 fixed-n construct–predictive correlations survives BH adjustment, so the observed rank relationships remain descriptive.The exact permutation p-values are reported for κ, accuracy, and macro-F1 comparisons.
- 7.2 Divergence at the case level: At the case level, the LLM instrument’s event-minus-baseline negativity change correlates with event-window CAR at ρ = −0.3106, but the strongest instrument is not established.The comparison uses n = 134 cases; the ordering is underpowered and sensitive to how the case-level measure is constructed, while restricting to relevant messages makes every instrument nonsignificant.
8 What the Leading Signal Responds To
The leading market association is concentrated in news-related content rather than clearly reflecting investor-expression polarity. FinBERT produces the largest conventional one-day estimate despite near-chance agreement with human polarity, while news accounts have the strongest account-level association.
- 8 What the Leading Signal Responds To: The relationship between construct validity and the one-day lead depends on both the sampling convention and score representation.The paper examines two pieces of evidence about what the observed leading associations reflect.
- 8 What the Leading Signal Responds To: FinBERT has the largest conventional one-day point estimate despite near-chance agreement with human polarity on securities-fraud tweets.Because FinBERT is pretrained on financial news, the paper interprets this pattern as potentially detecting adverse-news vocabulary rather than investor-expression polarity.
- 8 What the Leading Signal Responds To: News accounts have the strongest association with returns at ρ = −0.261, followed by cross-case broadcasters at −0.161 and other multi-case accounts at −0.115.Bot/spam accounts (+0.064) and law firm solicitation (+0.040) are not significant.
9 Robustness
Robustness checks support the central daily-panel findings while delimiting their scope: filtering is nonessential, volume is uninformative, and contamination and generalization remain unresolved.
- 9 Robustness: The LLM’s daily associations are similar across archive and recent cohorts, but the small case-level subsample leaves contamination unresolved.The daily panel is the basis for the paper’s claims; the case-level result is treated as corroborative.
- 9 Robustness: Filtering to relevant messages shifts same-day correlations by at most 0.0164, leaving the central result intact.The relevance labels come from the LLM instrument, but the robustness result reduces concern that filtering creates the finding.
- 9 Robustness: Message volume is uncorrelated with settlement amount, relevant-message volume, and negative-message volume across 133 resolved cases.Median settlements also do not differ across volume tertiles, supporting the paper’s distinction between content and attention.
- 9 Robustness: The corpus’s return relationships are sensitive to sampling and representation, so human agreement establishes semantic validity but does not determine predictive instrument rankings.The comparison is descriptive and should report horizon, score representation, and zero-day inclusion alongside both validity measures.
- 9 Robustness: Human validation covers polarity and relevance through a single annotator, not intensity, emotion, severity, topic, or inter-annotator reliability.The reported κ values therefore have no measured human-to-human ceiling.
- 9 Robustness: The findings may not generalize beyond securities-class-action discourse and may be attenuated by unresolved archive tickers and possible survivorship bias.This setting contains unusually prevalent solicitation, repeated headlines, and cashtag piggybacking.
Ethics Statement
The released materials protect privacy by distributing identifiers and annotations rather than message text, while the study uses public posts and court records.
- Ethics Statement: The corpus uses public X/Twitter posts and public court filings without author-profile lookups or personally identifying information in released artifacts.Account categories are derived from posting behaviour and attribution results are reported only at the aggregate class level.
- Ethics Statement: The reproducibility archive includes identifiers, annotations, sentiment series, abnormal returns, and analysis code rather than message text.The abnormal-return panel is included directly because 74 of 273 archive tickers no longer resolve through the original price API.
A Relevance Filtering
Relevance filtering does not generate the daily result, although it makes case-level windows sparse; the supporting lag tables also warn that coefficients are not cross-instrument comparable.
- A Relevance Filtering: Filtering to LLM-labelled relevant messages shifts same-day correlations by at most 0.0164 and never reverses a sign.This supports the conclusion that the daily finding is not produced by relevance labels, despite their potential circularity.
- A Relevance Filtering: Distributed-lag coefficients are not comparable across instruments because their underlying signals differ in scale; only sign and significance transfer.The complete battery is reported with all-message and relevant-message sample sizes.
- A Relevance Filtering: On the common relevant-message panel, FinBERT is strongest same day on the coarse measure, while Twitter-RoBERTa leads on the graded measure and VADER remains indistinguishable from zero.Table 11 repeats the common-day intersection using relevant messages only.
E Retained Null Results
The retained results show an information shock and contemporaneous market association, while placebo failures, weak multiple-testing evidence, and null volume effects limit causal interpretation.
- E Retained Null Results: Total, relevant, and negative message volume fail to predict settlement size or case-level price damage.This null is retained alongside insignificant fear, panic, and outrage shares and weakened event-window significance after controls.
- F Event Validation and Local Projections: 46.7% negative messages appear in disclosure windows versus 11.8% in baseline windows, while controlled class-period negativity predicts maximum drawdown.The disclosure-window shift is statistically strong, and the controlled drawdown model reports adjusted R2 = 0.60.
- F Event Validation and Local Projections: Local projections peak contemporaneously at h = 0, remain significant at h = 1, and are indistinguishable from zero by h = 2.Negative-horizon coefficients are also significantly negative, and the reverse next-day correlation is present, so causal interpretation is limited.
- E Retained Null Results: Across agreement measures, alignment is closest with graded same-day associations; one-day relationships are weaker or weakly negative, and none survives multiple-testing correction.The five-instrument comparisons are descriptive, with exact permutation tests used for the rank correlations.
- E Retained Null Results: Negative-heavy days cluster around major price declines while message volume remains similar in negative and quiet periods.The Target Corporation example illustrates that classified sentiment, rather than counting alone, aligns with price movements.