Source-linked AI summary
One AI Signal, Many Human Judgments: A Bayesian Cascade Analysis of AI-based Credibility Indicators in Online Information Spread
Zhuoran Lu, Weilong Wang, Yangyang Yu, Xinru Wang, Zhuoyan Li, Zhiwei Liu, Sophia Ananiadou
TL;DR
AI credibility indicators create a sequential human-AI problem because users observe judgments shaped by the same AI prediction. The paper extends Bayesian cascade theory with a shared public signal and calibrates it with human-subject data and simulations. The analysis finds a preservation–correction trade-off, heterogeneous reliance including cascade-prone users, and benefits from diversifying AI signals across users.
Problem
Existing human-AI decision-making theory largely studies individual responses to AI advice, leaving under-explored how shared AI signals reshape sequential judgments and public history.
Method
The paper extends the Bayesian cascade model by treating the AI credibility indicator as a shared public signal, then uses behavioral calibration and simulations to study its effects.
Results
The Gateway creates a preservation–correction trade-off, while calibration finds average users weight AI below private impressions but above several peer judgments and simulations identify weak-AI over-reliance as especially harmful.
Takeaways & Limitations
Diversifying AI signals across users can better keep crowds informative, while expert review is mainly useful for slowly forming cascades in the learning regime.
Takeaways & Limitations
The empirical calibration reuses one experiment, platform paradigm, news corpus, and subject pool, so it is evidence for the mechanism rather than a definitive population-level test.
Abstract
from arXiv · showhide
Social media platforms increasingly use AI-based credibility indicators to help users judge misinformation. Unlike individual human-AI decision-making, these indicators are embedded in information spread: users see both an AI prediction and earlier judgments shaped by the same AI, and their own judgments may then enter the public history. Yet how to analytically characterize this process remains under-explored. We therefore introduce a social-learning lens for this setting by extending the classical Bayesian cascade model with the AI indicator as a shared public signal. The resulting Gateway condition compares the evidence from the AI prediction with users' private impressions. Through this view, we show that AI changes what public history means. Crowd agreement may reflect accumulated independent human evidence, or repeated dependence on the same AI prediction. This creates a preservation-correction trade-off: stronger reliance on AI can preserve correct predictions, but can also lock in incorrect ones by blocking corrective private impressions. We calibrate the model using human-subject data on news veracity judgments. Although the AI outperforms human users, the average user weights it below her own impression but above several peer judgments, while individual users vary from discounting the AI to relying on it enough to cascade. Simulations show that over-reliance on a weak AI is especially harmful, and that diversifying AI signals across users can better keep the crowd informative. We conclude with implications for understanding human-AI interaction in information spread and designing misinformation interventions.
1 Introduction
The paper reframes AI credibility indicators as shared signals embedded in sequential human judgments. Its Gateway analysis shows that AI reliance can preserve correct predictions or lock in incorrect ones, with simulations and calibration revealing heterogeneous user responses.
- Motivation: AI indicators reshape sequential judgment because users observe earlier judgments influenced by the same prediction, unlike one-shot human-AI advice.Existing human-AI decision-making theory largely studies individual revision after AI advice, leaving this sequential setting under-explored.
- Approach: The model extends Bayesian social learning by combining each user’s private impression, the shared AI signal, and earlier judgments on a common log-odds scale.This framework formally links individual reliance on AI to crowd-level dependence.
- Gateway mechanism: When AI accuracy α meets or exceeds private-impression accuracy q, the Gateway causes users to follow the AI regardless of their own impressions.The AI prediction then absorbs the crowd, preventing private impressions from entering the public record.
- Implications: Shared-signal absorption creates a preservation–correction trade-off: it filters noisy impressions when AI is correct but blocks correction when AI is wrong.Stronger AI improves the favorable side but also makes incorrect predictions harder to overturn.
- Empirical calibration: The average user weights AI below her own impression but above several peer judgments, while users range from AI discounting to immediate cascading.The calibration uses sequential human-subject news-veracity judgments and identifies a cascade-prone minority.
- Simulations: Over-reliance on a weak AI is more damaging than under-reliance on a strong one, while diversified AI signals more reliably preserve informative crowds.Expert review helps mainly for slowly forming cascades when users remain in the learning regime.
2 Related Work
Prior work explains individual reliance on AI and social learning separately, but does not fully characterize how AI-shaped judgments propagate through public histories. The paper connects these traditions by treating credibility indicators as platform-deployed shared signals.
- Misinformation spread: Social platforms spread misinformation through sequential user judgments, while individual truth discernment is only modestly better than chance.Users encounter stories, assess veracity, and react through visible platform interactions.
- Human-AI decision making: Human-AI interaction research studies when people accept, discount, or over-rely on algorithmic advice in individual decisions.Credibility indicators and warning labels can move beliefs and sharing decisions about misinformation.
- Research gap: In information spread, AI reliance enters public history through visible judgments and changes what later users learn from the crowd.This sequential dimension remains under-explored by individual human-AI decision-making theory.
- Theoretical bridge: Bayesian log-odds provide a common currency for private impressions, AI predictions, and earlier judgments in social-learning models.The framework draws on established information-cascade machinery rather than introducing a new equilibrium concept.
3 Model and Dynamics
The model adds one common AI prediction to a classical Bayesian cascade, making it an initial public condition while human judgments update the public log-odds. Learning persists only below private-impression boundaries; beyond them, judgments become absorbing cascades.
- Model setup: Each user observes a private impression, one common AI signal, and earlier public judgments before choosing a binary veracity judgment.The AI signal is computed once per story and shown unchanged to every user.
- Bayesian updating: The posterior log-odds combine public evidence with a private contribution, where signal accuracy ρ contributes magnitude L_ρ = log(ρ/(1−ρ)).Private impressions and the AI prediction contribute ±L_q and ±L_α, respectively.
- Public dynamics: The shared AI signal enters public log-odds once and is carried forward, changing the initial level without adding a transition term.Conditional on veracity and the AI signal, the process remains Markovian in the current public log-odds.
- Equilibrium structure: The public-log-odds process is a Markov chain with a transient learning region and two absorbing cascade regions.Once either boundary is reached, the process remains absorbed.
- Learning and cascades: While |P_i| < L_q, the user’s judgment reveals her private impression; at P_i ≥ L_q or P_i ≤ −L_q, the corresponding cascade ignores it.Up-cascades always produce judgment 1, whereas down-cascades always produce judgment 0.
- Gateway: The Gateway condition L_α ≥ L_q determines whether the sequence begins with an immediate cascade rather than learning.In the strong-AI case, τ = 1 and the final cascade direction follows the AI signal.
4 Main Results
The Gateway determines whether a shared AI signal lets private impressions enter the public record. Above it, the AI immediately directs the cascade; below it, early private evidence can still shape outcomes, while visible consensus may overstate independent information.
- The Gateway: The Gateway is reached when AI accuracy α is at least private-impression accuracy q, equivalently when one AI signal weighs at least one private impression in log-odds.The relative weight is k*=L_α/L_q, and the learning region occurs only when α<q.
- Above the Gateway: When α≥q, the cascade forms at the first user and every judgment matches the AI prediction, whether or not that prediction is correct.No private impression is revealed through the public history, so correctness reduces to whether the AI prediction matches the truth.
- Below the Gateway: Below the Gateway, early private impressions can still enter the public history, biasing cascade timing and direction without letting the AI determine either alone.At α=q, this learning phase vanishes; below it, a wrong AI prediction can still be corrected or a correct one overturned.
- External confidence: After a cascade, a Bayes-correct observer’s evidence remains bounded, whereas a naïve observer’s excess log-odds grows linearly with N at slope L_q.The naïve observer treats repeated judgments as fresh private signals; above the Gateway, only the shared AI signal is independently informative.
- Preservation–correction trade-off: A shared AI signal creates a preservation–correction trade-off: it can preserve correct predictions, but stronger reliance can block corrective private evidence and lock in errors.For a single shared indicator, raising AI accuracy trades preservation against correction rather than improving both.
- Structural levers: Diversifying AI signals across users raises the information available from AI predictions with N, while cascades still eventually cap aggregation at an N-independent level.The paper presents diversification as a structural lever for escaping the single-shared-indicator frontier.
5 Empirical Calibration of the Theoretical Analysis
The empirical calibration tests whether users behave like calibrated Bayesians and estimates how they weight AI, private, and public signals. Users depart from the strict benchmark: the average user remains in the learning regime, but a heterogeneous minority is cascade-prone.
- Data and design: The dataset contains 480 news-veracity sequences, with 11 participants per sequence making judgments after viewing private impressions, public history, and treatment-specific AI information.Treatments were Control, AI-before, and AI-after; empirical AI accuracy was 0.750 and private-impression accuracy was approximately 0.515.
- Q1: Benchmark: The calibrated-Bayesian benchmark predicts immediate AI-following cascades, but observed users do not always follow the AI and full-room agreement is rare.The realized information plateau also remains below the strong-AI ceiling of I(c;θ)=0.131.
- Behavioral model: The behavioral model estimates separate log-odds weights for private impressions, AI indications, and net peer judgments, using ratios because logit temperature is not identified.The AI-private ratio is identified from AI-after, while AI-before compares AI and public-history weights.
- Behavioral Gateway: An estimated AI-private ratio at least 1 places a user on the behavioral cascading side of the Gateway; otherwise the chain begins in the learning region.This threshold compares perceived signal weights, not the signals’ true accuracies.
- Q2: Pooled calibration: The average user assigns the AI-private ratio 0.49, weighting one AI prediction about half as much as one private impression but roughly 3.5 times as much as one peer judgment.The behavioral estimate is approximately 3–6 peer judgments, and the average user therefore remains in the interior learning regime.
- Q3: Heterogeneity: Approximately 20% of subjects are assigned to a component weighting the AI at least as much as their own impression, while about 46% strongly favor their own impression.The remaining approximately 34% are not cleanly classified into either regime.
6 Simulations Extended from the Calibration
The simulations separate perceived trust from AI accuracy: crossing the behavioral Gateway determines whether cascades follow the AI, while true accuracy determines whether that direction is correct. Diversification is more robust than expert review across the tested regimes and behavioral weights.
- 6.1 Miscalibrated trust: the perceived Gateway: The behavioral Gateway is determined by perceived AI accuracy, so over-trusting a weak AI can push consensus toward near-chance accuracy.Above the perceived Gateway, the cascade follows the AI prediction; below it, private impressions continue entering public history.
- 6.1 Miscalibrated trust: the perceived Gateway: The experimental population occupies different regimes under objective calibration and behavioral weighting, with the gap between placements forming the main empirical finding.Its objective placement is in the strong-AI regime, whereas its behaviorally inferred placement falls in the interior regime.
- 6.1 Miscalibrated trust: the perceived Gateway: Figure 6 shows that P(D∞=c) reaches 1 when perceived accuracy crosses q, whereas P(D∞=θ) follows true AI accuracy.For αtrue=0.55, over-trusting the AI reduces consensus accuracy from approximately 0.77 to approximately 0.55.
- 6.2 Exploratory Analysis: simulations of AI diversification and expert review interventions: Expert review improves accuracy mainly below the behavioral Gateway, because immediate cascades prevent slow-cascade warnings from triggering review.In the tested settings, diversification remains more reliable across AI strengths and behavioral weightings.
- 6.2 Exploratory Analysis: simulations of AI diversification and expert review interventions: Diversification never performs worse than no AI and approaches perfect consensus accuracy as AI weighting increases, unlike the shared-AI baseline for weak indicators.The shared baseline can fall below the no-AI line when a weak indicator is over-trusted.
7 Discussion
The discussion frames the AI indicator as a shared public signal that changes how sequential judgments and crowd agreement should be interpreted. It identifies preservation–correction trade-offs, recommends information-structure interventions, and bounds the conclusions through model and empirical limitations.
- 7 Discussion: Treating the AI prediction as an exogenous common signal relocates where the public process starts and can create an AI-shaped information ceiling.The same AI output enters every posterior as a constant term rather than arising from the population.
- 7 Discussion: The Gateway depends on the accuracy users believe, and system evaluation therefore requires separate preservation and correction measures.Higher AI accuracy can improve preservation while reducing correction, making the regime points mutually non-dominating.
- 7 Discussion: A large AI-aligned majority may contain little independent human evidence, so crowd agreement should not be treated as straightforward validation.Repeated responses to one AI signal can resemble independent human judgments in the public history.
- 7 Discussion: Provenance disclosure and diversified AI signals address common dependence, while cascade-triggered expert review is mainly useful below the behavioral Gateway.Diversification is the most reliable intervention in the simulations, including for over-trusting crowds, although immediate cascades limit its benefit.
- 7 Discussion: The model assumes binary veracity, conditionally independent private impressions, and Bayesian or behaviorally quasi-rational users, simplifying modern misinformation settings.The empirical calibration relaxes strict Bayesian behavior but retains conditional independence and AI exogeneity as constraints.
- 7 Discussion: The empirical calibration reuses one experiment, platform paradigm, news corpus, and subject pool, so it is evidence for the mechanism rather than a definitive population-level test.Repeated subjects across rooms also require accounting for dependence across sequences.
- 7 Discussion: The sequential-chain lens omits network structure, strategic responses, optimal disclosure, and several AI-specific properties such as calibration, explanations, provenance, and perceived authority.Future work is proposed to extend the model and test interventions in larger-scale or field experiments.
8 Conclusion
The paper concludes that a shared AI signal reshapes Bayesian cascades: agreement can reflect repeated AI dependence rather than independent human evidence, creating a preservation–correction trade-off. Boundary tie-breaking changes limited timing details while leaving the main information and confidence conclusions intact.
- Conclusion: The shared AI signal makes crowd agreement potentially reflect repeated dependence on one prediction rather than accumulated independent human evidence.The information in the public record is limited to the shared signal and finite pre-cascade private impressions.
- Conclusion: The paper’s broader conclusion is that AI reliance can preserve correct predictions but also block corrective private impressions when the prediction is wrong.This trade-off motivates interventions such as diversifying AI signals across users.
- Boundary conventions: The conventional tie-breaking rule matters only at the measure-zero edges α∈{1/2,q}; elsewhere cascade entry is strict and tie-breaking is irrelevant.For other α, absorption occurs with |P|>L_q.
- Boundary conventions: At α=1/2, public and random tie-breaking conventions differ in cascade-formation time because random ties can reveal additional private impressions.Under the public rule τ=2, whereas under the random rule τ takes values in {2,3,…} with finite expected additional revelations.
- Conclusion: Both tie-breaking conventions preserve an N-independent information ceiling and the naive observer’s excess confidence slope L_q.The public rule reveals exactly one private impression; the random rule reveals only a few extra impressions in expectation.
B Continuous-Valued AI Indicators
The continuous-indicator extension replaces the binary AI likelihood contribution with a real-valued likelihood ratio while retaining the same cascade thresholds. First-user cascading becomes probabilistic according to the indicator’s resolving power, and information remains bounded by the AI signal plus the finite pre-cascade record.
- Continuous-valued indicators: A continuous AI indicator uses a real-valued output c and replaces the binary log-likelihood term λ(c;α) with a log-likelihood ratio under a monotone likelihood-ratio condition.The conditional density ratio f_1(c)/f_0(c) is strictly increasing in c.
- Cascade formation: With probability G_q=P(|Λ(c)|≥L_q), the continuous indicator causes a cascade at the first user.Conditional on |Λ(c)|≥L_q, every judgment follows the sign of Λ(c).
- Comparative statics: Greater separation of the AI’s conditional distributions increases G_q toward 1, whereas more accurate private impressions raise L_q and reduce G_q.First-user cascading is more likely when the AI indicator’s resolving power dominates the private impression’s.
- Information aggregation: The information bound becomes I(J̄_N;θ)≤I(c;θ)+I(σ<τ;θ|c), collapsing to I(c;θ) under almost-sure AI dominance.The continuous extension therefore preserves the discrete information-ceiling logic.
- Asymmetric priors: With an asymmetric prior, the first-user cascade threshold depends on the prior and realized AI indication: aligned priors lower it, while opposed priors raise it.A uniform condition for both indications is L_α≥L_q+|ℓ_0|.
D Additional results and figures for the main theorems
Additional results characterize cascade timing, direction, discontinuities, and prior effects. Correct AI predictions shorten cascades, incorrect predictions lengthen them, and the direction probability jumps at the Gateway threshold α=q.
- Cascade timing: On the interior α∈(1/2,q), cascade timing depends on q and whether the AI prediction is correct, not on α itself.The chain dynamics within this interval have no α-dependence.
- Cascade timing: P(τ=2|c=θ)=q, while P(τ=2|c≠θ)=1−q.These first-step probabilities distinguish correct from incorrect AI predictions.
- Asymmetric priors: With an asymmetric prior, the threshold varies with the realized indication: a prior aligned with c lowers α*, while an opposed prior raises it.The threshold is determined by L_α*=L_q−ℓ_0(2c−1).
- Cascade timing: E[τ|c≠θ]−E[τ|c=θ]=(2q−1)/(1−q+q^2)>0, so incorrect AI predictions produce longer expected cascade formation.The asymmetry is governed by q alone.
- Discontinuity: The probability of a first-user cascade is identically 1 for α≥q and jumps to 1 at α=q rather than approaching it continuously from below.The interior limit remains strictly below 1.
E Dataset and experimental design
The experiment models each room as an eleven-user sequential cascade with a news item’s fixed veracity, a fixed binary AI prediction, and judgments that enter the visible public history. Balanced treatments vary whether the AI indicator is absent, shown before the private impression, or shown afterward.
- Dataset: 538 subjects contributed about ten judgments each, with no subject spanning more than one treatment.Subjects rotated across rooms, preserving the in-room sequential independent-signal assumption while leaving rooms non-independent across the corpus.
- Dataset: Item-level analysis pools the two AI treatments, with each of the 40 items observed in 8 rooms.The treatment design holds α fixed across treatments.
- Dataset and model: Each room contains N=11 sequential users judging one news item whose veracity is fixed and known to the experimenter.The AI prediction is generated once per item and held fixed within the room.
- Experimental treatments: The three balanced treatments are Control with no AI, AI-before shown before the subject’s initial impression, and AI-after shown alongside public history afterward.Each treatment contains 160 rooms and uses the same 40 news items.
- Experimental treatments: In the AI-after treatment, the recorded initial impression separately measures the private signal before the published judgment incorporates the AI and public history.This treatment makes the private impression observable independently of the final judgment.
F.1 Strict-benchmark diagnostics
The behavioral calibration tests whether users’ sequential judgments match the strict Bayesian benchmark and estimates how users weight public history, private impressions, and the AI signal. Results show aggregate AI influence with substantial heterogeneity, while item-level clustering remains outside the per-user model.
- Strict-benchmark diagnostics: 0.644 and 0.650 are the empirical first-user cascade rates in the AI-before and AI-after treatments, versus the strict model’s prediction of 1 at every position.Across all eleven positions, the per-position rate remains between 0.56 and 0.71.
- Strict-benchmark diagnostics: 0.04–0.06 nats is the AI-after information plateau, below the strong-AI tight value of 0.131 nats at b_α = 0.75.Control remains at ≤0.045 nats and AI-before at ≤0.034.
- Strict-benchmark diagnostics: ≈0.03 per user is the OLS slope of naive external log-odds growth across all three groups, near b̄L_q = 0.061.The shortfall is attributed to imperfect cascade alignment.
- Behavioral weights: 5.9 and 3.5 are the temperature-free AI-to-public per-signal ratios in AI-before and AI-after, respectively.These ratios imply that one AI prediction contributes roughly as much as three to six public-history judgments.
- Behavioral weights: 0.49 is the AI-after AI-to-private threshold ratio, while 3.35 is its AI-to-public ratio; the control public-to-private ratio is 0.12.The AI-after AI-to-private confidence interval is [0.33, 0.67], below 1 in every bootstrap resample.
- Model scope: Per-user parameters cannot explain the observed item-level clustering, where agreement with the AI signal ranges from 0.12 to 1.00 across 40 items.Between-item over-dispersion is approximately 3.3 times the within-item binomial benchmark.
I Delayed exposure: when to attach the indicator
Delayed attachment changes the relationship between the crowd’s eventual direction and the majority record. A strong AI can redirect the cascade from any attachment position, yet the recorded majority follows it only when exposure begins before halfway through the sequence.
- Delayed exposure: k ≤ N/2 is the condition under which the delayed strong AI wins the majority sign(ȲJ_N).For later attachment, the recorded majority remains controlled by the pre-AI herd even though the final cascade direction follows c.
- Welfare consequences: Delaying a strong AI leaves cascade welfare at the above-Gateway point (1, 0), while record welfare slides from (1, 0) toward the classical (q, q).The two welfare objects diverge because delay separates D_∞ from sign(ȲJ_N).
- Delayed exposure: A strong AI with α = 0.85 > q sets D_∞ = c at every attachment position k, whereas a weak AI with α = 0.62 < q does neither.The result concerns the final cascade direction, not necessarily the aggregate majority.
- Welfare consequences: A strong AI introduced late is worse on both record-welfare coordinates than a weak AI shown from the start.The below-Gateway point strictly dominates the delayed strong-AI path when welfare is scored from the aggregate record.
- Mechanism: The shared AI signal shifts the initial public log-odds, while private impressions continue to enter only before the cascade absorbs.This is why attachment timing affects the recorded majority without changing the underlying cascade’s absorbing behavior.
J.7 Proof of Theorem 8 (AI-Mediated Agreement and External Overconfidence)
The proof shows that a shared AI signal bounds the independent information contained in the judgment sequence, even as naive external confidence grows with cascade length. Once a cascade forms, later repeated judgments add no private evidence.
- Information aggregation: I(ȲJ_N; θ) is bounded uniformly in N because the judgment sequence is determined by the finite pre-cascade prefix and the shared AI signal.For α ≥ q, the prefix is empty and equality holds at I(ȲJ_N; θ) = I(c; θ).
- Information aggregation: Only judgments before τ can reveal private impressions; the forced judgment at τ and all later cascade judgments reveal none.Data processing reduces the sequence’s information to variables determined by (c, σ_<τ, τ).
- External overconfidence: N·L_q + O(1) is the expected naive log-odds projection along the consensus direction, because post-cascade judgments are repeatedly counted as fresh signals.The Bayes-correct observer’s contribution remains bounded, producing the external-overconfidence gap.
- External overconfidence: N·L_q − L_α is the external-overconfidence gap when α ≥ q and the AI signal is shown from the first position.In the α = 1 benchmark, the corresponding gap is N·L_q − L_q.