Source-linked AI summary
The Software Supply Chain as a Market for Lemons: A Multivocal Review of Trust Signal Collapse
Ranindya Paramitha, Christian Kästner, Laurie Williams
TL;DR
Cheap dependency-adoption trust signals may no longer reliably distinguish trustworthy from untrustworthy packages, but the breadth of their collapse and ecosystem response was not previously synthesized. The paper conducts a multivocal review of Google Search and Reddit sources across eight signals. It finds collapse driven by adversarial manipulation, indistinguishable legitimate-looking behavior, and non-adversarial AI-driven inflation, while documented responses remain mostly advice rather than action.
Problem
Prior work documented individual signal gaming, but the broader collapse landscape across dependency-adoption signals and the ecosystem’s response remained unexplored.
Method
The study conducts a multivocal review of 252 Google Search sources and 870 Reddit threads covering eight dependency-adoption trust signals.
Results
Cheap trust signals collapse because faking costs less than earning, legitimate and adversarial actors can produce identical outputs, and some degradation occurs without attackers.
Takeaways & Limitations
The ecosystem is drifting toward a market for lemons, where downstream users cannot distinguish good dependencies from bad and proposed cheap-signal substitutes also fail.
Takeaways & Limitations
The corpus underrepresents non-English communities, private organizational grey literature, unindexed sources, and deleted or private Reddit threads.
Abstract
from arXiv · showhide
Practitioners evaluating open-source dependencies rely on cheap trust signals, e.g., stars, download counts, and contributor activity, as substitutes for direct code inspection, assuming those signals reflect genuine trustworthiness. Prior work has documented individual signal gaming, but the landscape of collapses across all dependency-adoption signals, as well as the ecosystem's response, remains unexplored. The goal of this study is to aid software practitioners in understanding the reliability of dependency adoption trust signals, such as download counts and contributor activity, by conducting a multivocal review of 252 Google Search sources and 870 Reddit threads. After coding the corpora, we find that cheap trust signals collapse under three simultaneous forces: adversarial manipulation, gaming techniques indistinguishable from legitimate behavior, and non-adversarial AI-driven inflation. The documented responses are more advice than actual action: 54.6% of Google Search sources contain advice on what practitioners should do, with no actual action taken. Responses proposed substituting one cheap signal for another or aggregating multiple signals, which are now also gameable. Non-adversarial inflation, i.e., degradation caused by the emergence of legitimate AI tooling, lacks documented actual behavior change in either corpus. The gap between known remedy and actual practice points toward a market for lemons: when faking signals costs less than earning them, good and bad dependencies become indistinguishable. Relying on individual practitioners to verify the cheap signals is not sustainable. Costlier signals, such as cryptographic attestation, should be made mandatory so that they become the default for all, not a voluntary choice for the few.
1 INTRODUCTION
The study examines how AI, attackers, and legitimate tooling undermine cheap dependency-adoption trust signals, making trustworthy and untrustworthy packages harder to distinguish. It uses a multivocal review to characterize collapse mechanisms and ecosystem responses across eight signals.
- Motivation: AI can cheaply produce convincing stars, commit histories, documentation, and community endorsements at scale, narrowing the cost gap between honest and malicious signal production.The paper frames this convergence as movement toward a market for lemons.
- Study Scope: The study reviews 252 Google Search sources and 870 Reddit threads covering eight dependency-adoption trust signals.The signals include stars, downloads, documentation, contributor activity, badges, issue response, and community endorsement.
- Study Scope: The analysis codes manipulation techniques, adversarial intent, ecosystem responses, and changes in reliance on signals.Google Search sources receive deductive two-pass coding, while both corpora receive separate LLM-assisted thematic analyses.
- Findings: 54.6% of Google Search sources provide practitioner recommendations without describing practitioners actually taking action.The documented response is therefore predominantly advice rather than enacted behavior.
- Contributions: The study contributes a corpus, codebooks, and role-based recommendations for practitioners, platform operators, tool builders, and policymakers.Its findings cover manipulation mechanisms and ecosystem responses across eight signals.
2 BACKGROUND AND RELATED WORKS
The background explains why dependency reuse requires trust signals and why signal gaming threatens their informational value. The study extends prior work from isolated incidents to a cross-signal synthesis of manipulation and ecosystem response.
- Trust and Signals: Modern applications use hundreds of transitive dependencies, making direct verification impractical and encouraging reliance on observable signals as trust proxies.Signaling theory predicts informative signals when producing them is relatively less costly for high-quality producers.
- Trust and Signals: When low-quality producers can generate signals at comparable cost, the separating equilibrium collapses into pooling, leaving signal outputs unable to distinguish dependency quality.The paper connects this mechanism to Akerlof’s market-for-lemons equilibrium.
- Signal Gaming: Prior studies document isolated manipulation of stars, downloads, documentation, co-maintainer status, and retrieval-oriented user-generated content.These examples address individual signals, mechanisms, or incidents rather than the full landscape.
- Signal Gaming: This study’s novelty is a cross-signal synthesis of manipulation mechanisms, actor intentions, and ecosystem responses recorded across the literature.The synthesis addresses a gap left by prior signal-specific studies.
- Research Gap: Existing empirical studies describe dependency-adoption practices but do not examine responses after practitioners learn that signals can be gamed.Existing attack taxonomies instead emphasize artifact-level compromises and technical safeguards.
- Research Gap: A multivocal literature review is appropriate because practitioner evidence about gaming, pricing, incidents, and platform responses appears largely in grey literature.The method combines formal academic publications with informal practitioner sources.
3 METHODOLOGY
The methodology combines academic and practitioner evidence in a multivocal review, constructs an eight-signal corpus, and applies separate coding schemes to mechanisms and responses. Corpus retrieval and coding are bounded by explicit inclusion, coverage, and reliability criteria.
- Review Design: The review combines formal publications with practitioner sources because evidence about signal gaming and platform responses is concentrated in grey literature.Sources include blog posts, forum threads, vendor reports, public videos, and other informal materials.
- Corpus Construction: The corpus is organized around eight trust signals identified through preliminary software-engineering literature review.Google Search retrieval is followed by inclusion and exclusion filtering, with separate Reddit retrieval for additional grey literature.
- Corpus Construction: 252 Google Search sources and 870 Reddit threads form the final analyzed corpora after collection and filtering.The Reddit set was selected from 95,190 retrieved threads using LLM relevance judgments and a positive post-score requirement.
- Coding: The Google Search corpus uses deductive two-pass coding for manipulation mechanisms and ecosystem responses, while both corpora receive separate thematic analysis.RQ1 captures manipulation or structural degradation; RQ2 captures ecosystem responses.
- Coding: All RQ1 closed-code fields exceeded the 70% agreement threshold after three rounds of prompt calibration.The study reports Krippendorff’s alpha, Cohen’s kappa, and percentage agreement, using percentage agreement as the primary threshold.
4 RQ1 RESULTS: SIGNAL COLLAPSE MECHANISMS
Cheap dependency-adoption trust signals collapse through adversarial manipulation, behavior indistinguishable from legitimate activity, and non-adversarial AI-driven inflation. Sophisticated campaigns can manipulate multiple signals simultaneously, undermining composite checks.
- Mechanism landscape: 66.3% of Google Search sources documented a specific manipulation mechanism, spanning a heterogeneous landscape across trust signals.The corpus identified ten manipulation mechanism categories across eight signals.
- Cost collapse: Faking costs less than earning signals because commodity manipulation, AI generation, and automated activity have reduced the original cost asymmetry.Stars cost $0.03–$0.85 from publicly listed vendors, while documentation can be generated by AI in seconds.
- Ambiguous behavior: 38.3% of documented mechanisms were structurally indistinguishable from legitimate behavior at the signal level.Examples include soliciting genuine community stars, optimizing content for AI retrieval, and using AI to generate documentation.
- Non-adversarial inflation: 14.4% of Google Search sources documented non-adversarial structural inflation caused by legitimate AI tooling, without a deceptive actor.AI coding agents generate 275 million GitHub commits per week, while commits rose 180% and releases rose 30%.
- Detection limits: 52.7% of documented manipulation in the Google Search corpus would remain outside meaningful detection even for a perfect attacker detector.Honest effort, legitimate tooling, and deliberate manipulation can produce the same observable signal properties.
- Cross-signal attacks: Sophisticated attacks reconstruct multiple trust signals rather than manipulating only one, rendering composite signal checks ineffective.Documented campaigns combined fabricated documentation, spoofed histories, credential compromise, package publication, and sustained community activity.
5 RQ2 RESULTS: RESPONSES TO SIGNAL COLLAPSES
The documented response to trust-signal collapse is primarily advice rather than enacted behavior, with practitioners typically downweighting signals and adding costly verification. Proposed substitutes and signal aggregation remain vulnerable to gaming, while cryptographic alternatives and responses to legitimate AI-driven inflation remain uncommon.
- Documented response patterns: 54.6% of Google Search sources with documented responses provide advice or concerns without any actual enacted action.Only 1 in 6 documents an action that was actually carried out.
- Documented response patterns: 55% of sources with documented trust direction record signal downweighting, compared with 23.3% describing replacement and 1.2% abandonment.Practitioners generally continue consulting unreliable signals while trusting them less.
- Verification costs: Added verification can make dependency evaluation more costly than writing code, including high-cost internally audited mirrors and growing human-review bottlenecks.The documented overhead includes automated scanning, composite evaluation, delayed updates, and manual contributor-history inspection.
- Direct verification: Minority sources recommend bypassing signals through source inspection, maintainer-history audits, or sandboxing, but this remains a minority practice in both corpora.These approaches add direct or practical verification rather than relying solely on observable adoption signals.
- Self-defeating remedies: Aggregating stars, documentation, activity, and community endorsement offers only short-term protection because coordinated campaigns can reproduce the full signal envelope simultaneously.The aggregation strategy assumes that manufacturing multiple signals is prohibitively costly, an assumption refuted by documented coordinated campaigns.
- Self-defeating remedies: 49 of 163 sources propose alternative signals, and 65.3% of those alternatives are cheap, observable metrics such as fork-to-star ratios, issue quality, and commit frequency.These substitutes are documented as failing through the same mechanisms as the signals they replace.
- Costlier alternatives: Only 34.7% of sources propose cryptographic infrastructure, while commit signing is adopted by 10% of GitHub repositories.Examples include commit signing, SBOMs, SLSA attestation, and Sigstore provenance.
- Non-adversarial inflation: Among 24 non-adversarial inflation sources, 33% document no response and only 3 document actual behavior change.Legitimate AI tooling and mirror infrastructure can inflate activity or downloads without deceptive intent, while existing detection tools target adversarial actors.
6 LIMITATIONS
The study’s findings are bounded by retrieval coverage, rapidly changing ecosystem conditions, LLM-based filtering, self-reported sources, and uneven signal coverage. Some signal-specific conclusions are therefore less generalizable than others.
- Corpus coverage: The Google Search corpus excludes sources not indexed by Google and underrepresents non-English communities and private organizational grey literature.Its retrieval window ended on July 19th, 2026.
- Corpus coverage: The Reddit corpus is bounded by automated LLM-based relevance filtering, subreddit coverage, and the absence of non-English, deleted, or private threads.The threads cover 2023 to 2026, and borderline-relevant threads were excluded regardless of content.
- Evidence character: The corpus primarily contains experiential and self-reported accounts, and coding captures recorded mechanisms and responses rather than independently verified ground truth.Sources are also filtered through their own framing and incentives.
- Temporal scope: Findings about tool names, pricing figures, and adoption rates may become outdated because the manipulation landscape and ecosystem responses are evolving rapidly.This limitation particularly affects time-sensitive specifics rather than the entire synthesis.
- Signal coverage: Documentation and platform-awarded-achievement signals have thinner coverage than several other signals, making their specific findings less generalizable.The Google Search corpus includes n=14 documentation sources and n=11 platform-achievement sources.
7 DISCUSSIONS, IMPLICATIONS, AND CONCLUSIONS
The review finds that cheap dependency trust signals are collapsing because mimicry, ambiguous behavior, and coordinated attacks undermine their ability to distinguish trustworthy packages. Documented responses largely remain advice or recycle gameable signals, while structurally stronger alternatives are known but rarely adopted.
- Why the Current Trajectory Is Unsustainable: 55.2% of documented Google Search trust direction involves downweighting signals, while 12.9% reaffirms reliance on gameable signals.These patterns indicate different stages of awareness: informed participants reduce reliance, while others continue using the signals.
- Implications for Platform Operators: PyPI’s removal of download-count displays is the only documented platform retirement of a gameable signal, supporting broader platform-level signal retirement.The review identifies platform action as a structural response rather than continued reliance on voluntary practitioner verification.
- Why Signal Substitution and Aggregation Cannot Break the Cycle: Signal substitution fails because alternatives are also cheap and gameable, while aggregation fails because patient attackers can reproduce multiple signals simultaneously.Both responses therefore retain the structural vulnerability that enables trust-signal collapse.
- Why the Current Trajectory Is Unsustainable: 38.3% of documented mechanisms are indistinguishable from legitimate behavior, limiting detection as a response to trust-signal manipulation.Disposable identities and low consequences further reduce deterrence, allowing exposed actors to reappear under fresh pseudonyms.
- Structurally Sound Alternatives: 34.7% of proposed alternatives are cryptographic attestation, provenance signing, or reproducible builds, but only 10% of repositories adopt them.These alternatives are causally tied to the quality they certify and therefore resist ordinary signal inflation, but limited adoption leaves a two-tier market.
- Implications for Policy Makers: Platform-level enforcement of signed provenance is proposed as the software supply-chain counterpart to mandatory email-authentication baselines.Making costlier signals mandatory would address the adoption gap that voluntary use has not resolved.