Source-linked AI summary
The Economics of Social Data
Dirk Bergemann, Alessandro Bonatti, Tan Gan
TL;DR
The paper studies how the social nature of individual data affects privacy, data acquisition, information sharing, and surplus. It models data intermediation among consumers, firms, and platforms, showing that social correlations can lower acquisition costs and that anonymization is optimal exactly when it raises social surplus.
Problem
Digital platforms’ low-cost acquisition and selective use of consumer data challenge privacy regulation, while individual data also predict the behavior of other consumers.
Method
The paper analyzes a data intermediary’s collection, transmission, aggregation, and precision choices in product and advertising markets involving consumers and firms.
Results
The intermediary collects anonymous data if and only if transmitting identities reduces total surplus, while strong preference correlation can lower acquisition compensation but still make sharing welfare-reducing.
Takeaways & Limitations
Social data can magnify individual-data value and facilitate large-scale acquisition, but profitable intermediation and social-welfare gains need not coincide.
Takeaways & Limitations
The intermediary cannot commit to withholding information from the producer and chooses its data-outflow policy before consumers’ data are realized.
Abstract
from arXiv · showhide
A data intermediary acquires signals from individual consumers regarding their preferences. The intermediary resells the information in a product market wherein firms and consumers tailor their choices to the demand data. The social dimension of the individual data -- whereby a consumer's data are predictive of others' behavior -- generates a data externality that can reduce the intermediary's cost of acquiring the information. The intermediary optimally preserves the privacy of consumers' identities if and only if doing so increases social surplus. This policy enables the intermediary to capture the total value of the information as the number of consumers becomes large.
1 Introduction
The paper explains how social data create externalities that shape data acquisition, information sharing, privacy, and welfare. It studies how a data intermediary’s market power affects these trade-offs across data and product markets.
- Motivation: Social data are informative about both the contributing consumer and other consumers with similar characteristics or behaviors.This predictive spillover creates a data externality whose sign and magnitude depend on the data structure and use of information.
- Framework: The framework models consumers, firms, and data intermediaries interacting in linked data and product markets.The intermediary acquires individual demand information, shares some information with consumers, and sells some to the producer.
- Value of Social Data: Collecting many signals can filter idiosyncratic errors or demand shocks and reveal common fundamentals or noise components.Which component is learned depends on whether fundamentals or errors are more correlated across consumers.
- Welfare and Market Power: The data externality can make information acquisition cheaper for the intermediary than its social value, creating a wedge between profitable and socially efficient data use.Profitable intermediation can harm consumers when correlated preferences and precise signals make acquisition inexpensive but producer information reduces welfare.
- Privacy and Anonymization: The intermediary collects anonymous data if and only if revealing identities reduces total surplus when consumers are homogeneous ex ante.Anonymization prevents identity-linked pricing while preserving the socially preferred information policy in this setting.
- Privacy and Anonymization: As the number of consumers increases under anonymized intermediation, each consumer contributes less to aggregate information, widening the gap between social data value and individual data payments.The resulting decline in individual payments can also reduce total payments to consumers.
- Extensions: With heterogeneous consumers, the intermediary aggregates at least to the coarsest homogeneous group, while product personalization can preserve recommendations without personalized prices.The model also allows further aggregation when consumer numbers are small.
2 Model
The model represents social data as correlated consumer signals traded by a monopolist intermediary between consumers and a producer. Its timing separates contracting, signal realization, information transmission, and product-market pricing.
- Environment: The environment contains many consumers, one data intermediary, and one producer interacting through data and product markets.Consumers choose quantities, while the producer sets unit prices and the intermediary controls information flows.
- Data Environment: Each consumer’s signal combines willingness to pay with noise, while fundamentals and errors may be correlated across consumers.The information structure permits arbitrary symmetric distributions and correlation patterns subject to the model’s independence restrictions.
- Data Environment: When fundamentals are common and errors independent, averaging signals identifies common willingness to pay as the consumer population grows.When fundamentals are independent and errors common, averaging instead identifies the common error and helps recover individual willingness to pay.
- Data Market: The monopolist intermediary chooses how to collect and share signals and therefore faces both information-design and information-pricing problems.Contracts specify data policies and fees for consumers and the producer before demand shocks are realized.
- Data Policies: Anonymized collection prevents the producer from matching signals to consumers and is equivalent here to aggregate information about average willingness to pay.Sharing information with consumers also affects their demand and therefore the producer’s willingness to pay for data.
- Equilibrium and Timing: The sequential game has contracting first, signal realization and transmission second, and producer pricing followed by consumer purchase decisions.The analysis characterizes Perfect Bayesian Equilibria over policies, prices, quantities, and participation decisions.
- Model Features: The model assumes the intermediary cannot commit to withholding information from the producer after consumers are enlisted.This captures limited ability to write advertising contracts contingent on later platform activity.
3 Value of Social Data
Data outflow creates value through consumers learning about their preferences and producers improving pricing, but producer information can reduce consumer and social surplus. The sign of the data externality depends on how consumers’ signals reveal information about one another.
- Data outflow policy: The intermediary’s data outflow policy determines both the producer’s information and the information available to consumers.The intermediary chooses the fee and information flow after observing the data inflow.
- Welfare effects: Consumers and social surplus increase with information learned by consumers but decrease with information learned by the producer.Consumer learning makes demand more responsive to willingness to pay, whereas producer learning enables price responses that reduce quantity responsiveness.
- Welfare effects: If consumers cannot learn from one another, any data sharing reduces consumer and social surplus; if individual signals are uninformative, any sharing improves both.These cases show that the welfare effect depends on the informativeness of initial signals and cross-consumer learning.
- Welfare effects: Social surplus is maximized by sharing all collected signals with consumers and none with the producer.The first-best allocation follows from the opposing welfare effects of consumer and producer information.
- Data externality: A market-power intermediary implements complete data outflow, sending the entire realized data inflow to the producer and all consumers.Because intermediary profits rise with producer information while each consumer receives at least as much information as the producer, complete outflow is optimal.
- Data externality: The data externality is negative when other consumers’ signals help predict a consumer’s willingness to pay but provide little additional learning for that consumer.It is positive when other consumers’ signals help filter common noise without improving the producer’s prediction of that consumer’s willingness to pay.
4 Optimal Data Intermediation
The intermediary’s profit equals the social-surplus effect of data sharing net of cross-consumer data externalities. Profitability therefore need not coincide with welfare improvement, and depends on whether other consumers’ signals substitute for each individual signal.
- Profitability condition: The intermediary’s profit equals the effect of data sharing on social surplus net of data externalities across consumers.Negative externalities reduce compensation owed to consumers, while positive externalities increase it.
- Profitability condition: Negative data externalities can make intermediation profitable but welfare reducing, whereas positive externalities can make welfare-enhancing intermediation unprofitable.The intermediary’s private objective differs from the social planner’s because acquisition costs reflect cross-consumer effects.
- Profitability condition: Data intermediation is profitable if and only if signals from other consumers generate at least 1/3 of the variance in willingness to pay explained by the entire data vector.Other signals act as substitutes for an individual signal, lowering acquisition costs; with independent fundamentals, this condition fails and intermediation is not profitable.
Corollary 2 (Common Preference)
With common preferences and independent errors, intermediation can be profitable while reducing social surplus; with independent fundamentals and common errors, sharing can improve surplus without generating intermediary profits. Anonymization reduces acquisition costs by limiting individual price discrimination.
- Common Preference: When fundamentals are perfectly correlated and errors are independent, data intermediation is always profitable for large N but can reduce social surplus when σ is sufficiently small.The negative welfare effect arises despite the intermediary’s profitable access to highly correlated information.
- Common Preference: When fundamentals are independent and errors are perfectly correlated, data sharing increases social surplus for sufficiently large σ but intermediation is never profitable.Independent fundamentals prevent the intermediary from acquiring data at reduced cost through cross-consumer substitution.
- Anonymization: Anonymized data induces a uniform price across participating consumers while retaining information about aggregate demand.It limits the producer’s ability to extract surplus from individual consumers and supports third-degree price discrimination across total-demand realizations.
- Anonymization: Anonymization profitably reduces the intermediary’s data-acquisition costs even though market-demand data are less valuable to the producer than individual-demand data.The lower value of anonymized information is offset by lower costs of obtaining fine-grained consumer data.
Proposition 3 (Optimality of Data Anonymization)
The intermediary strictly prefers anonymized consumer data because removing identities lowers acquisition costs without reducing consumers’ own learning, while preserving useful aggregate demand information. This trade-off makes anonymization central to profitability, though its optimality depends on the product-market setting.
- Proposition 3 (Optimality of Data Anonymization): The intermediary does not elicit consumer identities, so producers use variable market-demand prices rather than personalized prices.This conclusion holds independently of the distributions of the fundamental and noise components within the model’s policies.
- Proposition 3 (Optimality of Data Anonymization): The anonymization result is bounded: heterogeneous consumers and alternative product-market specifications can make aggregation short of complete privacy optimal.For heterogeneous responsiveness to advertising, coarser-than-complete anonymization may be optimal.
- Proposition 3 (Optimality of Data Anonymization): Anonymization leaves consumers’ learning and the data-externality term unchanged because posterior beliefs do not depend on other consumers’ identities.The producer’s inference is invariant to permutations of other consumers’ signals under the model’s symmetry assumptions.
- Proposition 3 (Optimality of Data Anonymization): Anonymization increases the intermediary’s profits by reducing information-acquisition costs relative to lost producer revenue.The reduction in producer information does not reduce consumers’ own learning, so total surplus terms and intermediary profits increase.
- Proposition 3 (Optimality of Data Anonymization): With many consumers, anonymization enables explosive profitability because each consumer’s marginal contribution and required compensation decline.Under complete sharing, the intermediary’s profits are not amplified in the same way because identity-revealing payments remain positive.
4.2 Large Markets
The large-market analysis studies how increasing participation changes the social efficiency and price of data. Its additive structure separates common and idiosyncratic components while holding pairwise correlations fixed as the market grows.
- 4.2 Large Markets: Large markets can make data intermediation increasingly profitable because product-market opportunities and demand information both expand with the number of consumers.Potential revenue can grow without bound while each consumer contributes only a small marginal amount of information.
- 4.2 Large Markets: The model represents each consumer’s willingness to pay as the sum of common and idiosyncratic components.The common component is shared across consumers, whereas the idiosyncratic component is consumer-specific.
- 4.2 Large Markets: The comparative statics hold pairwise correlations in fundamentals and noise constant as the number of consumers changes.The model denotes these correlations by α for willingness to pay and β for error terms.
- 4.2 Large Markets: The analysis first establishes when complete data sharing is profitable for large markets, then separates intermediary revenue from total acquisition cost.This decomposition supports the paper’s large-market results on profitability and compensation.
Proposition 4 (Profitable Intermediation of Anonymized Data)
With any positive correlation in consumers’ willingness to pay, anonymized data intermediation becomes profitable once the market is sufficiently large. As participation grows, individual compensation vanishes and total compensation stays bounded while intermediary revenue and profit grow linearly.
- Proposition 4 (Profitable Intermediation of Anonymized Data): For any α > 0, anonymized data sharing is profitable when N exceeds a sufficiently large threshold N* .Under the optimal policy, correlated consumers’ anonymized signals become sufficiently close substitutes as N grows.
- Proposition 4 (Profitable Intermediation of Anonymized Data): The large-market profitability result assumes independent error terms, although the authors expect similar results under more general correlated errors.The stated result also uses an additive data structure and a sample-average argument to bound learning from N − 1 signals.
- Proposition 4 (Profitable Intermediation of Anonymized Data): As N →∞, each consumer’s compensation m_i* converges to zero, total consumer compensation remains bounded, and intermediary revenue and profit grow linearly in N.The bounded total compensation follows from each additional consumer’s rapidly decreasing marginal information value.
- Proposition 4 (Profitable Intermediation of Anonymized Data): Anonymization allows total acquisition costs to converge to a constant while intermediary revenue grows linearly with the number of consumers.Consequently, per-capita intermediary profit converges to the per-capita profit attainable when anonymized data are freely available.
Proposition 6 (Asymptotics with Complete Sharing)
Complete identity-revealing data sharing does not produce the same increasing returns to scale as anonymized sharing. When consumers’ idiosyncratic fundamentals vary, individual payments remain positive asymptotically, causing total payments to grow with market size.
- Proposition 6 (Asymptotics with Complete Sharing): Under complete identity-revealing sharing, asymptotic individual compensation remains bounded away from zero when var[θ_i] > 0.The lower bound is strictly positive under the additive data structure.
- Proposition 6 (Asymptotics with Complete Sharing): Complete data sharing makes total consumer payments grow linearly in N, preventing per-capita profits from reaching the full value of information.Anonymization is therefore critical for increasing returns to scale in data intermediation.
- Proposition 6 (Asymptotics with Complete Sharing): The contrast with anonymized sharing is illustrated by an example where acquiring a larger anonymized dataset costs less than acquiring a smaller one, unlike complete data.The example uses normally distributed fundamentals and errors.
4.3 Unique Implementation
The paper shows that a divide-and-conquer contracting scheme guarantees a unique equilibrium outcome, although it raises total compensation relative to the intermediary’s most preferred equilibrium. As the market grows, this additional cost disappears on a per-capita basis.
- Unique implementation: The divide-and-conquer scheme sequentially conditions each consumer’s compensation on earlier consumers’ acceptance, guaranteeing a unique equilibrium outcome.The first consumer receives compensation equal to her entire surplus loss, while later consumers receive the baseline compensation corresponding to their position in the sequence.
- Cost and profits: The scheme’s data-acquisition cost is strictly higher than in the intermediary’s most preferred equilibrium.Despite this higher cost, the effect of unique implementation on per-capita profits vanishes in the limit.
- Cost and profits: Under divide and conquer, total consumer payments do not converge to a finite constant as the number of consumers grows.Their growth rate remains much smaller than the producer’s willingness to pay for data, so per-capita profits converge to the benchmark level when anonymized data are available.
5 Implications for Consumer Privacy
The paper examines how privacy-preserving data policies change with consumer heterogeneity, pricing instruments, product characteristics, and commitment. It finds that anonymization or aggregation can be privately optimal when it raises social surplus, while richer data structures may support limited group-level segmentation or personalized recommendations without personalized prices.
- Anonymization: The baseline anonymization result relies on homogeneous consumers and unambiguous welfare effects from data sharing.Under these assumptions, anonymized data improve consumer and social surplus relative to complete data intermediation.
- Anonymization: With homogeneous consumers, anonymized data intermediation is more profitable than complete data intermediation if and only if anonymization increases social surplus.This result generalizes beyond the linear-pricing model and identifies social surplus as the intermediary’s criterion for anonymization.
- Scope and limitations: The intermediary’s private objective aligns with social welfare only for the anonymization decision, not for the equilibrium information flow generally.The welfare ranking of group and uniform pricing is ambiguous, and market segmentation can be driven purely by data externalities outside the conditions of Proposition 8.
- Heterogeneous consumers: The intermediary anonymizes all signals within each consumer group while potentially revealing group identities, thereby permitting pricing across groups but not within them.For sufficiently large markets, group-level pricing is more profitable than uniform pricing; with few consumers and weak correlation, pooling all signals can instead reduce sourcing costs.
- Heterogeneous consumers: The value of the marginal consumer remains large as the market grows because larger datasets let the intermediary segment demand more precisely and extract more surplus.The resulting data externality can make acquiring additional data cheaper as the intermediary exploits richer demand structure.
- Recommender systems: With vertical and horizontal preference dimensions, the optimal policy anonymizes the vertical component but preserves the horizontal component for targeted product recommendations.This policy enables recommendations matched to tastes without allowing personalized pricing.
- Commitment: Stronger contracts can implement a socially efficient policy that shares signals among participating consumers but not with the producer.However, this commitment-based equilibrium does not capture the role of large online platforms, and the socially efficient policy need not maximize intermediary profits.
6 Conclusion
The conclusion argues that social data can make precise information cheap to acquire while creating welfare risks that individual control rights alone do not resolve. It emphasizes group- and market-level pricing, collective bargaining, privacy managers, and data-outflow taxation as possible policy responses.
- Main implications: In large markets, near-zero compensation can induce consumers to relinquish precise preference information even when firms later use it to extract surplus.The paper therefore concludes that giving consumers control rights over their data is insufficient to ensure efficient information use.
- Main implications: The privacy consequences of data aggregation extend beyond individualized prices to group-level and dynamic market-level pricing.The paper points to potentially significant welfare effects from prices that respond in real time to demand across groups, locations, or periods.
- Policy responses: The paper identifies consumer groups or unions, privacy managers, and taxes on targeted advertising as possible responses to data externalities.Taxing data outflow would limit both efficient and inefficient intermediation while affecting the intermediary’s data-policy choice.
- Model scope: The model treats the intermediary as collecting and redistributing data without directly mediating consumer-producer interactions, unlike platforms that auction access to consumers.This distinction separates the paper’s intermediary from product and social data platforms that provide services while selling information to third parties.
7 Appendix A
The appendix formalizes the signaling environment in which consumers and the producer receive different data outflows and choose demands and prices. Its proofs establish how anonymization affects posteriors, producer profits, consumer surplus, and the profitability of intermediation under different correlation structures.
- Equilibrium setup: The intermediary chooses an information inflow from consumers and an outflow to consumers and the producer before equilibrium pricing and demand are determined.The intermediary’s outflow policy maximizes the producer’s ex ante expected payoff within the signaling equilibrium.
- Equilibrium setup: Consumers form demand using their private signals, intermediary data, reported data, and observed prices, while the producer chooses prices from the information he receives.The appendix also compares on-path and off-path prices under alternative inflow policies.
- Correlation and welfare: The appendix shows that collecting additional signals can raise the intermediary’s value of information when fundamentals or noise are correlated across consumers.When fundamentals are independent across consumers, intermediation is always unprofitable; when initial signals are sufficiently precise, data sharing can hurt consumers.
- Correlation and welfare: Consumer surplus rises when signals are sufficiently noisy because additional information can improve consumers’ knowledge of their own demand.The appendix states this condition as var[E[wi|Si]] being close to zero.
- Anonymization proof: Anonymization leaves each consumer’s posterior about her willingness to pay unchanged relative to non-anonymized data for every realization of the underlying variables.This equivalence holds both on and off the equilibrium path.
- Anonymization proof: Anonymization is more profitable than complete sharing, and strictly more profitable whenever it makes estimation less precise.The appendix derives this comparison from the producer’s equilibrium payoff under the alternative information policies.
- Product differentiation: Aggregating the vertical preference component while preserving the horizontal component is optimal when the latter increases welfare without changing the relevant producer and consumer payoff terms.The resulting policy supports targeted product characteristics but not personalized prices.
Appendix B
Appendix B characterizes optimal privacy-preserving data inflow policies when the intermediary adds common and idiosyncratic noise to consumers’ signals. The policy uses no idiosyncratic noise, may use aggregate common noise, and can increase profits by reducing acquisition costs while preserving information sold to the producer.
- Optimal noise structure: The intermediary adds no idiosyncratic noise under the optimal data inflow policy.
- Optimal noise structure: The optimal policy adds weakly positive aggregate noise.
- Profitability: The threshold for profitable intermediation decreases with N and is independent of the initial-noise correlation coefficient β.
- Profitability: The intermediary’s profits are strictly positive if and only if the profitability condition in Proposition 11 holds.
- Profitability: Positive profits can be obtained by making correlated noise sufficiently large.
- Economic mechanism: Additional common noise reduces consumer compensation and revenue, while allowing the intermediary to hold producer-facing information constant and lower acquisition costs.Common noise makes each individual signal less valuable at the margin for estimating average willingness to pay than under idiosyncratic noise.
- Economic mechanism: If identity-revealing data sharing is required, supplemental idiosyncratic noise can instead be optimal when consumers’ initial signals are sufficiently precise.