Source-linked AI summary
A Survey on Data Pricing: from Economics to Data Science
Jian Pei
TL;DR
Data pricing is needed to assess the value of digitally distributed information goods across reuse, exchange, and business settings. The survey unifies work from economics and data science, covering pricing principles and models for digital and data products. It synthesizes established approaches and identifies open challenges, including interdisciplinary coordination and computational limits in data valuation.
Problem
Data pricing must value shared and reused information goods, but relevant research is dispersed across multiple disciplines and pricing settings.
Method
The paper presents a comprehensive interdisciplinary survey of data pricing, covering economics, pricing principles, digital products, data products, marketplaces, and privacy.
Results
The survey organizes data-pricing research around fundamental properties including truthfulness, fairness, revenue maximization, arbitrage-freeness, privacy preservation, and computational efficiency.
Takeaways & Limitations
Data pricing remains a fast-growing research area with practical demands and unresolved questions requiring broader research effort.
Takeaways & Limitations
Interdisciplinary communication and dialogue among economics, marketing, electronic commerce, data management, data mining, and machine learning need strengthening.
Abstract
from arXiv · showhide
Data are invaluable. How can we assess the value of data objectively, systematically and quantitatively? Pricing data, or information goods in general, has been studied and practiced in dispersed areas and principles, such as economics, marketing, electronic commerce, data management, data mining and machine learning. In this article, we present a unified, interdisciplinary and comprehensive overview of this important direction. We examine various motivations behind data pricing, understand the economics of data pricing and review the development and evolution of pricing models according to a series of fundamental principles. We discuss both digital products and data products. We also consider a series of challenges and directions for future work.
1 Introduction
The introduction frames data pricing as an interdisciplinary response to the need to value shared, reused digital information and surveys pricing for both digital products and data products.
- Motivation: Data reuse across applications creates economic implications and a need to measure data value through scalable, widely acceptable prices.Data pricing is presented as setting a sale or purchase price for data.
- Motivation: Data pricing is difficult because “price of data” can refer to transmission, digital-product content, or data-product services.Examples include mobile data packages, streamed movies, and weather-forecast subscriptions with different granularities and horizons.
- Scope: The survey focuses on digitally distributed information goods, covering digital products and data products while linking pricing ideas across both categories.Digital products include e-books, music, advertisements, and coupons; data products include datasets and services derived from them.
- Related Surveys: The literature spans economics, marketing, electronic commerce, data management, data mining, and machine learning, but interdisciplinary survey coverage has been limited.The article presents itself as an effort to provide a comprehensive picture across these fields.
- Organization: The survey organizes pricing research around economics, fundamental pricing principles, digital-product models, and data-product marketplaces.It reviews versioning, truthfulness, fairness, revenue maximization, arbitrage-freeness, privacy, auctions, and marketplace pricing.
2 Economics of Data Pricing
Information goods differ economically from physical products because search, production, replication, transportation, and tracking costs can be substantially reduced. These properties enable new business models while creating pricing, privacy, copyright, and governance challenges.
- Cost reductions: Information goods can sharply reduce search, production, replication, transportation, and tracking and verification costs relative to physical products.Digital economics examines how economic models adjust when these costs fall.
- Search costs: Low search costs improve discovery, price comparison, long-tail sales, platform matching, and variety, although online consumption can also exhibit echo chambers.The cited examples contrast online and offline prices and consumption patterns.
- Production costs: Reduced production costs and near-zero unit costs through sharing support customization, pay-as-you-go models, query-based consumption, and long-tail products.These reductions arise in materials, semi-finished products, customization, and sharing.
- Replication and consumption: Non-rivalry and zero marginal costs make bundling economically attractive for large collections of information goods.Many products can be sold together without substantial cost increases, helping accommodate diverse preferences.
- Risks and governance: Low replication and distribution costs expand access and innovation but also challenge copyright enforcement and can facilitate privacy breaches, spamming, and online crime.Government-mandated open data is specifically associated with possible data leakages and privacy breaches.
- Transportation costs: Near-zero transport costs can weaken local effects on adoption, although tastes, regulation, copyright, and regional availability may remain location-sensitive.The section notes both flat-world effects and continuing local variation in music and content consumption.
- Tracking and verification costs: Low tracking costs enable personalized markets, behavioral price discrimination, and personalized advertising, while raising serious privacy concerns.Auctions are discussed as mechanisms for pricing advertisements and discovering information-good prices.
2.2 Differences between Digital Products and Data Products
Digital products usually have fixed, independently consumed units, whereas data products are flexibly aggregated, reused, transformed, and resold, producing distinct pricing requirements and sometimes blurred category boundaries.
- Units and consumption: Digital products commonly have well-defined units consumed independently, while data pricing and consumption units vary by customer and use case.Data may be combined, aggregated, and consumed at a granularity different from its basic records.
- Aggregateability: Data products have strong aggregateability that creates business opportunities and technical challenges such as arbitrage-freeness.Customers can aggregate datasets across different dimensions.
- Reuse and resale: Data can be reused for different purposes, processed into new forms, and resold in ways that are harder to detect than whole-product reuse or resale.This reuse flexibility distinguishes data sets from digital products.
- Category boundary: The same information can function as a digital product in one setting and a data product when systematically collected and analyzed in another.Tweets may be consumed individually or sold as a processed dataset for event detection, customer profiling, or recommendation.
2.3 Summary
Information goods differ from physical products through major reductions in search, production, replication, transportation, and tracking and verification costs, with significant consequences for pricing.
- Summary: Information goods reduce search, production, replication, transportation, and tracking and verification costs, shaping their pricing economics.The survey identifies these cost reductions as a central distinction from traditional physical products.
3 Fundamental Principles of Data Pricing
This section introduces versioning as a fundamental framework for designing and pricing information goods, then reviews important cost-model properties for digital and data products.
- Versioning is presented as a fundamental framework for designing information goods.
- The section connects versioning with the pricing of information goods.
- It also reviews important properties in cost models for digital and data products.
3.1 Versioning
Versioning links prices to customer-perceived value by offering differentiated versions of information goods. Versions can vary by features, timeliness, convenience, comprehensiveness, and customer needs, while relational views support data-product versions.
- Very low replication costs make information goods inexpensive but also enable competitors to enter markets easily.
- Versioning links price to value by offering different versions for customers who value different features.
- Versions can differ in timeliness, convenience, and comprehensiveness.
- Free versions let potential customers test information goods and can support awareness, follow-on sales, networks, attention, and competitive advantages.
- The number of versions depends on how many ways the information can be used and how much customer valuations vary.
- Relational views provide a flexible technical means to create data-product versions, but pricing raises challenges including arbitrage, updates, integration, and competing sources.
3.2 Important Desiderata in Data Pricing
Data-pricing models pursue objectives and desiderata including truthfulness, revenue, fairness, arbitrage resistance, privacy preservation, and computational efficiency. The survey describes both formal requirements and practical challenges in achieving them.
- Truthfulness requires buyers to offer prices that maximize their true utility, supporting mechanisms such as auctions.
- Pricing models may optimize cost, profit, sales, or revenue, with revenue maximization receiving special attention.
- Fair allocation requires balance, symmetry, zero payment for no contribution, and additivity across tasks.
- The Shapley value is the unique allocation satisfying the stated fairness requirements.
- Low replication costs create a fairness challenge because sellers can produce more identical units to obtain larger Shapley values and payments.
- Pricing models must address arbitrage, privacy disclosure, and computational scalability, including efficient price computation for many goods and buyers.
3.3 Summary
Versioning links prices of information-goods variants to customer values, while pricing models face requirements for truthfulness, revenue, fairness, arbitrage freedom, privacy preservation, and computational efficiency.
- Versioning links prices of different information-goods versions to the values placed on them by customer groups.
- Important pricing requirements include truthfulness, revenue maximization, fairness, arbitrage-free pricing, privacy preservation, and computational efficiency.
- These requirements create technical challenges for pricing models.
4 Pricing Digital Products
The survey reviews pricing digital products because their pricing ideas can inform data-product pricing, whose boundary with digital products can be blurry.
- The survey briefly reviews digital-product pricing because general pricing ideas can be borrowed and extended to data products.
- Digital products and data products sometimes have an indistinct boundary.
- The section covers revenue streams, bundling and subscription, and auctions as major digital-product pricing topics.
4.1 Streams of Revenues
Digital products generate revenue through money, information/privacy, and time/attention streams, often combining them through pricing structures such as tiers, subscriptions, bundles, or micropayments.
- Money revenue comes from selling content or services, while information/privacy revenue comes from collecting and selling customer information.
- Time/attention revenue comes from selling advertising space, although advertising effects are difficult to measure accurately.
- These revenue streams are interdependent and require tradeoffs when firms design a combined revenue model.
- Common pricing structures include fixed prices, tiers, subscription durations, freemium models, micropayments, and usage-dependent or usage-independent assessment bases.
- Retargeted advertising combines customer behavior data with advertising opportunities to improve advertising effectiveness.
- Digital products generate revenue through money, information/privacy, and time/attention streams.
4.2 Bundling and Subscription Planning
Bundling and subscription pricing exploit the economics of digital products but require solving difficult combinatorial pricing problems, with several results providing revenue guarantees and revealing design limits.
- Bundling combines products or services into one package, a practice enabled by the low replication costs of information goods.
- Selling items separately or as one grand bundle achieves at least a constant fraction of optimal revenue under independent product-value distributions.
- A random single price achieves expected revenue within a logarithmic factor for customers with general valuation functions when supplies are unlimited.
- Subscription pricing charges for customer-platform interaction over time while accounting for heterogeneous usage rates and product values.
- Subscription fees proportional to product-set cardinality achieve 1/4 log 2m+log n of optimal revenue for n customer types and m product types.
- Dynamic pricing for bundles and subscriptions, including promotions and coupons, has received little attention.
4.3 Auctions
The survey presents auction formats, valuation models, and digital-product auction mechanisms, emphasizing challenges from unlimited supply and tradeoffs among truthfulness, competitiveness, and envy-freeness.
- Auction formats: The survey reviews ascending-bid, descending, first-price sealed-bid, and second-price sealed-bid auctions.
- Valuation models: Auction valuation models include private values, common values, and combinations of private and common values.
- Auction principles: Revenue equivalence states that under specified risk-neutral, independent-private-value assumptions, auction mechanisms yield the same expected revenue.
- Auction principles: Empirical studies show that assuming independent valuations can cause inefficient auctions in e-commerce.
- Unlimited supplies: Unlimited digital-product supply creates challenges because the marginal winning bid in a generalized second-price auction can approach zero.
- Unlimited supplies: Random-sampling auctions estimate price thresholds from one bidder sample and apply them to the other sample, producing competitive auctions for unlimited-supply digital goods.
- Unlimited supplies: For multiple unlimited-supply products, optimal assignments and sale prices can be determined through an integer-programming formulation.
- Auction tradeoffs: No auction can be truthful, competitive, and envy-free simultaneously, so randomized mechanisms relax one property probabilistically to retain the others.
4.4 Summary
Pricing digital products centers on revenue generation across money, information/privacy, and time/attention, with bundling, subscriptions, and auctions as recurring pricing mechanisms.
- Digital-product pricing generates revenue through money, information/privacy, and time/attention.These are identified as the three major revenue streams for digital products.
- Low replication costs create opportunities and challenges for bundling and subscription planning.
- Auctions are widely used to price digital products, including basic auction types and their applications.
5 Pricing Data Products
Pricing data products spans market structure and technical problems including arbitrage, revenue maximization, fairness, truthfulness, privacy, and dynamic or federated settings. The surveyed approaches use formal pricing models and approximation methods, while practical constraints include computational complexity, privacy loss, and overlapping queries.
- 5. Pricing Data Products: Data marketplaces involve sellers, buyers, products, purposes, and pricing strategies across multiple market designs.The survey frames data markets by asking what is sold, for what purposes, and who participates.
- 5. Pricing Data Products: Data ownership categories include open, public, and private data, with legal protection constrained by data’s non-rivalrous nature.
- 5. Pricing Data Products: Exact data Shapley values can be computationally prohibitive for large datasets and sophisticated models such as deep neural networks.Monte Carlo and gradient-based methods are developed to estimate them.
- 5. Pricing Data Products: Uniform query pricing is arbitrage-free but achieves only a logarithmic approximation to maximum-revenue arbitrage-free pricing.A greedy non-uniform design retains arbitrage-freeness while iteratively increasing revenue when possible.
- 5. Pricing Data Products: 2 ≤ q(x) ≤ p(x), and q(x) is computable by dynamic programming in O(n^2) time for n interpolated price points.
- 5. Pricing Data Products: Fair and truthful marketplace models can provide polynomial randomized ϵ-approximation algorithms for Shapley-fair payment division with probability 1 − δ.
- 5. Pricing Data Products: Approximation methods reduce utility evaluations for Shapley values, including an O(N(log N) log(log N)) evaluation algorithm under approximate sparsity.
6 Discussion and Open Challenges
The survey identifies open challenges at both macro and micro levels, emphasizing missing end-to-end data supply-chain research and stronger coordination across disciplines. It also calls for practical validation, value assessment, auditing, domain-specific mechanisms, and dynamic pricing.
- Macro-level challenges: Data marketplaces lack systematic investigation of data supply chains and end-to-end solutions connecting production and consumption.The proposed supply-chain view links providers, processors, analysts, consumers, and other roles through value contributions, rewards, and pricing feedback.
- Macro-level challenges: Interdisciplinary communication must improve because data pricing spans economics, marketing, electronic commerce, data management, data mining, and machine learning.
- Micro-level challenges: Most studies propose relative prices, while theoretical models remain weakly connected to absolute pricing practice, marketing effects, and experimental validation.The survey argues that experimental studies should be connected to theoretical investigations because user behavior is difficult to model completely.
- Micro-level challenges: Data marketplaces need systematic value-assessment principles that account for differing valuations among providers, owners, users, and brokers.The survey notes that value assessment and negotiations among marketplace participants remain largely unanalyzed in detail.
- Micro-level challenges: Accounting and auditing procedures remain underdeveloped despite their importance for transparency, marketplace efficiency, transaction verification, and adversary detection.Future work should establish principles, quality guarantees, and operational designs for these procedures.
- Micro-level challenges: Future pricing research must address domain-specific mechanisms and constraints alongside changing demand, supply, and data values.The survey specifically calls for application-aware pricing and mechanisms that capture and monitor dynamic marketplace conditions.