Source-linked AI summary

Your browsing behavior for a Big Mac: Economics of Personal Information Online

Juan Pablo Carrascal, Christopher Riederer, Vijay Erramilli, Mauro Cherubini, Rodrigo de Oliveira

arXiv:1112.6098v1cs.HCcs.CYcs.SI

TL;DR

Online services monetize personal information, but evidence is limited on how users value different information types in browsing contexts and perceive that monetization. Using Experience Sampling, browser popups, and a truth-telling auction with 168 participants, the paper measures these valuations and perceptions. Users value offline-identity information more than browsing behavior, value Finance and Social information highly, and accept service improvement more readily than provider monetization.

  • Problem

    Evidence is limited on how users value different types of personal information during web browsing and perceive providers’ economic use of it.

  • Method

    The study combines Experience Sampling through browser popups with a reverse second-price auction to measure context-specific monetary valuations and perceptions among 168 participants.

  • Results

    Users valued offline-identity information at €25 median versus €7 for browsing history, while Finance and Social information received high median bids of €15.5 and €12.

  • Takeaways & Limitations

    Users generally support using personal information to improve services but are overwhelmingly negative about providers monetizing that information.

  • Takeaways & Limitations

    The study categorizes websites monolithically, so services hosting multiple content categories may be assigned only their first Alexa category.

Abstract

from arXiv · show

Most online services (Google, Facebook etc.) operate by providing a service to users for free, and in return they collect and monetize personal information (PI) of the users. This operational model is inherently economic, as the "good" being traded and monetized is PI. This model is coming under increased scrutiny as online services are moving to capture more PI of users, raising serious privacy concerns. However, little is known on how users valuate different types of PI while being online, as well as the perceptions of users with regards to exploitation of their PI by online service providers. In this paper, we study how users valuate different types of PI while being online, while capturing the context by relying on Experience Sampling. We were able to extract the monetary value that 168 participants put on different pieces of PI. We find that users value their PI related to their offline identities more (3 times) than their browsing behavior. Users also value information pertaining to financial transactions and social network interactions more than activities like search and shopping. We also found that while users are overwhelmingly in favor of exchanging their PI in return for improved online services, they are uncomfortable if these same providers monetize their PI.

INTRODUCTION

The introduction frames personal information as the good exchanged in a two-sided online market and asks how users value it during web browsing. It motivates context-sensitive measurement and examines users’ perceptions of providers’ economic use of their information.

  • Online services provide free access while collecting, aggregating, and monetizing users’ personal information, creating privacy concerns.
  • Users need to understand the value they assign to personal information before weighing privacy loss against online-service benefits.
  • Valuations can vary with browsing context, interaction type, and demographics, while prior survey-based studies often fail to capture context.
  • The study uses Experience Sampling and a browser plugin to collect users’ responses about personal-information value while they browse different content and services.
  • The analysis focuses on users’ monetary valuations of personal information, rather than other notions of value such as satisfaction or happiness.
  • RELATED WORK: The related-work review covers monetary valuation of personal information, users’ perceptions of monetization, and gaps between stated privacy attitudes and behavior.

METHODOLOGY

The methodology combines Experience Sampling with browser instrumentation to ask privacy and valuation questions in users’ browsing contexts. Websites are classified into eight categories aligned with online advertising categories.

  • Experience Sampling asks participants to report experiences at specific points throughout the day, here using a browser plugin to observe browsing contexts.
  • The plugin logs visited websites and classifies them into eight categories: EMAIL, ENTERTAINMENT, FINANCE, NEWS, SEARCH, SHOPPING, SOCIAL, and HEALTH.
  • Auction popup: The auction popup identified each auction by sequential number and date and let participants bid or decline participation.
  • The eight categories were chosen to correspond closely to categories used by online advertising networks.
  • When users changed browsing context, the plugin triggered questions about privacy perceptions and personal-information valuation.

Participants

The study recruited Firefox users for a two-month experiment involving baseline monitoring, active context-sensitive popups, and a post-study questionnaire. Participants received fixed and auction-based incentives and could withdraw from the approved study.

  • From 279 recruits, 168 participants installed the Firefox plugin and completed the study; participants ranged from 18 to 58 years old.
  • The study lasted two months, with one inactive baseline week, four weeks of active popups, and a final questionnaire.
  • Baseline: During the baseline week, the plugin silently recorded browsing behavior to measure normal browsing and assess whether popups altered it.
  • Active study: During the experiment, popups asked about monetization perceptions and auctioned the minimum amount participants would accept for specified personal information.
  • Incentives: Participants received a €10 gift card and could add winnings from auctions, subject to a stated maximum of €3000.
  • Participant protections: The experiment received ethical and legal approval, and participants could disable or remove the plugin or leave without consequences.

Auction game

The auction game uses a reverse second-price mechanism to elicit participants’ monetary valuations of personal information. Participants could submit precise bids, decline participation, and were told whether their information would be used.

  • The reverse second-price auction selects the lowest bidder as winner and pays that participant the second-lowest bid.
  • The mechanism was chosen because truthful bidding is the best strategy, it had prior valuation use, and it was simple to explain.
  • Participants could submit nonnegative bids with cent-level precision or choose not to participate in an auction.
  • Winners were notified that their exact bid-on information would be used, while losers were told it would not be used.

Apparatus

The study apparatus combined a browser plugin with a server to collect browsing activity, categorize visited sites, present auctions and perception questions, and process participant responses.

  • System components: The system used a browser plugin and web server to capture browsing context and exchange configuration and response data.The server received bids and questionnaire responses and stored them in a database.
  • Browsing capture: The plugin recorded URLs, access times, and a browser identifier while omitting events such as file uploads and text highlighting.
  • Site categorization: Websites were assigned to eight categories using a hard-coded list of 1,184 popular Spanish sites and a keyword-similarity fallback for unlisted sites.The categorization relied on Alexa data and the Adnostic folksonomy approach.
  • Study materials: The apparatus used Table 1 to organize the questions asked during the study’s different phases.
  • Participant interface: Two independent pop-ups presented auction questions and privacy-preference questions, with controls for switching each pop-up on or off.The auction interface highlighted the PI being traded.
  • Auction execution: Auctions ran automatically for each category and PI type whenever 20 bids accumulated, with pooled auctions conducted daily.Participants received winner and loser results by email.

Measures

The measures separated privacy knowledge, PI valuation, and attitudes toward monetization, while preserving the browsing context in which responses were collected.

  • Measure design: The study used recruitment and Experience Sampling questions to measure privacy-related knowledge, auction valuations, and perceptions of PI exploitation.
  • Auction measures: Questions a1–a4 measured PI valuations: a1 concerned offline identity, a2 browsing history, and a3–a4 category-specific PI.
  • Information sources: The study distinguished third-party browsing information from category-specific information generally held by the service provider delivering that service.
  • Perception measures: Questions ap1–ap4 assessed awareness and comfort with monetization, exchanging PI for enhanced services, and personalized advertisements.
  • Context: The researchers called a1 context independent and a2–a4 context dependent because the latter related to the website or category being visited.
  • Analysis: Nonparametric analysis accommodated ordinal variables, non-normal continuous variables, and missing category observations from natural browsing.

AUCTION AND SURVEY RESULTS

The auction study measured participants’ monetary valuations across PI types and browsing categories, finding distinct values for identity, browsing, and category-specific information.

  • Auction outcomes: The median winning bid was 5 cents of Euro, while the median payout was 45 cents of Euro under the reverse second price auction.Winning bids ranged from 0 to 2.29 Euros, and payouts ranged from 0.01 to 5.69 Euros.
  • Sample and categories: 168 participants browsed naturally across categories, but health pages were visited by only 2%, so comparisons used seven categories.Search and entertainment were each visited by 82%, while email was visited by 64%.
  • Context-independent PI: € 25 was the overall median bid for context-independent offline-identity PI, with no significant category difference (p = .702).This PI included age, gender, address, and bank balance.
  • Browsing behavior: € 75 was the overall median bid for context-dependent browsing clicks, with no significant category difference (p = .569).
  • Category-specific PI: FINANCE had the highest median bid for category-specific PI at 15.5, followed by SOCIAL at 12 and EMAIL at 6, with significant differences across categories (p < .001).SHOPPING was 5, while NEWS, ENTERTAINMENT, and SEARCH were each 2.
  • Information quantity: One and 10 pieces of category-specific PI received similar median bids: 5 for question a3 versus 5 for question a4 (p = .59).The result indicates that the amount of information sold was not a significant bidding factor.
  • Demographic associations: Age was negatively correlated with median bids for Social-a3 and for the combined Social-a3 and Social-a4 measures.The reported correlations were ρ = −.276 and ρ = −.287, respectively.
  • Privacy associations: Concern about online data protection was positively correlated with higher bids for context-independent PI in entertainment, finance, and search.The reported correlations were ρ = .252, .278, and .23, respectively.

Results for RQ2

Participants generally knew that their PI could generate revenue and wanted services improved with PI, yet they were uncomfortable with providers monetizing it.

  • Response handling: The study used participants’ first answers to perception questions to prioritize initial opinions over possible effects of prolonged study exposure.
  • Knowledge of monetization: Participants were aware that websites could use shared PI to generate revenue, with a median rating of 4 across categories.Awareness did not differ significantly by category (p = .107).
  • Comfort with monetization: Participants were uncomfortable with websites extracting revenue from their PI, reporting a median rating of 2 across categories.The category difference was not significant (p = .429).
  • Service improvement: Participants wanted online companies to use their PI to improve web services, with a median rating of 4 across categories.The category difference was not significant (p = .869).
  • Personalized advertising: Participants were indifferent toward personalized advertisements based on PI, with a median rating of 3 across categories.The category difference was not significant (p = .686).

DISCUSSION

Users place higher economic value on offline identity information than on PI generated by online behavior. The authors distinguish this valuation comparison from disclosure strategies and suggest lower valuation of online PI may reflect limited awareness of how browsing data can be profiled and linked to offline identities.

  • Users consistently value offline identity PI, including age, gender, address, and economic status, more highly than PI related to online behavior.
  • The comparison concerns economic value attached to offline versus online-created PI, not users’ disclosure strategies.
  • Lower valuation of browsing information may reflect that users find offline PI more explicit and the consequences of continuous tracking harder to understand.

Users do not distinguish between quantity of PI, but type

Users’ valuations depend more on the type and context of PI than on its quantity. They value finance and social information more highly than search and shopping information, while attitudes toward monetization, service improvement, and personalized advertising diverge.

  • Median bids for a3 and a4 show little or no difference when the type and context of PI remain constant but quantity changes.
  • Users value FINANCE and SOCIAL PI more highly than SEARCH, SHOPPING, and other categories.
  • Age shows a high negative correlation with category-specific PI valuations for SOCIAL, ENTERTAINMENT, and NEWS, especially for bulk information.
  • Users are overwhelmingly negative about monetization, prefer PI use for service improvement, and are indifferent to personalized advertising.
  • Users may feel unfairly treated when providers gain from PI, possibly because they lack awareness of how the ecosystem operates.

IMPLICATIONS FOR DESIGN AND FUTURE RESEARCH

The findings motivate market and design proposals that give users more control over PI transactions and make the economic exchange more explicit. The paper presents valuation differences as inputs for service design and future ecosystem modeling.

  • The study proposes implications for monetizing PI and for designing new online services.
  • An open PI market could let users choose what information to sell and receive monetary compensation, potentially addressing concerns about inadequate compensation.
  • Measured valuations provide an empirical foundation for modeling PI markets and can help providers distinguish economically different information types.
  • A photo-selling market could let users select commercially usable photos, receive compensation after hosting costs, and sell the same photos to multiple sites.
  • Explicitly stating that free services collect and monetize specified PI is presented as an alternative to relying solely on complicated privacy policies.

Bulk data mechanism

The paper proposes bulk-oriented PI mechanisms because users’ valuations do not vary substantially with the quantity of otherwise similar information. Its broader conclusion emphasizes that PI type matters more than amount and that service improvement is preferred to revenue generation.

  • The proposed bulk-data mechanism responds to participants assigning similar values to one piece and ten pieces of the same information.
  • The study examines monetary PI valuation and attitudes toward collection and exploitation in context.
  • The refined Experience Sampling Method and truth-telling auction mechanism were intended to address gaps between reported preferences and actual privacy behavior.
  • Users give more importance to offline-identity PI than online-behavior PI, care about information type more than amount, and dislike revenue-generating use.
  • The authors frame the study as helping explain the economic mechanics of increasingly connected online privacy exchanges.
Loading 1112.6098v1…