Source-linked AI summary
Big data, bigger dilemmas: A critical review
Hamid Ekbia, Michael Mattioli, Inna Kouper, G. Arave, Ali Ghazinejad, Timothy Bowman, Venkata Ratandeep Suri, Andrew Tsou, Scott Weingart, Cassidy R. Sugimoto
TL;DR
Big Data research and practice span many domains, but their debates remain fragmented and share unresolved conceptual and practical problematics. This paper synthesizes scientific, humanistic, policy, and trade literatures, situating Big Data within broader socio-economic developments. It identifies recurring dilemmas involving autonomy, opacity, generativity, disparity, and futurity, while highlighting gaps between technological visions and practical constraints.
Problem
Existing Big Data debates are fragmented by disciplinary and domain-specific perspectives, creating a need for critical synthesis of their common problematics and dilemmas.
Method
The paper critically synthesizes writings from the sciences, humanities, policy, and trade literature, organizing recurring issues across technical, ethical, legal, and political dimensions.
Results
The synthesis identifies common dilemmas across Big Data domains and highlights attributes requiring greater attention, including autonomy, opacity, generativity, disparity, and futurity.
Takeaways & Limitations
Big Data’s promised transformations must be considered alongside gaps between visions and realities and their implications for knowledge, privacy, intellectual property, and social equity.
Takeaways & Limitations
Big Data research remains constrained by sampling and selection biases, variable access levels, and poor replicability in sources such as Twitter data.
Abstract
from arXiv · showhide
The recent interest in Big Data has generated a broad range of new academic, corporate, and policy practices along with an evolving debate amongst its proponents, detractors, and skeptics. While the practices draw on a common set of tools, techniques, and technologies, most contributions to the debate come either from a particular disciplinary perspective or with an eye on a domain-specific issue. A close examination of these contributions reveals a set of common problematics that arise in various guises in different places. It also demonstrates the need for a critical synthesis of the conceptual and practical dilemmas surrounding Big Data. The purpose of this article is to provide such a synthesis by drawing on relevant writings in the sciences, humanities, policy, and trade literature. In bringing these diverse literatures together, we aim to shed light on the common underlying issues that concern and affect all of these areas. By contextualizing the phenomenon of Big Data within larger socio-economic developments, we also seek to provide a broader understanding of its drivers, barriers, and challenges. This approach allows us to identify attributes of Big Data that need to receive more attention--autonomy, opacity, and generativity, disparity, and futurity--leading to questions and ideas for moving beyond dilemmas.
INTRODUCTION
Big Data has rapidly expanded across academic, professional, and policy domains, generating enthusiasm alongside diverse perspectives and critical questions. The paper argues that these discussions require a critical synthesis organized around recurring dilemmas.
- Growing interest: Big Data interest grew continuously across scholarly, trade, and mass-media publications between 2008 and 2013.Figure 1 tracks publications containing “big data” in titles or keywords across five academic databases.
- Growing interest: Business, industry, and government framed Big Data as an opportunity for commerce, innovation, and social engineering, encouraging infrastructure and policy development.These developments also stimulated public debate and reflective engagement with associated risks and questions.
- Diverse perspectives: Academic, industry, and other perspectives differ in whether Big Data represents a new science, new methods, or a source of risks and pitfalls.Practitioners draw on established computing processes and techniques when interpreting Big Data practices.
- Critical synthesis: The paper synthesizes diverse literatures by treating recurring cross-domain themes as dilemmas rooted in digital technologies and broader socio-historical developments.The inquiry spans perspectives from multiple institutional settings and identifies common underlying frictions.
CONCEPTUALIZING BIG DATA
Big Data lacks consensus about its definition and scope, while competing perspectives emphasize data attributes, computational processes, or human cognitive limits. The paper also situates these perspectives within institutional alliances and gaps between technological visions and practical realities.
- Definition and scope: Only 10% of 150 OECD delegates in 2013 were prepared to define Big Data, illustrating uncertainty about its policy scope and character.Academic writing likewise describes Big Data as a moving, relative term that can be “big” in different ways.
- Product-oriented perspective: The product-oriented perspective defines novelty through data attributes such as volume, velocity, variety, and unprecedented scale.The Sloan Digital Sky Survey collected more data in its first weeks than astronomy had previously accumulated, while a successor was expected to acquire 140 terabytes every five days.
- Process-oriented perspective: The process-oriented perspective emphasizes storage, management, searching, and analysis because Big Data’s opacity, noise, and relationality challenge existing technologies.These challenges motivate advances in infrastructure, programming, computation, and statistical analysis.
- Cognition-oriented perspective: The cognition-based perspective focuses on the mismatch between Big Data’s complexity and human cognitive capacities, requiring collective or computational analysis.A terabyte may be manageable for a national security agency but overwhelming for an individual.
- Social movement perspective: The perspectives are distinct but overlapping, and none fully captures Big Data’s socio-economic, cultural, and political foundations.The phenomenon is also organized through alliances among businesses, universities, governments, technology companies, and open-source communities.
- Social movement perspective: Big Data’s ambitious visions coexist with significant gaps between expected transformations and realities on the ground.The paper examines these gaps as dilemmas spanning epistemology, methodology, technology, law, ethics, and political economy.
EPISTEMOLOGICAL DILEMMAS
The paper places Big Data’s epistemological dilemmas within longstanding debates about how data relate to knowledge, phenomena, and appearances. It emphasizes a shift from causal explanation toward predictive modeling while retaining concerns about interpretation, perspective, and error.
- Data and knowledge: Big Data renews the longstanding problem of how observations and instruments become knowledge about the world.The paper argues that these epistemological issues require philosophical contextualization rather than treatment as wholly new problems.
- Causal relations and correlations: The Appearance-from-Reality criterion requires theories not only to predict appearances but also to provide mechanisms that produce them.The paper presents Big Data as threatening both this criterion and the Common Cause Principle.
- Causal relations and correlations: Big Data intensifies the debate between data-intensive approaches that privilege correlation and approaches that seek causal explanation or coherent models.The disagreement concerns whether science can advance without mechanistic explanations, unified theories, or models.
- Causal relations and correlations: Scientific perspectivism warns that representations are limited and biased, rejecting the idea that science can transcend the human perspective.It recommends a more modest view of science as an engineering practice rather than a God’s-eye account.
- Predictions and productions: Big Data expands predictive modeling across social, biological, and cultural phenomena, including attempts to infer public mood from social-media data.The Emotive program analyzed two thousand tweets per second and classified expressions into eight emotions.
- Predictions and productions: Threat-identification systems may function erratically when emotions are reduced to fixed categories and assumed to be expressed consistently across contexts.The passage questions both the adequacy of eight emotion categories and cross-contextual consistency.
- Predictions and productions: The paper concludes that Big Data contributes to a broader shift in scientific success criteria from causal explanations toward predictive modeling and simulation.This shift pushes earlier breaks between phenomena and appearances toward their logical limit.
METHODOLOGICAL DILEMMAS
Big Data challenges the quantitative–qualitative divide by exposing how human judgment enters data generation, cleaning, selection, and analysis. It redefines longstanding concerns about relevance, validity, generalizability, and replicability rather than eliminating them.
- Methodological Dilemmas: Big Data research blurs the boundary between quantitative analysis and qualitative interpretation.Large-scale visualization and human decisions in sampling, cleaning, and statistical analysis complicate claims of purely objective data.
- Methodological Dilemmas: Data generation involves multiple social agents, remains opaque, and can produce incomplete or skewed data.Calibration, standards, infrastructure design, and other upstream choices shape the resulting dataset.
- Data Cleaning: To Count or Not to Count?: Cleaning Big Data requires selecting which attributes and variables to retain, combining mechanized labor with interpretative human judgment.These choices become especially consequential for personal data because de-identification can be undermined by re-identification.
- Statistical Significance: To Select or Not to Select: Big Data amplifies concerns that significance testing and exploratory analysis can produce patterns that do not represent anything real.Apophenia, data dredging, and cherry picking describe risks of finding apparent patterns in enormous datasets.
- Statistical Significance: To Select or Not to Select: Twitter sampling frames vary in coverage and access, making sampling bias, selection bias, and replication difficult while limiting generalization from users to people.The firehose excludes private and protected tweets, while gardenhose and spritzer access provide different proportions of public tweets.
- Methodological Dilemmas: Big Data research redefines longstanding dilemmas concerning relevance, validity, generalizability, and replicability.The paper argues that these questions remain central even as Big Data urges scholars beyond traditional quantitative and qualitative categories.
AESTHETIC DILEMMAS
Big Data visualization creates tensions between accuracy, transparency, interpretation, and aesthetic appeal. Visual encodings and algorithmic choices shape what users see, while the paper argues that digital complexity tilts the balance between truth and beauty toward beauty.
- AESTHETIC DILEMMAS: Big Data visualization can improve comprehension and decision making, but its algorithmic processes may remain hidden from users and experts.This opacity makes the accuracy or significance of aesthetically driven results difficult to evaluate.
- Data Visualization: Accuracy or Aesthetics?: Visualization requires mapping data variables to visual features such as position, size, shape, and color.The resulting display reflects conversion rules as well as the underlying data.
- Data Mapping: Arbitrary or Transparent?: For a given dataset, many possible visual encodings exist, and arbitrary design decisions can influence how the data are interpreted.The generative contribution of visualization is often hidden or dismissed as non-analytical.
- Data Mapping: Arbitrary or Transparent?: Data visualization’s mapping metaphor is contestable because the relevance of visual features is not self-evident and layouts are often arbitrary.This challenges the relationship between the represented phenomenon and its visual appearance.
- Data Visualization: Accuracy or Aesthetics?: Visualization design can favor aesthetic appeal over direct data accuracy.Graph-drawing criteria such as minimizing edge crossings and favoring symmetry strongly shape appearance, potentially improving engagement while taking longer to understand.
- Data Visualization: Accuracy or Aesthetics?: The aesthetic-information matrix places visualizations along accuracy-to-aesthetic axes for data representation and communication of meaning, but lacks a clear practical rubric.Its value lies in making the underlying tension visible, even though users may not recognize the concessions involved.
TECHNOLOGICAL DILEMMAS
Big Data technologies balance continuity with innovation as they scale storage, processing, and real-time analytics. This development also changes human involvement, creating participation opportunities alongside concerns about infrastructure-based inequality, ownership, privacy, and security.
- Computers and Systems: Continuity or Innovation?: Real-time processing is a defining goal of Big Data analytics, requiring systems to handle massive data volumes while keeping pace with I/O growth.Achieving this goal depends on how architects and developers manage continuity alongside innovation.
- Computers and Systems: Continuity or Innovation?: Parallel, distributed, grid, cluster, and high-performance computing provide technological continuity for data-intensive applications.These approaches extend computing traditions that maximize storage, networking, and processing power.
- Computers and Systems: Continuity or Innovation?: Software-based solutions such as Beowulf clusters can work well, but hardware limitations constrain processing capacity as data accumulate.This limitation motivates hardware or cyberinfrastructure solutions, including supercomputers and commodity computers.
- Computers and Systems: Continuity or Innovation?: Data warehouses and related centralized, integrated infrastructures emerged as resource-intensive responses to the gap between commodity and supercomputer approaches.Organizations choose among infrastructures according to cost, performance needs, application type, and institutional support.
- The Role of Humans: Automation or Heteromation?: Participatory personal data practices let individuals contribute measurements through forms, trackers, sensors, apps, and games, potentially challenging automated classifications.Whether this soft resistance shifts power toward smaller players remains unresolved.
- The Role of Humans: Automation or Heteromation?: Big Data infrastructure creates advantages for institutions that can afford advanced computing while expanding shared participation in data construction.The resulting opportunities remain entangled with unresolved questions about ownership, privacy, and security.
LEGAL AND ETHICAL DILEMMAS
Big Data unsettles established legal and ethical frameworks by expanding privacy, ownership, and collective-action questions. Heterogeneous data can be generated and reused without affected people’s knowledge, producing risks including profiling, discrimination, surveillance, and loss of control.
- LEGAL AND ETHICAL DILEMMAS: Technological change and shifting social expectations create legal and ethical questions about privacy and intellectual property without a single balancing principle.Policymakers must negotiate among competing individual, industrial, and societal interests.
- Privacy: Big Data privacy risks arise when consumer, social-media, medical, genomic, email, and mobile data reveal information about individuals and related people.These examples span commercial prediction, public metadata, health-data distribution, and government monitoring.
- Privacy: Heterogeneous data may be generated, transferred, and analyzed without affected people’s knowledge, then used for unforeseen purposes.The resulting harms include profiling, tracking, discrimination, exclusion, surveillance, and loss of control.
- Intellectual Property: Conventional data often receive limited intellectual-property protection because factual data generally fail patent eligibility or copyright requirements.This has prompted database companies to seek sui generis protection for databases.
- Intellectual Property: Sampling, cleaning, masking, and other human judgments can make Big Data compilations more subjective and potentially more eligible for protection.The paper presents this as a challenge to conventional assumptions about data ownership and free riding.
- Intellectual Property: Big Data’s ethical and legal challenges suggest changes to dominant frameworks concerning privacy, data ownership, and collective action.The discussion specifically identifies free riding as a phenomenon requiring reconsideration.
POLITICAL ECONOMY DILEMMAS
The paper examines how Big Data’s growth intersects with wealth, capital, user participation, and social inequality. It identifies competing accounts of user contribution and links participation with simultaneous increases in wealth and poverty, producing both empowerment and anxiety.
- Data and Wealth: The explosive growth of data coincides with explosive growth of wealth and capital, raising questions about their relationship and the mechanisms sustaining it.The paper analyzes these mechanisms at psychological, sociocultural, and political levels.
- Data and Wealth: Internet data and capital flows concentrate in social networking, search, gaming, and scientific sites, exemplified by Facebook’s advertising growth from $300 million in 2008 to $4.27 billion in 2012.These trends reinforce the view of data as an economic asset.
- Data as Asset: Contribution or Exploitation?: Commentators identify users as a source of data value, while Facebook’s user growth accompanied revenue growth of about 1300% from December 2008 to December 2013.User-generated data are described as overproduced at almost no cost, concentrating wealth among technology-platform proprietors.
- Data as Asset: Contribution or Exploitation?: Accounts of user contribution range from exploitation through unpaid free labor to value based on affective investments and flexible mobility across networks.These perspectives offer different explanations for how users contribute to digitally mediated economic value.
- Data as Asset: Contribution or Exploitation?: Regardless of the explanation adopted, user participation and contribution correlate with the simultaneous rise of wealth and poverty.The passage states this as a correlation rather than a causal relationship.
- Social Engineering: Big Data’s autonomy, opacity, and generativity intensify social engineering’s benefits and pitfalls, leaving people ambivalently empowered while experiencing anxiety and confusion.The paper connects this condition to broader contemporary dilemmas involving mobility, privilege, and exclusion across networks.
DISCUSSION
The review identifies heterogeneous perspectives and recurring tensions in Big Data, including gaps between technological visions and practical realities, predictive benefits and opacity, and access and disparity.
- The synthesis finds diverse and heterogeneous perspectives among theorists and practitioners, producing different views of Big Data’s issues, solutions, and priorities.
- Big Data creates a gap between articulated visions and practical reality that varies across domains and requires policy innovation and practical compromise.
- Big Data’s predictive capacity can support preventative measures and better decisions, but data-driven approaches remain imperiously opaque across major domains.
- Balancing Big Data’s light and dark sides requires combining human judgment with technological prowess rather than denying either aspect.
- Big Data’s resource requirements create visible data divides among technology companies and users, governments and citizens, and large and small organizations.
BEYOND DILEMMAS: WAYS OF ACTING
The paper argues that Big Data dilemmas combine the dual character of digital technologies with greater scope, scale, and complexity. Moving beyond them requires historically informed analysis and concerted dialogue among academics, policymakers, and the public.
- BEYOND DILEMMAS: WAYS OF ACTING: Big Data dilemmas are partly continuous with earlier eras but novel in their scope, scale, and complexity.
- BEYOND DILEMMAS: WAYS OF ACTING: Dilemmas do not easily yield solutions, but understanding their origins, dynamics, drivers, and alternatives can reveal pathways for action.
- BEYOND DILEMMAS: WAYS OF ACTING: The review links Big Data to questions about theory’s role in structuring otherwise intractable information in data-driven environments.
- BEYOND DILEMMAS: WAYS OF ACTING: Social media provide a useful locus for investigating Big Data dilemmas because they combine vast data, ambiguous ownership and privacy, and widespread participation.
- BEYOND DILEMMAS: WAYS OF ACTING: Big Data raises unresolved questions about privacy, responsibility, equitable rewards, and alternatives to a heavily polarized economy.
- BEYOND DILEMMAS: WAYS OF ACTING: The human–machine division can affect human prosperity, dignity, and freedom because technologies may impose a logic on relationships between people.
- BEYOND DILEMMAS: WAYS OF ACTING: Big Data’s complexity and ubiquity call for systematic investigation and concerted conversation among academics, policymakers, and the public.