Source-linked AI summary

Undefined By Data: A Survey of Big Data Definitions

Jonathan Stuart Ward, Adam Barker

arXiv:1309.5821v1cs.DB

TL;DR

Big data lacks a consistent definition because its academic, industrial, and media origins produced diverse and contradictory accounts. This paper surveys established definitions and synthesizes them into a concise account focused on storing and analyzing large or complex datasets with multiple techniques. It concludes that big data is best described through the interplay of dataset characteristics and technologies, while conventional-capacity definitions remain temporally unstable.

  • Problem

    Big data has diverse and contradictory definitions, creating ambiguity that motivates a concrete definition.

  • Method

    The paper collates established definitions from organizations, literature, and Google Trends, then extrapolates their shared factors into a concise definition.

  • Results

    The proposed definition describes big data as storing and analyzing large or complex datasets using techniques including NoSQL, MapReduce, and machine learning.

  • Takeaways & Limitations

    Big data encompasses at least one of size, complexity, or technology, with most surveyed definitions combining at least two factors.

  • Takeaways & Limitations

    Definitions based on exceeding conventional computational capacity have continually moving boundaries as computation advances.

Abstract

from arXiv · show

The term big data has become ubiquitous. Owing to a shared origin between academia, industry and the media there is no single unified definition, and various stakeholders provide diverse and often contradictory definitions. The lack of a consistent definition introduces ambiguity and hampers discourse relating to big data. This short paper attempts to collate the various definitions which have gained some degree of traction and to furnish a clear and concise definition of an otherwise ambiguous term.

1. BIG DATA

Big data lacks a single definition because its academic, industrial, and media origins produced diverse and contradictory accounts. The paper surveys these accounts and proposes a definition centered on large or complex datasets, their storage and analysis, and associated technologies.

  • Big data’s shared provenance across academia, industry, and media has produced multiple, ambiguous, and contradictory definitions.
  • Existing definitions variously emphasize size, complexity, technology, data rates, formats, or the computational challenge of processing data.
  • Gartner’s three Vs—Volume, Velocity, and Variety—describe increasing data magnitude, production rate, and range of formats, but provide no numerical threshold.
  • Oracle defines big data through augmenting relational business decision-making with diverse unstructured sources and emphasizes technologies including NoSQL, Hadoop, HDFS, R, and relational databases.
  • Google Trends associates big data with analytics, Hadoop, NoSQL, machine learning, and other technologies, while indicating that one technology alone is insufficient.
  • A complexity-based account argues that even small datasets may be big when their permutations and interactions exceed conventional handling capabilities.
  • The conventional-capacity definition has moving goalposts because advances in computation can shrink what counts as big data.
  • The paper’s proposed definition describes storing and analyzing large or complex datasets using techniques including NoSQL, MapReduce, and machine learning.
Loading 1309.5821v1…