big data

Data at volume/velocity/variety exceeding conventional processing; increasingly a legacy framing.

Meanings by sector

Creative Industries

In media and entertainment practice, big data means behavioural audience data at scale: streaming and playback events, engagement and completion rates, programmatic-advertising auctions, and recommender interaction logs. It is operationalized through the decisions it feeds — commissioning, scheduling, personalization, and ad pricing — and its working unit is the event stream tied to a user or household profile. Competence means moving between these signals and editorial judgment without letting completion metrics stand in for cultural value.

In practice: Interrogate the audience analytics behind a commissioning or scheduling decision: identify what the event data measure, whom they exclude, and where editorial judgment must override the metric.

OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)

Financial Services

For financial-services engineering and risk teams, big data is data whose velocity and volume force architectural choices: transaction streams scored for fraud within milliseconds, tick-level market data, and alternative data too large or fast for batch relational processing. Operationally it is defined by latency budgets, streaming architectures, and distributed storage — and by the supervisory corollary that scale relaxes nothing: data feeding models must still be shown suitable, representative, and of known quality before the models they feed can be relied on.

In practice: Design and defend a data pipeline under explicit latency and volume constraints while documenting that model input data remain suitable, representative, and quality-assured at scale.

OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)

Healthcare

In clinical research and pharmacovigilance, big data has settled into meaning linked real-world data: electronic health records, insurance claims, disease registries, genomic data, and wearable streams joined at patient level across time. The operational criteria are linkage quality, longitudinal coverage, and fitness of routinely collected data for regulatory-grade evidence — not raw volume. Teams work the concept through record-linkage pipelines, phenotype definitions, and bias assessment for data that were never collected with research in mind.

In practice: Evaluate whether routinely collected, linked health data can support a research or regulatory question: check linkage quality, longitudinal coverage, phenotype validity, and collection-driven biases.

OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)

Public Administration

In official statistics, big data denotes non-survey sources — mobile network records, satellite imagery, web-scraped prices, retail scanner data — brought into statistical production. Statisticians operationalize the term through admission tests: documented access agreements with private data holders, quality frameworks extended to 'found' data, and confidentiality safeguards, since such sources arrive without any sampling design. The concept lives in pilot studies and trusted-smart-statistics programmes that decide whether a source can be allowed to carry an official figure.

In practice: Assess a non-survey data source for statistical production: secure a stable access agreement, test coverage and selectivity against benchmarks, and apply confidentiality rules before publication.

OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)

Public Administration

For oversight bodies, courts, and civil-society watchdogs, big data in government is operationalized as scale of linkage: the joining of tax, benefits, housing, and other administrative databases into risk scores applied to citizens. What counts is not volume but the shift it produces — from case-by-case assessment to population-wide suspicion — and the proportionality test it must survive: is the intrusion justified, are groups disparately targeted, can an individual contest a score. On this reading a modest dataset is still 'big' if linkage makes citizens legible in new ways.

In practice: Subject a proposed cross-database linkage or citizen risk-scoring scheme to a proportionality and disparate-impact review, and verify that affected individuals can learn of and contest their score.

OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)

Documented disagreement

Engineering and oversight communities bound the concept incompatibly. For financial-services engineers, data are big when volume and velocity force architectural choices — streaming pipelines, distributed storage, latency budgets — and the label carries no normative charge. For courts and watchdogs reviewing government data use, bigness is measured by linkage power: a modest dataset becomes big when joining it to other registries makes populations legible and scoreable in new ways, triggering proportionality and contestability requirements. Each criterion classifies the other side's paradigm cases as not big at all: a small linked welfare-scoring scheme has no engineering scale, while high-volume unlinked telemetry raises no linkage concern.

Machine-readable version (JSON-LD)