dataset

A curated collection of data treated as a unit.

Meanings by sector

Creative Industries

In media and creative production, a dataset — above all a training corpus — is first a rights object: a compilation of works, performances, and likenesses whose collection and reuse are governed by licences, Article 4(3) DSM text-and-data-mining reservations, and attribution expectations. What counts is not schema or splits but clearance status, recorded item by item in a rights ledger: which items are licensed, which were scraped under a claimed exception, which contain identifiable contributors. The AI Act's duties on general-purpose model providers — a copyright policy honouring reservations and a sufficiently detailed public summary of training content — turn corpus composition from an engineering detail into a disclosure obligation.

In practice: Establish, before use, the clearance status of every source in a creative training corpus, and record licences, opt-outs, and attribution obligations item by item.

OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)

Financial Services

Under model-risk management practice shaped by SR 11-7, a dataset is a controlled model input: an extract with documented source, ownership, refresh schedule, and quality checks, whose suitability and representativeness for the modeled portfolio must be evidenced in development documentation. Validators re-perform this assessment; a dataset without lineage back to systems of record, or whose vintage no longer matches current market conditions, is a validation finding regardless of model performance. Access, retention, and change control are governed like any other model artifact in the inventory.

In practice: Document a dataset's source systems, extraction logic, quality checks, and representativeness for the modeled portfolio so that an independent validator can re-perform the assessment.

OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)

Healthcare

In clinical research and medical AI development, a dataset is a fixed, versioned patient cohort extract: records selected under documented inclusion and exclusion criteria, drawn from named source systems, covered by a specific ethics approval and consent scope. Its composition — sites, devices, demographics, label provenance — is treated as a clinical-validity question, because a model is only credible for populations the dataset actually represents. Regulators expect training and test collections to be independent of each other and representative of the intended patient population, and any change to the cohort definition creates, in effect, a new dataset requiring re-review.

In practice: Trace a medical dataset to its cohort definition, consent scope, and source systems, and judge whether its composition supports claims about the intended patient population.

OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)

Healthcare

For teams engineering machine-learning models on health data, a dataset is a partitioned artifact: training, validation, and test splits with recorded hashes, split logic, and preprocessing state, matching the AI Act's distinction between data used to fit parameters, tune the learning process, and independently evaluate the system. The operational tests are leakage and lineage: the same patient must not appear across splits, and every transformed table must be reproducible from raw extracts. A dataset that cannot be rebuilt exactly from its sources is treated as unusable evidence for a performance claim.

In practice: Construct and document patient-level train, validation, and test splits, verify that no records leak across splits, and keep every dataset version reproducible from its raw sources.

OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)

Public Administration

In official statistics and government data management, a dataset is a registered information asset: it has a designated custodian, a documented legal basis, standardised metadata, a quality declaration, and a publication and revision policy. Statistical offices operationalize it through the release pipeline — a dataset exists once it is catalogued, disclosure-controlled, and versioned for dissemination, and citizens and courts can hold the administration to what a given release stated at a given time. Ad-hoc extracts that bypass the catalogue are treated as unofficial and cannot underpin administrative decisions.

In practice: Register a government dataset with custodian, legal basis, and quality metadata, and manage its releases so that any published figure can be traced to a dated version.

OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)

Machine-readable version (JSON-LD)