Information relating to an identified or identifiable person; GDPR Art. 4(1) anchor.
In journalism and creative production, personal data covers the people in the story and, increasingly, the people in the model: names, images, and voices of identifiable individuals. Newsrooms work under the journalistic exemption, balancing data-protection rights against freedom of expression case by case, while generative production has made likeness and voice operational personal data — facial images subjected to specific technical processing are biometric data, and cloning a performer's voice processes their personal data whether or not they are named. The practical test is recognisability to an audience, not the presence of formal identifiers.
In practice: Judge when identifiable people in reporting or generated content bring data-protection duties into play, and apply the journalistic exemption as a case-by-case balance, not a blanket shield.
OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)
For banks and insurers, personal data is operationalized through the customer file: KYC identity documents, account and transaction histories, credit scores, and behavioral fraud signals, all mapped to retention schedules where anti-money-laundering law compels keeping what data-protection law would otherwise delete. The concept's sharpest edge is Article 22: credit scoring and automated onboarding decisions with legal or similarly significant effect trigger rights to human intervention and contestation, so classifying a data flow as personal data immediately raises the question of which decisions it feeds and whether those decisions are automated.
In practice: Map customer data flows to legal bases, reconcile AML retention duties with data-protection deletion duties, and flag flows that feed automated decisions with significant effects.
OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)
In hospital data-protection practice, personal data is read maximally: anything in the record that relates to an identifiable patient — free-text notes, images, device identifiers, and pseudonymised research extracts — remains personal data, and health data sits in the special categories requiring an Article 9 condition on top of a legal basis. DPOs apply the 'means reasonably likely' test conservatively: as long as a re-identification key exists anywhere, or rich clinical detail could single a patient out, the data stay in scope, and sharing is governed by contracts, DPIAs, and consent or statutory research bases.
In practice: Classify clinical extracts by identifiability and special-category status, identify the Article 6 and Article 9 bases for each use, and treat pseudonymised research data as still personal.
OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)
Among health-data engineers preparing extracts for sharing, personal data is treated as a property to be measured out of the data: identifiability is quantified with k-anonymity-style metrics, motivated-intruder exercises, and re-identification testing, and an extract whose residual risk falls below the level agreed with governance is handled as anonymous and shareable. On this working view, anonymisation is an engineering outcome with a threshold, not a binary legal status — the same fields can be personal data in one release environment and anonymous inside a locked-down trusted research environment.
In practice: Quantify re-identification risk for each planned release environment, run motivated-intruder tests, and record the threshold and controls under which an extract is treated as anonymous.
OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)
In statistical offices, the boundary of personal data is managed as a measurable disclosure risk: microdata are personal until statistical disclosure control — aggregation, suppression, perturbation, safe-setting access — brings the risk of singling out a respondent below documented thresholds, after which outputs are released as anonymous statistics. Confidentiality is absolute as a principle, but its implementation is quantitative: cell-size rules, dominance checks, and output-checking protocols define in practice where personal data ends. This risk-based release machinery is what allows official statistics to be published at all.
In practice: Apply disclosure-control rules and risk thresholds to decide when microdata-derived outputs may be released as anonymous statistics, and document the assessment behind each release.
OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)
The communities disagree about where personal data ends. Hospital data-protection officers, reading Recital 26 conservatively, hold that pseudonymised or detailed clinical data remain personal as long as re-identification is conceivable by any party holding auxiliary information. Statistical-disclosure practice and health-data engineering instead define the boundary by measured residual risk: once quantified re-identification risk in a given release environment falls below an agreed, documented threshold, the data are handled as anonymous. Both sides claim fidelity to the same 'means reasonably likely' test.