Information relating to an identified or identifiable person; GDPR Art. 4(1) anchor.
In agri-data practice, personal data reaches much further than office files because most European holdings are sole proprietorships: parcel geometries, subsidy claims, yield maps, animal registers, and even machine telemetry relate to an identifiable person and follow GDPR rules. Location is the operational crux — a parcel polygon links through cadastre and public registers to its operator, so 'anonymized' agronomic data that still carries geometry usually is not. Practitioners classify datasets by identifiability chain: data about the land is treated as data about the farmer until the link is genuinely severed, which coarse aggregation, not mere name-stripping, is required to achieve.
In practice: Trace the identifiability chain from parcel and machine data to the natural person operating the holding, and treat geometry-bearing farm data as personal until aggregation genuinely severs the link.
OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)
In journalism and creative production, personal data covers the people in the story and, increasingly, the people in the model: names, images, and voices of identifiable individuals. Newsrooms work under the journalistic exemption, balancing data-protection rights against freedom of expression case by case, while generative production has made likeness and voice operational personal data — facial images subjected to specific technical processing are biometric data, and cloning a performer's voice processes their personal data whether or not they are named. The practical test is recognisability to an audience, not the presence of formal identifiers.
In practice: Judge when identifiable people in reporting or generated content bring data-protection duties into play, and apply the journalistic exemption as a case-by-case balance, not a blanket shield.
OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)
In security-service legal practice, personal data is the category whose reach into intelligence holdings is perpetually negotiated: watchlist entries, biometric enrollments, intercept products, and travel records are plainly information about identifiable persons, but the applicable regime depends on the mission, general data-protection law recedes where national-security exemptions apply, and specialized oversight, commissioners, tribunals, and courts, takes its place. Operationally, services run parallel compliance tracks: civil-security processing under data-protection law with its access and redress rights, and national-security processing under warrantry and proportionality regimes, with courts repeatedly redrawing the line by testing bulk holdings against necessity and safeguard requirements.
In practice: Classify each processing operation by its governing regime, apply the corresponding safeguards and oversight mechanisms, and track case law that moves the boundary between data-protection and national-security frameworks.
OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)
In education data-protection practice, personal data covers nearly everything an institution holds about learners, most of whom are minors: enrollment records, submitted work, photos, platform activity logs, and the especially sensitive layers of special-educational-needs and safeguarding notes. GDPR's child provisions structure daily practice: Article 8 sets the consent age for online services offered to children, and the recitals' insistence that children merit specific protection makes legitimate-interest and consent analyses run stricter than for adults, consent from a pupil in a compulsory setting is rarely freely given. Student work is itself personal data, a point institutions rediscover whenever it is proposed as training material.
In practice: Classify what the institution holds about each learner, including work, logs, and safeguarding notes, apply children's heightened protections, and identify the lawful basis for every use beyond teaching itself.
OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)
In plant data-protection practice, the working question is when machine data becomes people data: a vibration spectrum is not personal, but a cycle-time log joined to a login, an andon event tagged with a badge, or hall-camera footage is — and identifiability travels through joins, so a 'machine-only' stream becomes personal the day someone links it to the shift roster. Operationally, plants classify data flows by linkability to operators, apply the strict regime (legal basis, transparency, works-council agreement, retention limits) to the linkable ones, and design the rest to stay unlinkable. In small crews the line is stricter than it looks: 'the operator of press 4 on night shift' identifies one known person, whatever the field names say.
In practice: Classify every plant data flow by operator linkability including indirect joins and small-crew inference, and apply data-protection controls to linkable flows before analytics touch them.
OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)
For banks and insurers, personal data is operationalized through the customer file: KYC identity documents, account and transaction histories, credit scores, and behavioral fraud signals, all mapped to retention schedules where anti-money-laundering law compels keeping what data-protection law would otherwise delete. The concept's sharpest edge is Article 22: credit scoring and automated onboarding decisions with legal or similarly significant effect trigger rights to human intervention and contestation, so classifying a data flow as personal data immediately raises the question of which decisions it feeds and whether those decisions are automated.
In practice: Map customer data flows to legal bases, reconcile AML retention duties with data-protection deletion duties, and flag flows that feed automated decisions with significant effects.
OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)
In hospital data-protection practice, personal data is read maximally: anything in the record that relates to an identifiable patient — free-text notes, images, device identifiers, and pseudonymised research extracts — remains personal data, and health data sits in the special categories requiring an Article 9 condition on top of a legal basis. DPOs apply the 'means reasonably likely' test conservatively: as long as a re-identification key exists anywhere, or rich clinical detail could single a patient out, the data stay in scope, and sharing is governed by contracts, DPIAs, and consent or statutory research bases.
In practice: Classify clinical extracts by identifiability and special-category status, identify the Article 6 and Article 9 bases for each use, and treat pseudonymised research data as still personal.
OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)
Among health-data engineers preparing extracts for sharing, personal data is treated as a property to be measured out of the data: identifiability is quantified with k-anonymity-style metrics, motivated-intruder exercises, and re-identification testing, and an extract whose residual risk falls below the level agreed with governance is handled as anonymous and shareable. On this working view, anonymisation is an engineering outcome with a threshold, not a binary legal status — the same fields can be personal data in one release environment and anonymous inside a locked-down trusted research environment.
In practice: Quantify re-identification risk for each planned release environment, run motivated-intruder tests, and record the threshold and controls under which an extract is treated as anonymous.
OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)
A law firm's matter files are among the densest personal-data holdings in any sector: clients' affairs, opponents' finances, witnesses' statements, employees' conduct — much of it special-category or criminal-offence data processed under the legal-claims bases rather than consent. Counsel operationalize the category twice over: as advised subject matter, applying the expansive identifiability test to client processing, and as the firm's own compliance problem, where data subject access requests from opponents or former employees collide with privilege — the firm must locate and release the requester's personal data while withholding what privilege and third-party rights protect, an extraction exercise running through every matter the requester touches.
In practice: Identify the legal bases carrying personal data through each matter, and answer access requests by extracting the requester's data while defending privilege and third-party redactions line by line.
OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)
In transport operations, the trap is that operational data is personal data: a truck's GPS trace is the trace of its driver's working day, tachograph files are working-time records of named individuals, every parcel scan timestamps a courier's movements, and delivery data carries consignee names, addresses, phone numbers, signatures, and doorstep photos. The working classification runs through assignment: vehicle- and device-level telemetry counts as personal wherever rosters map machines to people on shifts, which in fleet practice is almost everywhere. Operators therefore govern telematics as employee data — lawful basis, purpose limits, works-council agreements — and delivery records as customer data with retention clocks, rather than as neutral machine exhaust.
In practice: Classify telemetry by whether roster or device assignment links it to individuals, govern the linked streams as employee or customer data with declared purposes and retention, and resist the machine-data framing.
OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)
The trades of this sector hold intimate data casually: the salon's client card noting a scalp condition and its medication, the restaurant's allergy list, the care file describing a client's home, habits, and body, the WhatsApp thread with a regular. Under GDPR nearly all of it is personal data, and much — health details in care and beauty records above all — is special-category data demanding explicit safeguards, however informal its container. The workers' own traces are personal data too: GPS trails, biometric check-ins, ratings. The operational shift the sector faces is realizing that familiarity is not a lawful basis, and a shoebox is not a filing system exempt from the law.
In practice: Inventory the client and worker data your business holds, identify what is health or other special-category data, secure it accordingly, and stop treating informality as exemption.
OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)
In statistical offices, the boundary of personal data is managed as a measurable disclosure risk: microdata are personal until statistical disclosure control — aggregation, suppression, perturbation, safe-setting access — brings the risk of singling out a respondent below documented thresholds, after which outputs are released as anonymous statistics. Confidentiality is absolute as a principle, but its implementation is quantitative: cell-size rules, dominance checks, and output-checking protocols define in practice where personal data ends. This risk-based release machinery is what allows official statistics to be published at all.
In practice: Apply disclosure-control rules and risk thresholds to decide when microdata-derived outputs may be released as anonymous statistics, and document the assessment behind each release.
OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)
In ad-tech compliance, personal data is read against Article 4(1) and Recital 30's online identifiers, which capture most of the stack's working currency: cookie IDs, mobile advertising IDs, hashed emails, IP addresses, and the inferred segments attached to them all relate to an identifiable person, because singling someone out for targeting is identification even without a name. The scope determines everything downstream — what needs a consent state, what a deletion request must reach across CDP, platforms, and clean rooms, what can be retained after opt-out. Practice therefore maintains an identifier inventory: every ID the stack mints or receives, mapped to its legal treatment.
In practice: Inventory every identifier the marketing stack mints or receives, classify each as personal data under the singling-out test, and ensure consent, access, and deletion mechanics reach all of them.
OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)
In research data protection, personal data reaches much further than most researchers' intuitions: anything relating to an identifiable person, where identifiability is judged by means reasonably likely to be used, keeps its status through key-coding, so pseudonymized research data remain personal while a key exists anywhere, and rich phenotypes, genomes, or free text can single out individuals without any name attached. The consequences are operational: an Article 6 basis plus, for health and genetic data, an Article 9 research condition, Article 89 safeguards, and sharing governed by agreements rather than goodwill. The working question for any extract is not whether names were removed but whether anyone, with likely means, could still resolve a row to a person.
In practice: Classify each research extract by realistic identifiability rather than by removed fields, secure the required legal bases for personal and special-category data, and re-assess when auxiliary data could change the answer.
OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)
In systems engineering, personal data is far broader than the users table: device identifiers, IP addresses, cookie IDs, precise location, and free-text logs all qualify once reasonably linkable to a person, and identifiability is assessed against available means, not intent. Operationally the concept lives as classification — schema-level tags that mark fields as personal or sensitive and flow into retention, masking, and access enforcement — and as plumbing: data-subject access and deletion must actually traverse microservices, caches, and backups, which makes right-to-erasure an architecture requirement that systems designed without it discover late and expensively.
In practice: Classify identifiability at the schema level, propagate tags into retention and access enforcement, and design deletion and export paths across services and backups before the first user record exists.
OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)
The communities disagree about where personal data ends. Hospital data-protection officers, reading Recital 26 conservatively, hold that pseudonymised or detailed clinical data remain personal as long as re-identification is conceivable by any party holding auxiliary information. Statistical-disclosure practice and health-data engineering instead define the boundary by measured residual risk: once quantified re-identification risk in a given release environment falls below an agreed, documented threshold, the data are handled as anonymous. Both sides claim fidelity to the same 'means reasonably likely' test.