Where experts disagree
- accountability — Communities disagree on where accountability's boundary lies. One school, rooted in GDPR Article 5(2) and audit practice, bounds it at demonstrable conformity: an actor is accountable when complete, retrievable evidence shows obligations were met and decisions can be reconstructed. Another school, rooted in administrative justice and frontline practice, holds that records alone cannot constitute accountability: the concept extends to a human who explains the decision to the affected person, faces consequences, and can provide remedy. Both camps value documentation; they dispute whether it is the whole of the concept or merely its precondition.
- accountability — In AI-assisted clinical care, one position holds that the treating clinician must remain the single, undiluted locus of accountability for every decision, on the ground that patient safety depends on one professional who cannot point elsewhere. The opposing position holds that accountability must be distributed across the value chain, spanning manufacturer, deploying organization, and clinician, in proportion to actual control, arguing that loading residual responsibility onto the frontline user is both unfair and unsafe because it shields the actors who control design, training data, and updates. The disagreement concerns what accountability arrangements should protect: the clarity of a single answerable professional, or the alignment of answerability with control.
- accuracy — Communities draw the boundary of 'accuracy' around different objects. Model-risk and machine-learning practice defines it as a distributional property of a system over a population — a rate monitored within tolerance bands, under which individual errors are expected and managed. Data-protection practice defines it as a property of each stored record about each person, enforceable individually through rectification rights. A system can satisfy one reading while violating the other.
- aggregation — Communities read the evidence on aggregation's protective power differently. Health-reporting practice treats thresholded aggregate tables as privacy-safe by construction, pointing to decades of routine releases without demonstrated harm. Statistical-disclosure specialists point to differencing and database-reconstruction attacks showing that aggregation without formal guarantees can be reversed at scale, and conclude that 'aggregate' cannot be equated with 'anonymous'.
- algorithm — Technical communities reserve 'algorithm' for a precisely specified, finite input-output procedure, kept distinct from the trained model that supplies parameters and from the deployed system around both. Creators, affected citizens, and platform-accountability reviewers use 'algorithm' for the entire sociotechnical decision or ranking system — code, models, tuning, policy, and operators — because only that whole is observable, governable, or contestable from where they stand. Both usages are internally coherent and serve real work in their communities, and the clash surfaces even within a single sector: feed engineers decompose exactly the pipeline that platform auditors insist on treating as one accountable unit.
- algorithm — Governance communities disagree on whether deterministic, human-authored decision rules belong inside formal algorithmic governance. Model-risk practice attaches its full validation apparatus to statistical and machine-learning estimation and leaves rule engines to lighter IT change control, a boundary echoed by the AI Act's exclusion of systems executing solely human-defined rules. Equity auditors and public-sector overseers instead scope governance by consequence, citing deterministic procedures — welfare debt formulas, exam standardization, race-corrected clinical equations — whose documented harms matched or exceeded those of learned models.
- alignment — One set of communities scopes alignment as a property of the model that can be specified, measured, and certified before deployment — evaluation suites, refusal metrics, adversarial-testing thresholds, auditable risk-mitigation artifacts — while another scopes it as an ongoing socio-political question about whose values govern the system, which no pre-deployment measurement can close because those values are contested and legitimately change through democratic and legal process. The same word names a test result for one side and a governance process for the other.
- anonymization — Communities draw the boundary of anonymized data in incompatible places. One position, anchored in the Article 29 Working Party's opinion, holds data anonymized only when re-identification is prevented against all means reasonably likely to be used, irreversibly and independent of context, with singling out, linkability, and inference all defeated. The other, anchored in risk-based regulatory guidance and statistical-disclosure practice, holds anonymization to be a contextual judgment: data counts as anonymous when residual risk is remote given the specific environment and controls around it, since zero risk is unattainable for any data that retains utility. The same boundary dispute recurs wherever pseudonymous identifiers circulate commercially: European regulators class hashed emails and rotating device identifiers as personal data, while industry contracts routinely label them anonymous.
- anonymization — Communities read the empirical record of re-identification attacks in opposite ways. Privacy engineers cite demonstrations such as the Netflix Prize linkage, credit-card metadata unicity, and genomic surname inference as proof that record-level anonymization of rich data fails structurally, so only mechanisms with formal guarantees deserve the name. Statistical-disclosure and health de-identification practitioners read the same attacks as strikes against naive public releases, pointing to decades of controlled official-statistics outputs without demonstrated re-identification harm as evidence that well-governed traditional methods remain sound.
- artificial intelligence — Communities disagree about whether 'artificial intelligence' names a bounded category with membership criteria or a region on a continuum of model complexity. EU legal practice, following AI Act Article 3(1), requires a yes/no classification because obligations attach to membership; model-risk practice in finance holds that inference-producing systems differ only in degree, and that a bright line between 'AI' and conventional statistical models is arbitrary and unstable as techniques evolve.
- audit — Communities disagree over what qualifies an examination as an 'audit'. The assurance tradition (internal audit, supreme audit institutions, certification bodies) reserves the term for independent examinations against defined criteria under a mandate, with auditee cooperation, professional standards and formal reporting duties. The investigative tradition in journalism and research applies it to adversarial external probes of system behavior conducted without access, agreed criteria, or the operator's consent, grounding authority in method transparency and public interest instead.
- automated decision-making — Communities disagree about where in an automated pipeline the 'decision' occurs and how much human involvement removes a decision from the automated category. Established lending compliance attaches Art. 22 duties at the final act on the customer, treating upstream scores as inputs and a trained human checkpoint as sufficient de-automation. The post-SCHUFA data-protection reading and administrative-law oversight practice instead look to determinative effect: a score or system output that is systematically followed is itself the decision, and pro-forma human review does not change its automated character.
- benchmark — Communities disagree about what a benchmark score is evidence of. Regulatory and procurement actors increasingly treat standardized benchmark results as operative evidence of system capability, sufficient to trigger legal classification, obligations, or acceptance decisions. Clinical researchers and critical observers counter that benchmark scores measure performance on a narrow, often contaminated proxy task and warrant no presumption about real-world capability without independent, context-specific validation. The same score is decisive evidence in one room and a marketing claim in the next.
- bias — Communities that build and communities that scrutinize public risk-scoring tools operationalize an unbiased score through incompatible criteria. Developers and vendors center calibration within groups: a given score must correspond to the same outcome probability for every group. Journalistic, equality-body, and civil-rights auditors center error-rate parity: false-positive and false-negative burdens must not fall unequally on protected groups. Impossibility results (Chouldechova 2017; Kleinberg, Mullainathan and Raghavan 2016) show that when base rates differ no score can satisfy both, so choosing a criterion is a normative choice about which errors matter, not a technical refinement.
- bias — Quantitative traditions — estimation theory, actuarial science, official statistics — draw the concept's boundary at systematic deviation of an estimate from a true value: directional, measurable, sometimes deliberately introduced (regularization, shrinkage), and not intrinsically about people or harm. Critical, audit-oriented communities draw it around patterned disadvantage produced by sociotechnical systems: bias is inseparable from power, institutions, and injury, and a purely numerical reading misses the object. NIST SP 1270's systemic/human/statistical taxonomy institutionalizes the broad boundary; estimation-theoretic practice institutionalizes the narrow one.
- big data — Engineering and oversight communities bound the concept incompatibly. For financial-services engineers, data are big when volume and velocity force architectural choices — streaming pipelines, distributed storage, latency budgets — and the label carries no normative charge. For courts and watchdogs reviewing government data use, bigness is measured by linkage power: a modest dataset becomes big when joining it to other registries makes populations legible and scoreable in new ways, triggering proportionality and contestability requirements. Each criterion classifies the other side's paradigm cases as not big at all: a small linked welfare-scoring scheme has no engineering scale, while high-volume unlinked telemetry raises no linkage concern.
- data governance — Communities disagree about what the term denotes. In finance and enterprise practice, data governance is an internal control discipline over data assets — named owners, lineage, quality thresholds — evidenced to auditors and supervisors. In health and public administration it names the institutional allocation of decision rights and answerability over data use, extending to patients and citizens; EU legislation (the Data Governance Act) has fixed this wider sense in law. Both meanings are current, and each side hears the other's usage as either too narrow or too vague.
- data literacy — Communities disagree about what data literacy is for and where it resides. Clinical and organizational framings treat it as an individual or workforce competence: skills that let a person interpret and use data and AI outputs correctly, deliverable through training and verifiable through records. Critical civic framings hold that this deficit model misplaces the burden: literacy is a collective, political capacity to question and reshape the data regimes that classify people, and training individuals to read outputs leaves the regimes themselves unexamined.
- data minimization — The communities disagree about when and how 'necessary' is established. Administrative lawyers, hospital data-protection officers, and ethics boards treat necessity as an ex-ante legal test that must be satisfied before any collection occurs: an item without a demonstrable basis is unlawful to gather, whatever its potential utility. Model developers treat necessity as an empirical property demonstrated during development: candidate data is explored under controls, and the minimized feature set is the output of ablation and pruning rather than a precondition of exploration.
- data provenance — Communities disagree about what data provenance covers. Engineering-oriented practice in healthcare and finance bounds provenance at the verifiable processing record: where data physically originated and which transformations produced its current form, judged by reconstructability. Rights- and consent-oriented practice — creative-sector chain-of-title and health-data access governance alike — bounds it at authorization: whose works, permissions, and consent scopes the data embodies, judged by whether uses stay inside that inherited envelope. Each side treats the other's criterion as outside the concept, so the same dataset can be provenance-complete and provenance-void simultaneously.
- data provenance — Communities that agree provenance documentation must exist disagree about whom it is for. Journalistic and public-accountability practice holds that provenance is only realized when origin can be stated to audiences, citizens, and courts, with secrecy as a per-element exception. Clinical stewardship and financial model governance hold that provenance records are confidential custody instruments for regulators and validators, since the trail itself exposes patient linkages or commercially sensitive sourcing. Both positions claim accountability while prescribing opposite disclosure defaults.
- data quality — Builder and analyst communities in healthcare and finance define data quality relationally, as fitness of data for a specified use, with checks and thresholds re-justified per task; integrity-focused audit communities in the same sectors define it as an intrinsic property of records and their production controls — attributability, traceability, reconciliation to authoritative sources — assessable and reportable without reference to any downstream use. Both usages are institutionally entrenched: fitness-for-use language dominates data-science and engineering standards, while inspection and control regimes (GxP data integrity, risk-data aggregation) codify the intrinsic reading.
- data quality — Statistical and corpus-building communities locate data quality in the aggregate product — reliable estimates, well-distributed corpora — and accept individual record errors that wash out at scale; caseworker, consumer-credit, and data-protection communities locate it in each individual record, because a decision about a specific person stands or falls on that person's data being correct. The two commitments assign opposite priorities to the same error: negligible to the aggregate, decisive to the individual, and each side's quality assurance leaves the other's core concern unmeasured.
- data sharing — Health-research communities and government data-protection practice want opposite defaults. Researchers treat sharing as a duty owed to participants and to science — non-sharing is waste that must be justified — and build repositories and access committees to make it routine. Public-sector legal practice treats every share as an exception demanding an identified gateway, purpose compatibility, and signed agreements before anything moves. Each default is principled: one maximizes value from data already collected, the other guards purpose limitation and citizens' expectations.
- discrimination — Actuarial and risk-modeling communities hold that differentiation is discriminatory only when it lacks statistical or actuarial justification: risk-based differentiation is the fair outcome, and forcing equal treatment of unequal risks creates cross-subsidy injustice. Legal and administrative communities hold that using protected grounds or their proxies is discrimination regardless of predictive validity: the wrong lies in the ground of differentiation itself, so actuarial soundness is no defense, as Test-Achats established for gender-based insurance pricing.
- explainability — Communities deploying AI in high-stakes settings disagree about whether current post-hoc explanation methods are faithful enough to support reliance on individual outputs. Frontline professional practice — clinicians reconciling case-level rationales with the chart, underwriters translating attribution-derived reason codes into adverse-action notices — treats case-level explanations as a workable and necessary ingredient of safe, accountable use, enabling users to check and override systems. The skeptical side cites fidelity failures — explanations insensitive to model internals, attributions that describe a surrogate rather than the deciding model, reason codes that reorder under retraining — and concludes that reliance should rest on rigorous validation or on inherently interpretable designs, with persuasive but unverified explanations regarded as a hazard rather than a safeguard.
- explainability — Sectors disagree about where explainability's boundary lies: whether it denotes a technical property of models realized by attribution methods and tooling, or the communicative and legal act of giving an affected person accurate, specific, contestable reasons for a decision. Builder communities bound the concept at the pipeline artifact — attributions computed, logged, and quality-checked — treating translation into reasons as someone else's task, while compliance, administrative-law, and advocacy communities bound it at the recipient: without legally and practically usable reasons there is no explainability, however sophisticated the tooling.
- fairness — Formal fairness criteria are mutually incompatible when base rates differ across groups: a risk score cannot in general be simultaneously calibrated within groups and equal in false-positive and false-negative rates across them. Credit-risk builders treat within-group calibration as the fairness baseline and log residual error-rate gaps as monitored differences; clinical and public-sector builders privilege error-rate parity because misclassification costs fall directly on patients and citizens. Each community's release gate would fail models the other would pass.
- fairness — Communities draw the boundary of fairness work at protected-attribute use in opposite places. For many European model-building teams, fairness and data-protection compliance mean excluding protected attributes and their proxies from collection and processing altogether. For health-equity and public-sector auditors, fairness work begins with collecting those same attributes under safeguards, because disparities cannot be measured, nor non-discrimination demonstrated, without disaggregated data. Each side treats the other's core practice as the problem itself.
- generative AI — Sectors disagree about when synthetic content must be disclosed. Public-integrity institutions hold that generated media should be machine-marked and labeled by default, because deception risk is structural and detection after distribution fails. Creative practice ties disclosure to deceptive intent and context: generative tools are legitimate craft, and blanket labeling stigmatizes lawful work while doing little against bad actors who will not comply. The AI Act institutionalizes the tension without resolving it — provider-side marking is unconditional while deployer-side labeling is relaxed for evidently artistic works.
- ground truth — Communities disagree about the epistemic status of ground truth itself. The operational-fixture position — dominant among model builders in healthcare and finance — treats the reference label set as fixed and definitive for training and scoring: disagreement among labelers is noise to be adjudicated away so optimization has a stable target. The constructed-reference position — held by clinical epistemologists, official statisticians, and equity auditors — insists every reference standard is itself a fallible, socially situated measurement with an error structure, and that treating it as truth launders those errors into every downstream metric.
- ground truth — Where the authoritative record and observed reality diverge, communities disagree about which one is ground truth. The legal-authority position — administrative decision-makers, and analogously model-risk governance in finance — holds that the designated register or approved outcome definition is definitive because legal certainty and mass administration require a single authoritative account, correctable only through due process. The empirical-verification position — caseworkers, auditors, journalists — holds that the record is evidence, not truth, and observed reality must be able to overturn it.
- hallucination — Communities shipping tightly scoped, retrieval-grounded applications read claim-level attribution evaluations as showing that fabrication can be engineered down to negligible levels for bounded tasks: constrain generation to an authoritative corpus, require citations, force abstention on low support, and measured unsupported-claim rates on curated benchmarks approach zero. Validation and clinical NLP communities read the same evaluation literature, together with the character of sampling-based decoding, as showing an irreducible floor: rates fall but never reach zero, curated benchmarks systematically undercount open-ended fabrication, and a certified zero is an artifact of the test set — so verification and monitoring controls remain permanently necessary parts of any deployment.
- hallucination — Within the creative sector, entertainment and game-development communities prize the generative behavior the term condemns: producing people, places, and events unconstrained by fact is the creative product, and hallucination misnames it as malfunction. Journalism and editorial-standards communities hold that any fabricated factual assertion in published matter is a professional breach whatever the tool, and object that the softer clinical-sounding term launders what their codes plainly call fabrication. The two camps do not disagree about what the systems do — they disagree about whether the word names a harm.
- human oversight — There is genuine disagreement about what establishes that human oversight exists. Compliance-oriented communities in finance and public administration operationalize oversight as demonstrable arrangements: designated competent reviewers, interface tools, override authority, and auditable logs — if the prescribed capabilities are designed in and documented, oversight is established. Clinically and critically oriented communities read the empirical record on automation bias, alert fatigue, and rubber-stamping as showing that such arrangements frequently fail to produce actual control, and therefore count oversight as existing only where humans are shown to detect, override, and change outcomes under realistic conditions.
- impact assessment — Communities disagree about what an impact assessment is for. In regulated-industry practice it is a compliance gate: an internal document that classifies risk, routes a change to the right controls and evidences diligence for supervisors, completed before approval and updated on change. A participatory tradition, rooted in environmental-assessment lineage and algorithmic-accountability scholarship, holds that assessments exist to surface harms visible only to affected people, so consultation, publication and ongoing reassessment are constitutive, not optional extras.
- inference — Communities draw the boundary of what the word inference names at three different places. Statistical and official-statistics practice reserves it for reasoning from data to population-level claims under explicit uncertainty; ML engineering and deployment practice uses it for the runtime execution of a trained model on individual cases; data-protection and administrative-law practice treats the inference as the derived personal datum itself — the new fact recorded about a person — whatever process produced it. Each usage is entrenched in standards, tooling, billing models, case law, and job descriptions, so no community treats its reading as metaphorical or secondary.
- informed consent — Communities in the creative industries disagree about what legitimates the use of creative works, voices, and likenesses in generative-AI development. Performer representatives, commissioning decision-makers, and many working creators hold that prior, specific, and typically compensated opt-in consent is required before material is used for training or replication. Model developers operationalize legitimacy as lawful access under text-and-data-mining exceptions combined with honoring machine-readable opt-outs, reserving affirmative consent for targeted replication of identifiable individuals. Both sides use the word consent for these incompatible mechanisms.
- informed consent — Sectors draw the boundary of what counts as informed consent in incompatible places. Financial-services audit practice treats consent as a documented compliance event: a well-formed, timestamped record satisfying legal elements settles the matter. Health-research ethics treats consent as an ongoing relational process in which signed forms are merely evidence, with comprehension and continued willingness as the real criteria. Critical public-sector stewards go further, holding that where refusal carries costs, consent is structurally unavailable regardless of record or process quality.
- interpretability — Communities disagree on where interpretability's boundary lies. One school holds it is an intrinsic property of transparent model classes whose parameters can be read directly, so post-hoc explanations of black boxes fall outside the concept. A second holds that a black-box model becomes interpretable once constraints, attribution methods, and stability tests characterize its behavior to a demonstrated standard; a behavioral variant of this school claims interpretability whenever output changes can be reliably predicted from defined changes to inputs, prompts, or model versions. The schools use the same term for different objects: the model itself versus the evidence assembled around it.
- interpretability — Communities disagree about whose understanding interpretability must serve. One position treats interpretability as satisfied when qualified professional audiences, such as validators, risk committees, hospital adoption committees, and clinician users, can understand and challenge the model, since they carry the control responsibility. The opposing position holds that interpretability is owed to the people affected by decisions: patients, citizens, and customers must be able to grasp and contest the grounds of an outcome, and expert-only understanding leaves the affected person's position unchanged.
- missing data — The communities disagree about what kind of object missing data is. Trial statisticians and model developers treat it as a property of a dataset: classify the mechanism, apply a licensed repair, document the effect, and the analysis can proceed. Equity practitioners and data journalists treat it as a property of the measurement system: absence records who institutions fail to see, so the correct response is to change what gets measured and to report the gap itself — repairing the dataset without naming the absence hides the finding.
- model — Governance and transparency communities define a model by decision influence: any quantitative or automated logic that turns inputs into estimates or decision-shaping outputs — including scorecards, spreadsheets, and hand-coded rules — so that oversight coverage is complete. Engineering communities define a model as the trained, versioned parameter artifact produced by a learning pipeline, categorically distinct from rules, thresholds, prompts, and surrounding code. Both usages are internally coherent but draw the term's boundary in incompatible places, and each community's inventories, registers, and controls presuppose its own boundary.
- model — Communities disagree about what evidence establishes that a model is adequate. The representational tradition, strong in official statistics and regulated quantitative modeling, requires that a model's assumptions about the data-generating process be stated, tested, and scientifically defensible; good fit alone is insufficient. The predictive tradition, strong in clinical and industrial machine learning, treats demonstrated out-of-sample performance on the intended population as the acceptance criterion, with mechanistic plausibility desirable but not required. The split mirrors Breiman's 'two cultures' of statistical modeling, and each side reads the other's evidence dossier as missing the point.
- open data — The sectors draw the boundary of 'open' incompatibly. Public administration operationalizes open data as a universal reuse right: anyone, any purpose, open licence, at most attribution — access controls disqualify the label. Health research and finance attach 'open' to governed access arrangements — controlled-access repositories and consented, authorization-gated APIs — where data are neither public nor freely reusable, and openness means a discoverable, rule-bound route in. Both usages are institutionally entrenched, one in open-government law and the Open Definition, the other in research-governance and open-banking regimes, so the same word licenses very different expectations about who may obtain and reuse the data.
- outlier — The communities draw the concept's boundary differently: for statistical producers an outlier is a data value threatening estimate quality, to be verified and contained by procedure; for fraud and clinical teams it is the signal of interest, the very thing detection exists to surface and investigate; for critics of administrative AI it is a person whose atypicality a system converts into suspicion or degraded performance. The same flagging operation is therefore quality control, detection success, or due-process harm depending on which reading holds.
- personal data — The communities disagree about where personal data ends. Hospital data-protection officers, reading Recital 26 conservatively, hold that pseudonymised or detailed clinical data remain personal as long as re-identification is conceivable by any party holding auxiliary information. Statistical-disclosure practice and health-data engineering instead define the boundary by measured residual risk: once quantified re-identification risk in a given release environment falls below an agreed, documented threshold, the data are handled as anonymous. Both sides claim fidelity to the same 'means reasonably likely' test.
- privacy — Communities disagree about when data about people has been made private enough to circulate. One position treats privacy obligations as discharged once a recognized anonymization or de-identification threshold is met and documented: the data ceases to be personal and leaves data-protection scope. The other holds that re-identification risk is continuous and never reaches zero — especially for rich clinical, genomic, longitudinal, or small-area data — so only quantified leakage bounds such as a differential-privacy budget, or adversarial attack evaluations, can substantiate a privacy claim; a threshold sign-off merely records an opinion about an unmeasured risk. Both positions read the same releases and the same re-identification literature and draw opposite conclusions about sufficiency.
- privacy — Communities disagree about privacy's normative weight when it collides with mandated or legally protected public-interest activity. Data-protection supervisors and civil-liberties stewards operationalize privacy as a fundamental right and structural check on institutional power whose every interference must be shown necessary and proportionate, with the burden on the intruding party. Editors and financial-crime officers operationalize privacy as one lawful interest among several, routinely and legitimately outweighed by freedom of expression in public-interest journalism or by statutory anti-money-laundering surveillance duties that customers cannot refuse. Each side regards its ordering as the legally and morally correct one, and each can cite binding law or documented harms in support, so the disagreement is about which values set the default, not about legal ignorance.
- profiling — Commercial personalization communities operationalize profiling as core value-creating infrastructure: better profiles mean better recommendations, experiences and yield, with law entering as a consent constraint on an otherwise legitimate activity. Rights-protective communities in public administration, courts and data-protection practice operationalize profiling as a presumptively dangerous exercise of classificatory power over people, tolerable only with explicit legal basis, demonstrated necessity and safeguards, and in some uses prohibited outright.
- prompt — Creative practitioners operationalize prompting as authorial craft — skilled, iterated labour that determines whether an output is usable and that merits credit, compensation, and protection as know-how — while copyright registrars and prevailing legal practice treat prompts as unprotected ideas or instructions to a machine, locating authorship only in demonstrable human expressive contribution beyond the prompt. The disagreement concerns where the boundary of authorship lies when expressive detail is produced by a model steered by human instruction, and it is sharpened rather than settled by every new registration decision.
- pseudonymization — The communities draw the boundary of personal data differently: research practice treats key-coded data as effectively de-identified for a recipient with no access to the key, grounding lighter handling of coded datasets, while the data-protection reading holds that data remain personal so long as anyone retains means to re-identify, because the test is whether the person is identifiable at all, not identifiable by the current holder. European litigation over coded data shared without keys has kept both readings live.
- risk — Regulatory product-safety regimes, including the EU AI Act, define risk exclusively as the combination of probability and severity of harm, so risk work means preventing and minimizing adverse outcomes. Enterprise risk-management traditions descending from ISO 31000, echoed in the NIST AI RMF and in commissioning practice in the creative economy, define risk as the effect of uncertainty on objectives, whose consequences can be negative or positive; there risk is deliberately taken and accepted, not only mitigated. Both usages are deeply institutionalized, and each community's documents presuppose its own scope without flagging it.
- risk — Quantitative modeling communities operationalize risk as a calibrated, testable number — predicted probabilities, expected loss, value-at-risk — because commensurability is what lets thresholds, prices, and capital consume it. Rights-focused stewards and affected creator and citizen communities counter that severity to fundamental rights and livelihoods is not commensurable: harms that concentrate on particular groups or destroy an individual practice cannot legitimately be netted against aggregate benefit, so risk must be assessed distributionally and sometimes treated as non-offsettable.
- robustness — Communities disagree about what robustness is a property of. Model-centric practitioners — clinical-ML and quantitative developers — operationalize it as a measurable attribute of the trained artifact: quantified degradation under perturbed, shifted, or stressed inputs, established by test batteries before release and versioned with the model. Governance and operations communities — production leads, public-sector auditors — operationalize it as an attribute of the whole deployed sociotechnical service: fallback workflows, monitoring, update governance, and the institutional capacity to detect and correct failure. Both use the same word for differently bounded referents, so a claim of robustness carries no fixed scope until the speaker's community is known.
- robustness — Communities that share the goal of robust systems disagree about what evidence demonstrates robustness. One school privileges designed worst-case examination before deployment: adversarial perturbation suites with certified bounds, and pre-specified adverse scenarios whose results gate approval. The opposing school holds that designed tests cannot anticipate real operating conditions — local case mix, policy changes, slow drift — so only ecological evidence counts: local prospective validation on the deploying institution's own population, and production monitoring such as drift detection and canary cohorts accumulated in situ over time.
- sampling — Communities disagree over what licenses inference from a sample to a population. Official statistics holds that only designed probability sampling with known inclusion probabilities yields defensible population estimates, and that massive found datasets remain biased regardless of size. Machine-learning practice, exemplified in healthcare AI, treats large convenience cohorts — corrected, externally validated, and checked for subgroup performance after the fact — as adequate warrant for deployment claims.
- surveillance — Public-health practice treats systematic population monitoring as a core protective duty, evaluated by how completely and quickly it detects health threats, while fundamental-rights practice treats the same monitoring as a presumptive interference with private life that is unlawful unless a necessity and proportionality case is made in advance. The word therefore carries opposite default valences: for one community more surveillance is prima facie better; for the other, less.
- synthetic data — Practitioner communities read the privacy evidence about data generation differently. Bank technology and vendor-management practice treats well-generated synthetic records as categorically non-personal: no record maps to a customer, so data-protection constraints fall away. Health-data researchers and statistical offices counter with empirical results — generative models memorize outliers, membership-inference attacks succeed against synthetic releases, and rare individuals can be reconstructed — so anonymity is a measured, per-release property rather than a consequence of the generation method itself.
- training data — Creative-industry communities disagree about the conditions under which existing works may legitimately become training data. Developer practice treats works that are lawfully accessible online as usable corpus material provided machine-readable opt-outs are honored, relying on the DSM Directive's text-and-data-mining exception with its Art. 4(3) rights-reservation mechanism and on analogous fair-use arguments elsewhere. Working creatives and the organizations that commission or license their work treat training as an act of exploitation requiring consent, licensing, credit and compensation per work, regardless of accessibility. Both sides agree the works were used; they disagree about what legitimate use requires.
- training data — Communities draw the boundary of 'training data' differently. Regulatory and audit practice follows the AI Act's definition — data used to fit a model's learnable parameters — keeping training, validation and testing data legally distinct so duties can attach precisely. Engineering practice in production settings bounds the term functionally, as the whole evolving data estate that shapes deployed model behavior: fitted partitions, tuning splits, retraining feeds and feedback loops. The statute itself concedes the line needs active maintenance — Art. 3(31) allows the validation data set to be a part of the training data set — and each side regards the other's boundary as unusable for its work: too broad to audit, or too narrow to govern.
- transparency — Communities divide over what transparency is transparency of. One cluster operationalizes it as disclosure of use: the affected person or audience must be told that AI generated or contributed to what they encounter, and transparency is achieved by an effective label or notice at the point of encounter. Another cluster operationalizes it as intelligibility of the system: documentation of intended use, performance, limitations, and working logic sufficient for a user or reviewer to interpret outputs — on this view a bare notice that AI was used achieves nothing. Both call their object transparency, and each treats the other's object as, at best, a component of the real thing.
- transparency — Parties agree that algorithmic systems must be transparent to someone but disagree about how far disclosure must extend beyond supervisors and internal reviewers. Institutional risk owners in banking and government defend graded disclosure — full access for validators, supervisors, and courts; calibrated reasons for affected individuals; restricted public detail — arguing that published logic invites gaming, evasion, and loss of protected assets. Consumer advocates, civil-society watchdogs, and some courts counter that transparency which never reaches the affected public is mere record-keeping, and that gaming and trade-secret claims must be proven rather than presumed.
- trustworthiness — Assurance-oriented communities (clinical-AI engineering, financial model validation) draw trustworthiness as a bounded, verifiable property of a system: a bundle of characteristics that can be specified, measured, evidenced, and certified. Relational communities (patients, media audiences and the editors answerable to them, citizens and agency leaders, retail-finance executives facing public scrutiny) draw the boundary around institutions and relationships: trustworthiness is standing earned over time through honest, answerable conduct, which certification evidence can support but cannot constitute. The line runs through sectors as much as between them — finance houses both a proceduralized assurance stance and a public-defensibility stance — so each side hears the other as either naive or evasive while using the same word for differently bounded objects.
- trustworthiness — Within healthcare, and with echoes in other regulated sectors, communities read the same validation evidence differently. Procurement decision-makers treat regulatory clearance and pivotal-trial results as sufficient warrant to deploy a system and direct staff reliance. Governance and audit stewards hold that such ex-ante evidence decays under distribution shift and site variation, so trustworthiness exists only while local post-deployment monitoring actively re-verifies it. The disagreement concerns what evidence warrants trust at a specific site and time, not what trustworthiness is for.
- uncertainty — The disagreement concerns what uncertainty information is owed to whom. Statistical producers hold that quantified uncertainty must accompany every published figure because suppressing it misleads users and corrodes trust. Clinical and operational decision contexts hold that uncertainty is only meaningful once translated into action thresholds, and that displaying raw probabilistic detail impairs timely, safe decisions. Both sides claim fidelity to the user; they want different goods — complete disclosure versus decidable workflows.
- validation — Across sectors, 'validation' names two different things. In machine-learning development practice — echoed in the EU AI Act's definition of validation data — it denotes a data split and tuning-and-testing stage inside model building, performed by the building team. In assurance regimes such as SR 11-7 model-risk management and medical-device regulation — echoed in the IEEE/ISO definition NIST's glossary carries — it denotes an independent, evidence-based confirmation that a system is fit for a specific intended use, producing an approval status with conditions. Both usages are entrenched in their communities, both are backed by authoritative texts, and neither reduces to the other.
- validation — Within healthcare AI, communities disagree on the evidence needed before a clinical model may be called validated. Model developers and much of the publishing research community treat strong retrospective external validation — stable discrimination and calibration on independently collected cohorts — as sufficient. Hospital deployment leaders and a growing clinical-safety school hold that only prospective, site-specific evaluation in the live workflow, on local data feeds and case mix, establishes fitness for use, because retrospective transportability has repeatedly failed to predict deployed performance.
- accountability — Communities disagree about whether accountability for system failure should terminate in a named individual who can face personal consequences or in an organization's demonstrated capacity to reconstruct, answer for, and correct its own behavior. Platform-engineering practice institutionalizes blameless postmortems, deliberately separating system accountability from individual fault on the ground that blame suppresses the reporting that makes systems safer. Senior-manager regimes in finance, signature chains in quality-managed manufacturing, and military command responsibility instead require an identified person who answers for each failure, with personal sanction available, and treat a harm with no named owner as itself the governance breach.
- accuracy — Two measurement traditions define what an accuracy claim consists of. Metrological practice in measurement-based research and production quality decomposes accuracy into trueness and precision, established by calibration against traceable reference standards, gauge studies, and a stated uncertainty budget, so a figure is meaningless without its reference procedure. Machine-learning platform practice defines accuracy as measured agreement with labels on a versioned held-out evaluation set — top-1, F1, exact match, pass@k — pinned to an evaluation harness rather than to any traceable standard. Each tradition's headline number fails the other's admissibility conditions, and methodological reviewers increasingly require authors to state which family they mean.
- accuracy — Communities working in the same monitoring and data chains attach accuracy to different objects. Map and model producers treat accuracy as a quantified statistical property of a product over a population: confusion matrices from probability samples, per-class rates with confidence intervals, and error-adjusted estimates, with individual misclassifications expected and priced into the design. Administrative and customs communities attach accuracy to the individual record a decision rests on: an inaccurate parcel flag, customs declaration, or tachograph entry makes the resulting decision against a named person unlawful regardless of the pipeline's aggregate statistics, so accuracy obligations bite record by record, with correction rights and contestation routes attached.
- algorithm — Communities draw the boundary of 'the algorithm' in workplaces incompatibly. Worker-facing communities in platform work and transport operationalize it as the whole apparatus of algorithmic management as experienced — dispatch logic, routing, scoring and deactivation rules, plus the business policies and human dispatcher discretion around them — because platform-work rights, disclosure duties, and works-council co-determination attach to that whole. Engineering and computational-research communities reserve the term for the specified procedure, deliberately kept distinct from the trained model and the deployed system, because testing, reproducibility, and liability analysis depend on the three-way separation. Where the boundary falls determines who is covered and what must be disclosed.
- alignment — Communities disagree about whose intent fixes the alignment target for optimization systems that manage work. Deployer-side practice audits an optimizer's objective and proxy metrics against what the business and its customers actually want — margin, return-rate, and lifetime-value counterweights, or the plant's accountable KPIs — and certifies alignment when the objective matches that intent within legal bounds. Worker-side practice in platform-mediated services judges alignment from below: a system is misaligned when its incentives make the people it manages work unsafely or unsustainably, whatever the deployer's objective specifies, and the operative test is watching what the incentive actually makes people do rather than auditing the objective on paper.
- alignment — For generative models, vendor and product-engineering communities operationalize alignment as conformance to the provider's policy: instruction tuning, refusal behavior, and jailbreak resistance, with success measured as pass rates on behavior suites and stability of refusal and tone across releases. Creative-production communities judge alignment as steerability to the brief — voice, style, format, and legitimately dark material within rights and disclosure boundaries — and experience the very tuning that satisfies the vendor's tests as misalignment when it flattens tone, refuses defensible content, or injects house-style blandness. The same model behavior is scored aligned by one community's instruments and misaligned by the other's.
- anonymization — Communities disagree about the unit anonymization must protect. Engineering and research-infrastructure practice operationalizes it as reduction of individual re-identification risk — singling out, linkage, inference — under a stated intruder scenario, with a release adequate when residual individual risk is remote and documented. A critical stewardship school in health-data governance and statistics policy holds that this protects the wrong unit: genomic and clinical data expose relatives and communities even when no individual is identifiable, and suppression practices erase small and marginalized populations from the statistics that justify their services, so anonymization decisions must also govern group-level harm, representation costs, and consultation with affected communities.
- audit — Communities want incompatible things from an audit's findings. The assurance and integrity traditions premise audit on findings that travel: platform engineering builds append-only evidence pipelines so external examiners can reconstruct events on demand, and research-integrity audits produce discrepancy findings that trigger public corrections, expressions of concern, or investigations. Legal practice operationalizes audit as an instrument of counsel: designed from the scoping letter to keep findings privileged work product, with purpose, direction, and custody fixed before any testing starts, and no audit commissioned without a plan for what happens if it finds something. The same examination cannot simultaneously be independently reportable and privilege-controlled.
- bias — Communities disagree over what an observed group difference proves. Assessment and clinical-audit communities treat an unexplained disparity as presumptively suspect: a validity threat that keeps results from standing, or a reportable finding that triggers corrective action, unless the instrument's operator supplies a documented justification — and aggregate accuracy is never accepted as evidence of absence. Research-methods and litigation communities allocate the burden the other way: a disparity alone establishes nothing until an identification strategy or the causal account demanded by the applicable disparate-impact or disparate-treatment standard connects it to the instrument, since the gap may faithfully record real differences. The same measured gap is a standing finding in one regime and an unestablished claim in the other.
- compliance — Communities draw the boundary of compliance in different places. Regtech and platform-engineering communities bound it at machine-verifiable conformity: obligations are encoded as rules and policy-as-code gates, evidence is generated automatically by the platform, and being compliant is a continuously monitored state of systems, with drift detected in CI rather than in audits. Legal-counselling and public-law communities hold that this captures only the codifiable fringe of the concept: compliance centrally includes judgment-demanding obligations — procedural fairness, reason-giving, proportionality, and a genuinely operating program whose existence mitigates sanctions after failure — which cannot be reduced to encoded checks, so a green control dashboard is neither necessary nor sufficient for being compliant.
- data governance — Communities disagree about whom data governance exists to serve. Platform and retail data-organization communities judge governance by the holding institution's control over its estate: catalogued assets with named owners, access policies enforced as code, lineage, retention schedules, and documented, auditable flows to third parties. Farm-data and small-service communities judge the same arrangements from the counterparty's side: governance is measured by whether the party who generated the data or whom it concerns can see who uses it, take a complete copy out, revoke access, and exit without loss — so a platform with impeccable internal controls can still constitute a governance failure through lock-in.
- data literacy — Communities disagree about what counts as evidence that data literacy exists. In education the term names a curriculum construct: a competence to be taught, progressed, and assessed against rubrics, from reading a bar chart to interrogating a claim's data source, and therefore certifiable through completed outcomes. Several practice communities explicitly reject that form of evidence for the competence they mean by the term: literacy is judged only in consequential, situated decisions — in the field, in the meeting, at the moment budget is reallocated — and is expressly not established by vocabulary, coursework, or certificates. The same training record establishes literacy under one reading and establishes nothing under the other.
- data minimization — Communities want the minimization principle to protect different goods. School and agricultural-monitoring practice applies a strict collection-time reading: every field is justified against the declared purpose before collection, possibly-useful-later fails review, derived indicators substitute for full-resolution raw streams, and records are deleted when the purpose ends rather than archived because storage is cheap. Research data-protection practice, pulled by FAIR mandates and the value of unforeseen reuse, operationalizes the same principle as safeguarded richness: variables are justified against the protocol, but the aim is the richest responsibly holdable record, with early pseudonymization and coarsening applied at release rather than at collection.
- data sharing — Communities disagree about the default posture the term implies. Research practice treats sharing as a professional norm with funder and career machinery behind it: deposit in repositories with persistent identifiers, tiered access proportionate to sensitivity, and unjustified non-sharing framed as scientific waste. Industrial supply-chain and transport communities treat the same act as a negotiated competitive disclosure: process data is process know-how and granular event streams expose networks, margins, and subcontractors, so the norm is to share the minimum derived form — conformity characteristics, computed milestones — under confidentiality terms, and demands for parameter-level or raw visibility are legitimately resisted rather than presumptively owed.
- discrimination — Communities draw discrimination's boundary in different places. Doctrine-anchored communities in legal services and education bound the concept by law: discrimination exists where differentiation on protected grounds, directly or through neutral-seeming practice, satisfies the elements of disparate treatment or disparate impact, including significance thresholds and justification tests; a disparity that never enters that machinery is exposure, not yet discrimination. Measurement-anchored communities in agri-environmental policy, health equity, security screening, and ML practice operationalize discrimination as any unexplained differential burden revealed by disaggregation, including along axes doctrine does not protect, such as holding size, deprivation, or insurance status, treating the measured gradient itself as the finding that demands design remediation.
- drift — Sectors agree that deployed models drift but disagree about what the response must protect. Operations-centered communities in MLOps, marketing, and logistics treat drift as a routine operational signal whose remedy is speed: scheduled retraining cadences, event-triggered rebuilds, and rollback, executed inside the team's own pipeline without external review. Approval-centered communities in public administration and clinical governance treat material drift as the invalidation of an authorization: the deployed model was approved on specific evidence, so drift reopens the impact assessment or governance review, and silent retraining, the other community's core hygiene practice, is itself the violation.
- explainability — Communities apply incompatible acceptance tests to the same explanation artifact. Fidelity-anchored communities in research and defense evaluation hold that an explanation is valid only if it demonstrably reflects the system's actual process, surviving re-seeding, resampling, and sanity checks, because an unfaithful account misleads whatever action it usefully prompts. Actionability-anchored communities in retail operations, manufacturing root-cause practice, agronomy, and dispatch hold that an explanation is valid only if its audience can act on it: turn a process variable, check a parcel against acquisition dates, override a route, script a retention call. These communities explicitly discount fidelity: a faithful attribution nobody can use does not count, and a usable reason is judged by the action it enables rather than by its relation to model internals.
- fairness — Communities bound fairness work differently: as a specification over model outputs or as a property of the decision process around them. ML practice operationalizes fairness as a family of formal criteria, such as demographic parity, equalized odds, and calibration within groups, implemented in evaluation harnesses, where the substantive act is the documented choice and per-release testing of a metric. Legal and examination communities operationalize fairness procedurally: notice, a hearing, published criteria applied alike, moderation, reasonable adjustments, appeal routes, and the capacity to justify this individual's treatment to them or to a tribunal. On that view parity statistics are evidence at most, never the definition, and a system can satisfy every metric while remaining unfair because a party cannot meaningfully contest it.
- fine-tuning — Communities disagree about what kind of event a fine-tune is. Engineering-anchored communities operationalize it as routine technical adaptation: a versioned campaign step or one option on a cost-and-control menu, judged by curated local data, held-out validation, and eval-gated release, with transfer error and fork maintenance as the operative liabilities. Legal-event communities operationalize the same act as a status transformation that precedes any metric: adapting a procured model on case files makes a deployer provider-like under the AI Act, tuning on matter documents moves client confidences into the weights, and tuning on mission data makes the weights a classified derivative. On that reading the fine-tune requires authorization before it runs, and the resulting artifact changes ownership, handling, and accountability.
- generalization — Two communities attach the generalization claim to different objects. Builder practice operationalizes it as an aggregate, distributional property: a model generalizes when performance on data the development process never touched, such as later periods, new users, and new tenants, matches what offline evaluation promised, with memorized duplicates treated as a hygiene problem inside an overall transfer claim. Creative-sector legal practice operationalizes it item-wise, as the boundary with memorization: a model generalizes only insofar as specific outputs arise from learned statistical structure rather than reproducing protectable expression, so a single elicited near-verbatim copy collapses the claim for that output into copying, whatever held-out benchmarks report.
- generative AI — Communities disagree about the quality regime under which generated content may reach its consequential audience. Professional-adoption communities in engineering, law, and freight operationalize generative AI as drafting technology whose every output is a standing-less draft until a named, qualified human reviews and adopts that specific artifact, a verification duty treated as non-delegable and enforced through release discipline and write-access boundaries. Production-machinery communities in software and commerce operationalize it as a feature pattern whose output quality cannot be asserted per item, only measured on samples: content ships behind eval suites, guardrails, and risk-tiered review, with individually unread outputs reaching customers by design at catalog and chat scale.
- ground truth — Communities disagree about what class of referent may bear the name and authority of ground truth. Observation-anchored sectors, including remote-sensing agriculture, battle-damage assessment, and freight operations, reserve the term for direct observation of the physical world: the surveyor standing in the parcel, the physically inspected launcher, the signed delivery in the yard; annotations and event streams are inference, and substituting them for observation must be flagged and negotiated. Procedure-anchored communities, including applied ML and technology-assisted review, operationalize ground truth as whatever the authorized labeling process produced: adjudicated annotations under versioned guidelines, or the supervising attorney's coding under a negotiated protocol, true because they survived the designated procedure, whose own error is carried as a quality budget.
- hallucination — Communities disagree about what kind of object a hallucination is and what closes the question when one occurs. Generative-AI engineering practice operationalizes it as a system-level reliability defect: a rate measured per release, budgeted per use case, and tracked through incidents, with residual fabrication reaching users acceptable inside the budget. Legal and research-integrity practice operationalizes it as an act of the person who files the output: a fabricated citation in a filing or manuscript is a sanctionable breach of candor or a research-integrity failure attributable to the professional, for whom the model's aggregate rate is neither excuse nor relevant mitigation and no nonzero budget is acceptable.
- human oversight — Communities draw the boundary of what arrangement counts as human oversight in incompatible places. Editorial and legal practice hold that oversight exists only where an accountable person actually reviews each consequential output before it takes effect, so any path by which generated content reaches an audience or a filing unread is an oversight failure. Manufacturing, commerce-automation, and research-workflow practice hold that per-item review of machine-speed or high-volume output is physically impossible and collapses into nominal review, so oversight is constituted instead by engineered exception routing, sampling audits with measured error-catching performance, and circuit breakers with named halt authority.
- impact assessment — Communities disagree about whether candor or protection should govern the assessment document itself. A participatory accountability tradition requires the assessment to be published and candid: its point is surfacing harms the deploying institution cannot see and giving affected communities a voice, so a form completed internally is a failed assessment however legally sufficient. Legal counselling practice operationalizes the same instrument as evidence drafted under adversarial constraint: producible to authorities, discoverable in litigation, and quotable back as an admission of every risk it names, so the candid analysis is structured under privilege and the disclosable document frames each risk against its mitigation.
- machine learning — Communities want opposite things from a deployed model's capacity to change. Safety-regulated engineering practice treats a learned model as an artifact to be locked and qualified like tooling: frozen versions with defined operating envelopes, requalification gates on any change, locked-or-adaptive status declared for accreditation, and continuously learning systems excluded from release and safety functions because a self-modifying artifact defeats change-control logic. Product-technology practice treats ongoing change as machine learning's defining virtue: the model exists inside a staffed lifecycle of monitoring and retraining, and a model that is not retrained rots silently as the world drifts, so currency rather than fixity is the mark of sound deployment.
- metadata — Communities draw the line between metadata and protected content in incompatible places, with legal consequences attached. Signals-intelligence practice operationalizes metadata as communications externals, who contacted whom, when, from where, over what channel, a category with legal force because collection authorities have historically treated externals as less protected than content, even as tradecraft exploits externals at scale for contact chaining and pattern-of-life reconstruction. Civil-liberties practice around government systems denies that the boundary holds at all: under data-protection law any information relating to an identifiable person is personal data whatever its technical layer, so metadata retention and analysis programs are assessed as surveillance measures, not administrative documentation.
- missing data — Communities read the same absence of data under opposite default presumptions. Statistical research and official-statistics practice starts from mechanism classification: absence is analyzable under stated assumptions, missing completely at random, at random, or not at random, which license documented repairs such as weighting and multiple imputation, disciplined by sensitivity analysis and declared nonresponse rates. Intelligence and litigation practice inverts that default: a gap is presumed informative, and possibly authored, until shown otherwise, whether an adversary's denial and deception or a party's spoliation, so gaps are never imputed but investigated, stated in the product, and litigated, with silence about a gap itself treated as the failure.
- model — Communities disagree about what a model must offer before its outputs may be relied on. In litigation practice, a model is operationalized through the properties an opponent will attack — training inputs, assumptions, known error rate, validation method — and an artifact whose error rate and validation cannot be stated is unusable whatever its benchmark performance elsewhere; defense accreditation doctrine similarly confines reliance to a verified, validated envelope, treating outputs beyond it as unaccredited conjecture. Marketing data science instead operationalizes a model by its decision role — it moves budget or selects customers — and routinely relies on platform-side black boxes that practitioners can neither inspect nor retrain, accepting opacity as the price of operational yield. The disagreement is about what reliance should be conditioned on: examinability under challenge, or demonstrated decision utility.
- open data — Communities disagree about what confers openness on material. In research and software practice, open is a licence property: a dataset is open only when its terms — CC0, CC-BY, an approved open licence — affirmatively grant reuse, share-alike and text-and-data-mining reservations are binding engineering constraints, and merely accessible or scraped material of unclear status is routed to legal review before any use. In intelligence practice, open data names open-source material: whatever can be observed or collected without privileged access — social media, ship and flight trackers, commercial imagery, corporate registries, even leaked datasets — is open take, and the operative discipline is provenance grading and verification against adversarial seeding, not licence clearance. The same scraped corpus is unusable pending a licence determination on one reading and routine collection material on the other.
- overfitting — Technical communities bound overfitting at the generalization gap: a model or analysis pipeline is overfitted when apparent performance exceeds what out-of-sample, out-of-time, or untouched-holdout evaluation sustains, and the remedies are evaluation discipline, penalization, and pre-commitment. The critical reading in algorithmic government extends the concept to models fitted to the enforcement process that generated their labels: a risk model can pass every holdout check — because the holdout records carry the same historical scrutiny patterns — and still be, on this reading, overfitted to who was previously investigated rather than to the conduct of interest. The technical test therefore clears models the critical test condemns, and the two communities are measuring different failures under one word: sample-versus-population optimism on one side, reconstruction of past enforcement priorities on the other.
- privacy — Communities disagree about what discharges privacy where monitoring targets people who depend on the monitoring party. One school, rooted in platform and marketing engineering, operationalizes privacy as faithful individual-control machinery: consent states propagated and enforced end-to-end through tags and vendors, purpose labels on datasets, deletion as an SLO — privacy exists when the plumbing verifiably honors each person's choices. The other school, rooted in monitored work, holds that where being watched is a condition of getting or keeping work, individual consent cannot do the legitimating: privacy must be operationalized as substantive limits on the monitoring itself — proportionality caps, purpose bans such as no performance evaluation from telemetry, co-determined works agreements — because a dependent worker's choice is void. Applied to the same tracking system, one operationalization ships it behind a compliant consent flow; the other forbids it regardless of any consent collected.
- profiling — Communities draw the category's boundary by different tests. Financial-crime operations hold that a behavioral baseline — an expected-activity envelope of volumes, counterparties, and geographies against which live transactions are scored — is not an evaluative judgment of the person but a statistical instrument, whose governance is alert precision, detection coverage, and investigator capacity rather than the profiling apparatus. Data-protection counsel and plant-level practice classify by function instead: any automated processing that evaluates personal aspects — reliability, performance, behavior — is profiling whatever its operator calls it, so a dashboard ranking technicians by repair time or driver scores labeled safety coaching fall inside the category and trigger transparency, DPIA, objection, and co-determination duties, while a system that merely logs stays outside. The envelope reading and the function reading assign the same pipelines to opposite sides of the line.
- reproducibility — Communities disagree about what evidence establishes that a result is reproducible. Administrative, legal, and settlement practice demand exact reconstruction of the particular figure or decision: the challenged decision re-derived from inputs as they stood at decision time under the rule version then in force, the forensic extraction repeated hash-for-hash by a second examiner, the on-time figure regenerated from the same agreed events — anything less leaves the claim indefensible before a tribunal, an opposing expert, or a counterparty. ML engineering, facing nondeterministic GPU kernels and vendor-hosted models that change under the caller, operationalizes reproducibility in explicit tiers: rebuild the artifact, match metrics within tolerance, or regenerate an equivalent model from the pipeline — treating bitwise identity as often unattainable and specifying what must match instead. Each community reads the other's standard as either impossible or insufficient for the same deployed system.
- risk — Engineering safety practice operationalizes risk as a computed magnitude: FMEA and machine-safety files rate probability and severity per failure mode, acceptability follows from whether the computed residual meets pre-set criteria, and the AI Act's probability-times-severity definition reads there as a restatement of what the risk assessment file already does. Legal practice operationalizes AI-regulation risk as a categorical classification: a system's tier follows from its intended purpose and context, attaches before any quantitative estimate exists, and determines the obligations that bind. Each community treats the other's object as, at most, an input to its own.
- robustness — Communities disagree over whether deliberate adversaries fall inside robustness's boundary. Engineering and agri-environmental practice, in the robust-design tradition, bound the concept at uncontrolled but unintentional variation: noise factors, contrasting seasons, sensor degradation, demonstrated by deliberately varying those factors in designed experiments across the natural envelope. Defense evaluation and litigation practice hold that robustness is constituted by performance against a motivated opponent probing the weakest case, from camouflage, jamming, and data poisoning to the cross-examiner's hypothetical, so a system characterized only under benign or naturally varied conditions is not robust at all, whatever its tested noise envelope shows.
- sampling — Communities disagree about what licenses reliance on a sample. Statistical and engineering-data practice holds that the warrant lies in the design itself: probability selection with known inclusion probabilities, documented frames, and propagated reweighting, so a sound, well-documented sample licenses inference whoever drew it and whenever. Litigation practice holds that a sample's force as evidence depends on its protocol being agreed or court-ordered before selection, with sizes, method, and acceptance thresholds fixed in advance between adverse parties, because a unilaterally drawn, after-the-fact sample persuades no court however sound its statistical design.
- surveillance — Communities draw surveillance's boundary by different tests. Asset-protection practice bounds it by declared object and purpose: monitoring aimed at cargo, routes, machines, and shipped products is a protective discipline whose intensity is a quality attribute, kept deliberately separate from watching people. Worker-side practice bounds it by reconstructive capability: because the parcel's journey and the worker's day are one data stream, and machine telemetry sits one join away from a person's pace and breaks, monitoring counts as surveillance of people whenever it could reconstruct an individual's behavior if queried, whatever its declared object.
- synthetic data — Communities that all document the sim-to-real gap disagree about what simulation-derived data may evidence. Defense test practice admits accredited synthetic scenarios into the acceptance case itself: rare and dangerous conditions that live trials cannot safely or affordably produce are injected synthetically, with mixing ratios and residual gap documented, and accreditation rests partly on that evidence. Quality-engineering and research practice hold the opposite rule: synthetic data may train models, augment scarce classes, and validate machinery, but qualification evidence and claims about the world must be established on real parts, real lines, and real data.
- transparency — Communities that agree transparency exceeds a bare AI-involvement label disagree about what discharges it. Builder and safety-engineering practice operationalizes transparency as a professional-grade artifact chain: model cards, instructions for use, and the technical file, complete enough for deployers, integrators, and authorities to interpret and audit the system. Education and agri-administrative practice operationalize it as an account delivered to the affected layperson: a grade explicable to a parent across a table, a flag presented with parcel-level evidence a farmer can contest, holding that a result the affected person cannot understand fails the standard whatever documentation exists.
- uncertainty — Communities impose incompatible forms on the same uncertainty. Quantitative reporting traditions require the number: propagated intervals, error budgets, and calibrated probabilities, because only a quantified statement can be tested against outcomes, and their institutions have formally repudiated compressing it into coarser categories. Intelligence and legal practice require the register: standardized estimative language and opinion-ladder terms whose defined institutional weight carries reliance, penalty protection, and a separate signal of evidentiary quality, and they treat a naked model percentage as false precision the sourcing cannot support, to be re-expressed into the controlled vocabulary or refused.
- uncertainty — Where AI components enter inspection and conformity roles, communities disagree about what counts as an uncertainty statement. Metrological practice holds it is a traceable budget: contributions from repeatability, reference standards, and environment, assembled along a traceability chain and feeding guard-band decision rules, so a model's confidence score, calibrated or not, is not a measurement uncertainty and cannot populate that calculation. ML engineering practice holds that a demonstrably calibrated score, monitored in production like any reliability property, is precisely an operational uncertainty quantity fit to threshold and route decisions, and that calibration testing is what makes it so.
- validation — Within the fitness-for-use tradition, communities disagree about what validation must establish. Operational-performance practice validates by confronting predictions with realized events under live conditions: hit-rates within tolerance per corridor and season, staged gates from offline evaluation through shadow and canary deployment, so demonstrated predictive performance in deployment conditions is what 'validated' asserts. Educational measurement validates an interpretation: evidence that scores mean what the decisions built on them assume, spanning content, internal structure, external relations, and consequences for the people scored, with accuracy testing subordinated to that argument, so an accurate predictor can remain invalid for a particular use.