Where experts disagree
- accountability — Communities disagree on where accountability's boundary lies. One school, rooted in GDPR Article 5(2) and audit practice, bounds it at demonstrable conformity: an actor is accountable when complete, retrievable evidence shows obligations were met and decisions can be reconstructed. Another school, rooted in administrative justice and frontline practice, holds that records alone cannot constitute accountability: the concept extends to a human who explains the decision to the affected person, faces consequences, and can provide remedy. Both camps value documentation; they dispute whether it is the whole of the concept or merely its precondition.
- accountability — In AI-assisted clinical care, one position holds that the treating clinician must remain the single, undiluted locus of accountability for every decision, on the ground that patient safety depends on one professional who cannot point elsewhere. The opposing position holds that accountability must be distributed across the value chain, spanning manufacturer, deploying organization, and clinician, in proportion to actual control, arguing that loading residual responsibility onto the frontline user is both unfair and unsafe because it shields the actors who control design, training data, and updates. The disagreement concerns what accountability arrangements should protect: the clarity of a single answerable professional, or the alignment of answerability with control.
- accuracy — Communities draw the boundary of 'accuracy' around different objects. Model-risk and machine-learning practice defines it as a distributional property of a system over a population — a rate monitored within tolerance bands, under which individual errors are expected and managed. Data-protection practice defines it as a property of each stored record about each person, enforceable individually through rectification rights. A system can satisfy one reading while violating the other.
- aggregation — Communities read the evidence on aggregation's protective power differently. Health-reporting practice treats thresholded aggregate tables as privacy-safe by construction, pointing to decades of routine releases without demonstrated harm. Statistical-disclosure specialists point to differencing and database-reconstruction attacks showing that aggregation without formal guarantees can be reversed at scale, and conclude that 'aggregate' cannot be equated with 'anonymous'.
- algorithm — Technical communities reserve 'algorithm' for a precisely specified, finite input-output procedure, kept distinct from the trained model that supplies parameters and from the deployed system around both. Creators, affected citizens, and platform-accountability reviewers use 'algorithm' for the entire sociotechnical decision or ranking system — code, models, tuning, policy, and operators — because only that whole is observable, governable, or contestable from where they stand. Both usages are internally coherent and serve real work in their communities, and the clash surfaces even within a single sector: feed engineers decompose exactly the pipeline that platform auditors insist on treating as one accountable unit.
- algorithm — Governance communities disagree on whether deterministic, human-authored decision rules belong inside formal algorithmic governance. Model-risk practice attaches its full validation apparatus to statistical and machine-learning estimation and leaves rule engines to lighter IT change control, a boundary echoed by the AI Act's exclusion of systems executing solely human-defined rules. Equity auditors and public-sector overseers instead scope governance by consequence, citing deterministic procedures — welfare debt formulas, exam standardization, race-corrected clinical equations — whose documented harms matched or exceeded those of learned models.
- alignment — One set of communities scopes alignment as a property of the model that can be specified, measured, and certified before deployment — evaluation suites, refusal metrics, adversarial-testing thresholds, auditable risk-mitigation artifacts — while another scopes it as an ongoing socio-political question about whose values govern the system, which no pre-deployment measurement can close because those values are contested and legitimately change through democratic and legal process. The same word names a test result for one side and a governance process for the other.
- anonymization — Communities draw the boundary of anonymized data in incompatible places. One position, anchored in the Article 29 Working Party's opinion, holds data anonymized only when re-identification is prevented against all means reasonably likely to be used, irreversibly and independent of context, with singling out, linkability, and inference all defeated. The other, anchored in risk-based regulatory guidance and statistical-disclosure practice, holds anonymization to be a contextual judgment: data counts as anonymous when residual risk is remote given the specific environment and controls around it, since zero risk is unattainable for any data that retains utility. The same boundary dispute recurs wherever pseudonymous identifiers circulate commercially: European regulators class hashed emails and rotating device identifiers as personal data, while industry contracts routinely label them anonymous.
- anonymization — Communities read the empirical record of re-identification attacks in opposite ways. Privacy engineers cite demonstrations such as the Netflix Prize linkage, credit-card metadata unicity, and genomic surname inference as proof that record-level anonymization of rich data fails structurally, so only mechanisms with formal guarantees deserve the name. Statistical-disclosure and health de-identification practitioners read the same attacks as strikes against naive public releases, pointing to decades of controlled official-statistics outputs without demonstrated re-identification harm as evidence that well-governed traditional methods remain sound.
- artificial intelligence — Communities disagree about whether 'artificial intelligence' names a bounded category with membership criteria or a region on a continuum of model complexity. EU legal practice, following AI Act Article 3(1), requires a yes/no classification because obligations attach to membership; model-risk practice in finance holds that inference-producing systems differ only in degree, and that a bright line between 'AI' and conventional statistical models is arbitrary and unstable as techniques evolve.
- audit — Communities disagree over what qualifies an examination as an 'audit'. The assurance tradition (internal audit, supreme audit institutions, certification bodies) reserves the term for independent examinations against defined criteria under a mandate, with auditee cooperation, professional standards and formal reporting duties. The investigative tradition in journalism and research applies it to adversarial external probes of system behavior conducted without access, agreed criteria, or the operator's consent, grounding authority in method transparency and public interest instead.
- automated decision-making — Communities disagree about where in an automated pipeline the 'decision' occurs and how much human involvement removes a decision from the automated category. Established lending compliance attaches Art. 22 duties at the final act on the customer, treating upstream scores as inputs and a trained human checkpoint as sufficient de-automation. The post-SCHUFA data-protection reading and administrative-law oversight practice instead look to determinative effect: a score or system output that is systematically followed is itself the decision, and pro-forma human review does not change its automated character.
- benchmark — Communities disagree about what a benchmark score is evidence of. Regulatory and procurement actors increasingly treat standardized benchmark results as operative evidence of system capability, sufficient to trigger legal classification, obligations, or acceptance decisions. Clinical researchers and critical observers counter that benchmark scores measure performance on a narrow, often contaminated proxy task and warrant no presumption about real-world capability without independent, context-specific validation. The same score is decisive evidence in one room and a marketing claim in the next.
- bias — Communities that build and communities that scrutinize public risk-scoring tools operationalize an unbiased score through incompatible criteria. Developers and vendors center calibration within groups: a given score must correspond to the same outcome probability for every group. Journalistic, equality-body, and civil-rights auditors center error-rate parity: false-positive and false-negative burdens must not fall unequally on protected groups. Impossibility results (Chouldechova 2017; Kleinberg, Mullainathan and Raghavan 2016) show that when base rates differ no score can satisfy both, so choosing a criterion is a normative choice about which errors matter, not a technical refinement.
- bias — Quantitative traditions — estimation theory, actuarial science, official statistics — draw the concept's boundary at systematic deviation of an estimate from a true value: directional, measurable, sometimes deliberately introduced (regularization, shrinkage), and not intrinsically about people or harm. Critical, audit-oriented communities draw it around patterned disadvantage produced by sociotechnical systems: bias is inseparable from power, institutions, and injury, and a purely numerical reading misses the object. NIST SP 1270's systemic/human/statistical taxonomy institutionalizes the broad boundary; estimation-theoretic practice institutionalizes the narrow one.
- big data — Engineering and oversight communities bound the concept incompatibly. For financial-services engineers, data are big when volume and velocity force architectural choices — streaming pipelines, distributed storage, latency budgets — and the label carries no normative charge. For courts and watchdogs reviewing government data use, bigness is measured by linkage power: a modest dataset becomes big when joining it to other registries makes populations legible and scoreable in new ways, triggering proportionality and contestability requirements. Each criterion classifies the other side's paradigm cases as not big at all: a small linked welfare-scoring scheme has no engineering scale, while high-volume unlinked telemetry raises no linkage concern.
- data governance — Communities disagree about what the term denotes. In finance and enterprise practice, data governance is an internal control discipline over data assets — named owners, lineage, quality thresholds — evidenced to auditors and supervisors. In health and public administration it names the institutional allocation of decision rights and answerability over data use, extending to patients and citizens; EU legislation (the Data Governance Act) has fixed this wider sense in law. Both meanings are current, and each side hears the other's usage as either too narrow or too vague.
- data literacy — Communities disagree about what data literacy is for and where it resides. Clinical and organizational framings treat it as an individual or workforce competence: skills that let a person interpret and use data and AI outputs correctly, deliverable through training and verifiable through records. Critical civic framings hold that this deficit model misplaces the burden: literacy is a collective, political capacity to question and reshape the data regimes that classify people, and training individuals to read outputs leaves the regimes themselves unexamined.
- data minimization — The communities disagree about when and how 'necessary' is established. Administrative lawyers, hospital data-protection officers, and ethics boards treat necessity as an ex-ante legal test that must be satisfied before any collection occurs: an item without a demonstrable basis is unlawful to gather, whatever its potential utility. Model developers treat necessity as an empirical property demonstrated during development: candidate data is explored under controls, and the minimized feature set is the output of ablation and pruning rather than a precondition of exploration.
- data provenance — Communities disagree about what data provenance covers. Engineering-oriented practice in healthcare and finance bounds provenance at the verifiable processing record: where data physically originated and which transformations produced its current form, judged by reconstructability. Rights- and consent-oriented practice — creative-sector chain-of-title and health-data access governance alike — bounds it at authorization: whose works, permissions, and consent scopes the data embodies, judged by whether uses stay inside that inherited envelope. Each side treats the other's criterion as outside the concept, so the same dataset can be provenance-complete and provenance-void simultaneously.
- data provenance — Communities that agree provenance documentation must exist disagree about whom it is for. Journalistic and public-accountability practice holds that provenance is only realized when origin can be stated to audiences, citizens, and courts, with secrecy as a per-element exception. Clinical stewardship and financial model governance hold that provenance records are confidential custody instruments for regulators and validators, since the trail itself exposes patient linkages or commercially sensitive sourcing. Both positions claim accountability while prescribing opposite disclosure defaults.
- data quality — Builder and analyst communities in healthcare and finance define data quality relationally, as fitness of data for a specified use, with checks and thresholds re-justified per task; integrity-focused audit communities in the same sectors define it as an intrinsic property of records and their production controls — attributability, traceability, reconciliation to authoritative sources — assessable and reportable without reference to any downstream use. Both usages are institutionally entrenched: fitness-for-use language dominates data-science and engineering standards, while inspection and control regimes (GxP data integrity, risk-data aggregation) codify the intrinsic reading.
- data quality — Statistical and corpus-building communities locate data quality in the aggregate product — reliable estimates, well-distributed corpora — and accept individual record errors that wash out at scale; caseworker, consumer-credit, and data-protection communities locate it in each individual record, because a decision about a specific person stands or falls on that person's data being correct. The two commitments assign opposite priorities to the same error: negligible to the aggregate, decisive to the individual, and each side's quality assurance leaves the other's core concern unmeasured.
- data sharing — Health-research communities and government data-protection practice want opposite defaults. Researchers treat sharing as a duty owed to participants and to science — non-sharing is waste that must be justified — and build repositories and access committees to make it routine. Public-sector legal practice treats every share as an exception demanding an identified gateway, purpose compatibility, and signed agreements before anything moves. Each default is principled: one maximizes value from data already collected, the other guards purpose limitation and citizens' expectations.
- discrimination — Actuarial and risk-modeling communities hold that differentiation is discriminatory only when it lacks statistical or actuarial justification: risk-based differentiation is the fair outcome, and forcing equal treatment of unequal risks creates cross-subsidy injustice. Legal and administrative communities hold that using protected grounds or their proxies is discrimination regardless of predictive validity: the wrong lies in the ground of differentiation itself, so actuarial soundness is no defense, as Test-Achats established for gender-based insurance pricing.
- explainability — Communities deploying AI in high-stakes settings disagree about whether current post-hoc explanation methods are faithful enough to support reliance on individual outputs. Frontline professional practice — clinicians reconciling case-level rationales with the chart, underwriters translating attribution-derived reason codes into adverse-action notices — treats case-level explanations as a workable and necessary ingredient of safe, accountable use, enabling users to check and override systems. The skeptical side cites fidelity failures — explanations insensitive to model internals, attributions that describe a surrogate rather than the deciding model, reason codes that reorder under retraining — and concludes that reliance should rest on rigorous validation or on inherently interpretable designs, with persuasive but unverified explanations regarded as a hazard rather than a safeguard.
- explainability — Sectors disagree about where explainability's boundary lies: whether it denotes a technical property of models realized by attribution methods and tooling, or the communicative and legal act of giving an affected person accurate, specific, contestable reasons for a decision. Builder communities bound the concept at the pipeline artifact — attributions computed, logged, and quality-checked — treating translation into reasons as someone else's task, while compliance, administrative-law, and advocacy communities bound it at the recipient: without legally and practically usable reasons there is no explainability, however sophisticated the tooling.
- fairness — Formal fairness criteria are mutually incompatible when base rates differ across groups: a risk score cannot in general be simultaneously calibrated within groups and equal in false-positive and false-negative rates across them. Credit-risk builders treat within-group calibration as the fairness baseline and log residual error-rate gaps as monitored differences; clinical and public-sector builders privilege error-rate parity because misclassification costs fall directly on patients and citizens. Each community's release gate would fail models the other would pass.
- fairness — Communities draw the boundary of fairness work at protected-attribute use in opposite places. For many European model-building teams, fairness and data-protection compliance mean excluding protected attributes and their proxies from collection and processing altogether. For health-equity and public-sector auditors, fairness work begins with collecting those same attributes under safeguards, because disparities cannot be measured, nor non-discrimination demonstrated, without disaggregated data. Each side treats the other's core practice as the problem itself.
- generative AI — Sectors disagree about when synthetic content must be disclosed. Public-integrity institutions hold that generated media should be machine-marked and labeled by default, because deception risk is structural and detection after distribution fails. Creative practice ties disclosure to deceptive intent and context: generative tools are legitimate craft, and blanket labeling stigmatizes lawful work while doing little against bad actors who will not comply. The AI Act institutionalizes the tension without resolving it — provider-side marking is unconditional while deployer-side labeling is relaxed for evidently artistic works.
- ground truth — Communities disagree about the epistemic status of ground truth itself. The operational-fixture position — dominant among model builders in healthcare and finance — treats the reference label set as fixed and definitive for training and scoring: disagreement among labelers is noise to be adjudicated away so optimization has a stable target. The constructed-reference position — held by clinical epistemologists, official statisticians, and equity auditors — insists every reference standard is itself a fallible, socially situated measurement with an error structure, and that treating it as truth launders those errors into every downstream metric.
- ground truth — Where the authoritative record and observed reality diverge, communities disagree about which one is ground truth. The legal-authority position — administrative decision-makers, and analogously model-risk governance in finance — holds that the designated register or approved outcome definition is definitive because legal certainty and mass administration require a single authoritative account, correctable only through due process. The empirical-verification position — caseworkers, auditors, journalists — holds that the record is evidence, not truth, and observed reality must be able to overturn it.
- hallucination — Communities shipping tightly scoped, retrieval-grounded applications read claim-level attribution evaluations as showing that fabrication can be engineered down to negligible levels for bounded tasks: constrain generation to an authoritative corpus, require citations, force abstention on low support, and measured unsupported-claim rates on curated benchmarks approach zero. Validation and clinical NLP communities read the same evaluation literature, together with the character of sampling-based decoding, as showing an irreducible floor: rates fall but never reach zero, curated benchmarks systematically undercount open-ended fabrication, and a certified zero is an artifact of the test set — so verification and monitoring controls remain permanently necessary parts of any deployment.
- hallucination — Within the creative sector, entertainment and game-development communities prize the generative behavior the term condemns: producing people, places, and events unconstrained by fact is the creative product, and hallucination misnames it as malfunction. Journalism and editorial-standards communities hold that any fabricated factual assertion in published matter is a professional breach whatever the tool, and object that the softer clinical-sounding term launders what their codes plainly call fabrication. The two camps do not disagree about what the systems do — they disagree about whether the word names a harm.
- human oversight — There is genuine disagreement about what establishes that human oversight exists. Compliance-oriented communities in finance and public administration operationalize oversight as demonstrable arrangements: designated competent reviewers, interface tools, override authority, and auditable logs — if the prescribed capabilities are designed in and documented, oversight is established. Clinically and critically oriented communities read the empirical record on automation bias, alert fatigue, and rubber-stamping as showing that such arrangements frequently fail to produce actual control, and therefore count oversight as existing only where humans are shown to detect, override, and change outcomes under realistic conditions.
- impact assessment — Communities disagree about what an impact assessment is for. In regulated-industry practice it is a compliance gate: an internal document that classifies risk, routes a change to the right controls and evidences diligence for supervisors, completed before approval and updated on change. A participatory tradition, rooted in environmental-assessment lineage and algorithmic-accountability scholarship, holds that assessments exist to surface harms visible only to affected people, so consultation, publication and ongoing reassessment are constitutive, not optional extras.
- inference — Communities draw the boundary of what the word inference names at three different places. Statistical and official-statistics practice reserves it for reasoning from data to population-level claims under explicit uncertainty; ML engineering and deployment practice uses it for the runtime execution of a trained model on individual cases; data-protection and administrative-law practice treats the inference as the derived personal datum itself — the new fact recorded about a person — whatever process produced it. Each usage is entrenched in standards, tooling, billing models, case law, and job descriptions, so no community treats its reading as metaphorical or secondary.
- informed consent — Communities in the creative industries disagree about what legitimates the use of creative works, voices, and likenesses in generative-AI development. Performer representatives, commissioning decision-makers, and many working creators hold that prior, specific, and typically compensated opt-in consent is required before material is used for training or replication. Model developers operationalize legitimacy as lawful access under text-and-data-mining exceptions combined with honoring machine-readable opt-outs, reserving affirmative consent for targeted replication of identifiable individuals. Both sides use the word consent for these incompatible mechanisms.
- informed consent — Sectors draw the boundary of what counts as informed consent in incompatible places. Financial-services audit practice treats consent as a documented compliance event: a well-formed, timestamped record satisfying legal elements settles the matter. Health-research ethics treats consent as an ongoing relational process in which signed forms are merely evidence, with comprehension and continued willingness as the real criteria. Critical public-sector stewards go further, holding that where refusal carries costs, consent is structurally unavailable regardless of record or process quality.
- interpretability — Communities disagree on where interpretability's boundary lies. One school holds it is an intrinsic property of transparent model classes whose parameters can be read directly, so post-hoc explanations of black boxes fall outside the concept. A second holds that a black-box model becomes interpretable once constraints, attribution methods, and stability tests characterize its behavior to a demonstrated standard; a behavioral variant of this school claims interpretability whenever output changes can be reliably predicted from defined changes to inputs, prompts, or model versions. The schools use the same term for different objects: the model itself versus the evidence assembled around it.
- interpretability — Communities disagree about whose understanding interpretability must serve. One position treats interpretability as satisfied when qualified professional audiences, such as validators, risk committees, hospital adoption committees, and clinician users, can understand and challenge the model, since they carry the control responsibility. The opposing position holds that interpretability is owed to the people affected by decisions: patients, citizens, and customers must be able to grasp and contest the grounds of an outcome, and expert-only understanding leaves the affected person's position unchanged.
- missing data — The communities disagree about what kind of object missing data is. Trial statisticians and model developers treat it as a property of a dataset: classify the mechanism, apply a licensed repair, document the effect, and the analysis can proceed. Equity practitioners and data journalists treat it as a property of the measurement system: absence records who institutions fail to see, so the correct response is to change what gets measured and to report the gap itself — repairing the dataset without naming the absence hides the finding.
- model — Governance and transparency communities define a model by decision influence: any quantitative or automated logic that turns inputs into estimates or decision-shaping outputs — including scorecards, spreadsheets, and hand-coded rules — so that oversight coverage is complete. Engineering communities define a model as the trained, versioned parameter artifact produced by a learning pipeline, categorically distinct from rules, thresholds, prompts, and surrounding code. Both usages are internally coherent but draw the term's boundary in incompatible places, and each community's inventories, registers, and controls presuppose its own boundary.
- model — Communities disagree about what evidence establishes that a model is adequate. The representational tradition, strong in official statistics and regulated quantitative modeling, requires that a model's assumptions about the data-generating process be stated, tested, and scientifically defensible; good fit alone is insufficient. The predictive tradition, strong in clinical and industrial machine learning, treats demonstrated out-of-sample performance on the intended population as the acceptance criterion, with mechanistic plausibility desirable but not required. The split mirrors Breiman's 'two cultures' of statistical modeling, and each side reads the other's evidence dossier as missing the point.
- open data — The sectors draw the boundary of 'open' incompatibly. Public administration operationalizes open data as a universal reuse right: anyone, any purpose, open licence, at most attribution — access controls disqualify the label. Health research and finance attach 'open' to governed access arrangements — controlled-access repositories and consented, authorization-gated APIs — where data are neither public nor freely reusable, and openness means a discoverable, rule-bound route in. Both usages are institutionally entrenched, one in open-government law and the Open Definition, the other in research-governance and open-banking regimes, so the same word licenses very different expectations about who may obtain and reuse the data.
- outlier — The communities draw the concept's boundary differently: for statistical producers an outlier is a data value threatening estimate quality, to be verified and contained by procedure; for fraud and clinical teams it is the signal of interest, the very thing detection exists to surface and investigate; for critics of administrative AI it is a person whose atypicality a system converts into suspicion or degraded performance. The same flagging operation is therefore quality control, detection success, or due-process harm depending on which reading holds.
- personal data — The communities disagree about where personal data ends. Hospital data-protection officers, reading Recital 26 conservatively, hold that pseudonymised or detailed clinical data remain personal as long as re-identification is conceivable by any party holding auxiliary information. Statistical-disclosure practice and health-data engineering instead define the boundary by measured residual risk: once quantified re-identification risk in a given release environment falls below an agreed, documented threshold, the data are handled as anonymous. Both sides claim fidelity to the same 'means reasonably likely' test.
- privacy — Communities disagree about when data about people has been made private enough to circulate. One position treats privacy obligations as discharged once a recognized anonymization or de-identification threshold is met and documented: the data ceases to be personal and leaves data-protection scope. The other holds that re-identification risk is continuous and never reaches zero — especially for rich clinical, genomic, longitudinal, or small-area data — so only quantified leakage bounds such as a differential-privacy budget, or adversarial attack evaluations, can substantiate a privacy claim; a threshold sign-off merely records an opinion about an unmeasured risk. Both positions read the same releases and the same re-identification literature and draw opposite conclusions about sufficiency.
- privacy — Communities disagree about privacy's normative weight when it collides with mandated or legally protected public-interest activity. Data-protection supervisors and civil-liberties stewards operationalize privacy as a fundamental right and structural check on institutional power whose every interference must be shown necessary and proportionate, with the burden on the intruding party. Editors and financial-crime officers operationalize privacy as one lawful interest among several, routinely and legitimately outweighed by freedom of expression in public-interest journalism or by statutory anti-money-laundering surveillance duties that customers cannot refuse. Each side regards its ordering as the legally and morally correct one, and each can cite binding law or documented harms in support, so the disagreement is about which values set the default, not about legal ignorance.
- profiling — Commercial personalization communities operationalize profiling as core value-creating infrastructure: better profiles mean better recommendations, experiences and yield, with law entering as a consent constraint on an otherwise legitimate activity. Rights-protective communities in public administration, courts and data-protection practice operationalize profiling as a presumptively dangerous exercise of classificatory power over people, tolerable only with explicit legal basis, demonstrated necessity and safeguards, and in some uses prohibited outright.
- prompt — Creative practitioners operationalize prompting as authorial craft — skilled, iterated labour that determines whether an output is usable and that merits credit, compensation, and protection as know-how — while copyright registrars and prevailing legal practice treat prompts as unprotected ideas or instructions to a machine, locating authorship only in demonstrable human expressive contribution beyond the prompt. The disagreement concerns where the boundary of authorship lies when expressive detail is produced by a model steered by human instruction, and it is sharpened rather than settled by every new registration decision.
- pseudonymization — The communities draw the boundary of personal data differently: research practice treats key-coded data as effectively de-identified for a recipient with no access to the key, grounding lighter handling of coded datasets, while the data-protection reading holds that data remain personal so long as anyone retains means to re-identify, because the test is whether the person is identifiable at all, not identifiable by the current holder. European litigation over coded data shared without keys has kept both readings live.
- risk — Regulatory product-safety regimes, including the EU AI Act, define risk exclusively as the combination of probability and severity of harm, so risk work means preventing and minimizing adverse outcomes. Enterprise risk-management traditions descending from ISO 31000, echoed in the NIST AI RMF and in commissioning practice in the creative economy, define risk as the effect of uncertainty on objectives, whose consequences can be negative or positive; there risk is deliberately taken and accepted, not only mitigated. Both usages are deeply institutionalized, and each community's documents presuppose its own scope without flagging it.
- risk — Quantitative modeling communities operationalize risk as a calibrated, testable number — predicted probabilities, expected loss, value-at-risk — because commensurability is what lets thresholds, prices, and capital consume it. Rights-focused stewards and affected creator and citizen communities counter that severity to fundamental rights and livelihoods is not commensurable: harms that concentrate on particular groups or destroy an individual practice cannot legitimately be netted against aggregate benefit, so risk must be assessed distributionally and sometimes treated as non-offsettable.
- robustness — Communities disagree about what robustness is a property of. Model-centric practitioners — clinical-ML and quantitative developers — operationalize it as a measurable attribute of the trained artifact: quantified degradation under perturbed, shifted, or stressed inputs, established by test batteries before release and versioned with the model. Governance and operations communities — production leads, public-sector auditors — operationalize it as an attribute of the whole deployed sociotechnical service: fallback workflows, monitoring, update governance, and the institutional capacity to detect and correct failure. Both use the same word for differently bounded referents, so a claim of robustness carries no fixed scope until the speaker's community is known.
- robustness — Communities that share the goal of robust systems disagree about what evidence demonstrates robustness. One school privileges designed worst-case examination before deployment: adversarial perturbation suites with certified bounds, and pre-specified adverse scenarios whose results gate approval. The opposing school holds that designed tests cannot anticipate real operating conditions — local case mix, policy changes, slow drift — so only ecological evidence counts: local prospective validation on the deploying institution's own population, and production monitoring such as drift detection and canary cohorts accumulated in situ over time.
- sampling — Communities disagree over what licenses inference from a sample to a population. Official statistics holds that only designed probability sampling with known inclusion probabilities yields defensible population estimates, and that massive found datasets remain biased regardless of size. Machine-learning practice, exemplified in healthcare AI, treats large convenience cohorts — corrected, externally validated, and checked for subgroup performance after the fact — as adequate warrant for deployment claims.
- surveillance — Public-health practice treats systematic population monitoring as a core protective duty, evaluated by how completely and quickly it detects health threats, while fundamental-rights practice treats the same monitoring as a presumptive interference with private life that is unlawful unless a necessity and proportionality case is made in advance. The word therefore carries opposite default valences: for one community more surveillance is prima facie better; for the other, less.
- synthetic data — Practitioner communities read the privacy evidence about data generation differently. Bank technology and vendor-management practice treats well-generated synthetic records as categorically non-personal: no record maps to a customer, so data-protection constraints fall away. Health-data researchers and statistical offices counter with empirical results — generative models memorize outliers, membership-inference attacks succeed against synthetic releases, and rare individuals can be reconstructed — so anonymity is a measured, per-release property rather than a consequence of the generation method itself.
- training data — Creative-industry communities disagree about the conditions under which existing works may legitimately become training data. Developer practice treats works that are lawfully accessible online as usable corpus material provided machine-readable opt-outs are honored, relying on the DSM Directive's text-and-data-mining exception with its Art. 4(3) rights-reservation mechanism and on analogous fair-use arguments elsewhere. Working creatives and the organizations that commission or license their work treat training as an act of exploitation requiring consent, licensing, credit and compensation per work, regardless of accessibility. Both sides agree the works were used; they disagree about what legitimate use requires.
- training data — Communities draw the boundary of 'training data' differently. Regulatory and audit practice follows the AI Act's definition — data used to fit a model's learnable parameters — keeping training, validation and testing data legally distinct so duties can attach precisely. Engineering practice in production settings bounds the term functionally, as the whole evolving data estate that shapes deployed model behavior: fitted partitions, tuning splits, retraining feeds and feedback loops. The statute itself concedes the line needs active maintenance — Art. 3(31) allows the validation data set to be a part of the training data set — and each side regards the other's boundary as unusable for its work: too broad to audit, or too narrow to govern.
- transparency — Communities divide over what transparency is transparency of. One cluster operationalizes it as disclosure of use: the affected person or audience must be told that AI generated or contributed to what they encounter, and transparency is achieved by an effective label or notice at the point of encounter. Another cluster operationalizes it as intelligibility of the system: documentation of intended use, performance, limitations, and working logic sufficient for a user or reviewer to interpret outputs — on this view a bare notice that AI was used achieves nothing. Both call their object transparency, and each treats the other's object as, at best, a component of the real thing.
- transparency — Parties agree that algorithmic systems must be transparent to someone but disagree about how far disclosure must extend beyond supervisors and internal reviewers. Institutional risk owners in banking and government defend graded disclosure — full access for validators, supervisors, and courts; calibrated reasons for affected individuals; restricted public detail — arguing that published logic invites gaming, evasion, and loss of protected assets. Consumer advocates, civil-society watchdogs, and some courts counter that transparency which never reaches the affected public is mere record-keeping, and that gaming and trade-secret claims must be proven rather than presumed.
- trustworthiness — Assurance-oriented communities (clinical-AI engineering, financial model validation) draw trustworthiness as a bounded, verifiable property of a system: a bundle of characteristics that can be specified, measured, evidenced, and certified. Relational communities (patients, media audiences and the editors answerable to them, citizens and agency leaders, retail-finance executives facing public scrutiny) draw the boundary around institutions and relationships: trustworthiness is standing earned over time through honest, answerable conduct, which certification evidence can support but cannot constitute. The line runs through sectors as much as between them — finance houses both a proceduralized assurance stance and a public-defensibility stance — so each side hears the other as either naive or evasive while using the same word for differently bounded objects.
- trustworthiness — Within healthcare, and with echoes in other regulated sectors, communities read the same validation evidence differently. Procurement decision-makers treat regulatory clearance and pivotal-trial results as sufficient warrant to deploy a system and direct staff reliance. Governance and audit stewards hold that such ex-ante evidence decays under distribution shift and site variation, so trustworthiness exists only while local post-deployment monitoring actively re-verifies it. The disagreement concerns what evidence warrants trust at a specific site and time, not what trustworthiness is for.
- uncertainty — The disagreement concerns what uncertainty information is owed to whom. Statistical producers hold that quantified uncertainty must accompany every published figure because suppressing it misleads users and corrodes trust. Clinical and operational decision contexts hold that uncertainty is only meaningful once translated into action thresholds, and that displaying raw probabilistic detail impairs timely, safe decisions. Both sides claim fidelity to the user; they want different goods — complete disclosure versus decidable workflows.
- validation — Across sectors, 'validation' names two different things. In machine-learning development practice — echoed in the EU AI Act's definition of validation data — it denotes a data split and tuning-and-testing stage inside model building, performed by the building team. In assurance regimes such as SR 11-7 model-risk management and medical-device regulation — echoed in the IEEE/ISO definition NIST's glossary carries — it denotes an independent, evidence-based confirmation that a system is fit for a specific intended use, producing an approval status with conditions. Both usages are entrenched in their communities, both are backed by authoritative texts, and neither reduces to the other.
- validation — Within healthcare AI, communities disagree on the evidence needed before a clinical model may be called validated. Model developers and much of the publishing research community treat strong retrospective external validation — stable discrimination and calibration on independently collected cohorts — as sufficient. Hospital deployment leaders and a growing clinical-safety school hold that only prospective, site-specific evaluation in the live workflow, on local data feeds and case mix, establishes fitness for use, because retrospective transportability has repeatedly failed to predict deployed performance.