missing data

Absent values and the mechanisms producing them (MCAR/MAR/MNAR); also an equity signal.

Meanings by sector

Creative Industries

In data journalism, missing data is a finding: gaps in official records — deaths no agency counts, categories no form collects, jurisdictions that fail to report — signal where institutions do not look, and closing the gap becomes the story. Newsroom practice operationalizes this by auditing what a dataset should contain against what it does contain, building crowdsourced or scraped counts where records are absent, and disclosing coverage limits in the published piece so audiences can judge what the numbers omit.

In practice: Audit official datasets for what they should contain but do not, report the absence itself as substantive news, and disclose coverage gaps in what you publish.

OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)

Financial Services

For credit-model developers, missing data is a documented preprocessing object with regulatory consequences: absent bureau records define thin-file applicants, unanswered application fields become explicit 'missing' categories or imputed values, and each treatment must be recorded because it changes score distributions and adverse-action reasoning. Model-risk expectations require developers to demonstrate that data are suitable and adjustments justified, so missingness handling appears in development documentation with its rationale and its tested effect on performance across segments — an engineering decision with an audit trail, not an ad-hoc cleaning step.

In practice: Document the treatment of every missing-value pattern in model development, test its effect on segment-level performance, and justify the chosen treatment to independent validators.

OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)

Healthcare

In clinical-trials statistics, missing data is handled as a threat to valid estimation, classified by mechanism: values missing completely at random, missing at random given observed covariates, or missing not at random. The classification determines the licensed repair — complete-case analysis, multiple imputation, or mechanism-explicit models — and since the estimand framework of ICH E9(R1), regulators expect the assumed mechanism to be stated against the treatment-effect question the trial claims to answer and stress-tested with sensitivity analyses, because dropout is rarely random with respect to outcome. A team reporting an effect without probing its missingness assumptions has not finished the analysis.

In practice: State the assumed missingness mechanism for each incomplete variable, choose imputation or modelling methods licensed by that assumption, and report sensitivity analyses under plausible alternatives.

OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)

Healthcare

In health-equity work, missing data is read as a trace of who the health system fails to see: unrecorded ethnicity fields, absent lab values for patients who could not attend, and conditions undiagnosed for lack of access are structured absences, not random noise. Imputing over them can launder inequity into apparently complete datasets — models trained on such records learn the system's blind spots as if they were patient properties. The operational response is measurement reform and disaggregated missingness reporting, not statistical repair alone.

In practice: Report missingness rates disaggregated by population group, treat structured absence as evidence about access and recording practice, and escalate gaps that statistical repair would conceal.

OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)

Public Administration

In official-statistics production, missing data is managed as nonresponse within a quality framework: unit and item nonresponse rates are standard quality indicators, imputation and weighting adjustments follow documented and regularly reviewed procedures, and imputed values are flagged so users can separate collected from constructed figures. The European Statistics Code of Practice requires sound methodology and systematic quality reporting, so missingness is neither hidden nor improvised — it enters the release as declared nonresponse rates, documented adjustment methods, and bias assessments such as post-enumeration studies.

In practice: Monitor and publish nonresponse rates, apply documented imputation and weighting procedures, flag constructed values, and assess nonresponse bias against independent benchmarks.

OmniGloss seed synthesis, 2026 (machine-drafted, pending expert validation)

Documented disagreement

The communities disagree about what kind of object missing data is. Trial statisticians and model developers treat it as a property of a dataset: classify the mechanism, apply a licensed repair, document the effect, and the analysis can proceed. Equity practitioners and data journalists treat it as a property of the measurement system: absence records who institutions fail to see, so the correct response is to change what gets measured and to report the gap itself — repairing the dataset without naming the absence hides the finding.

Machine-readable version (JSON-LD)