Health Data Lab: Never Overwrite Raw Data — Why Every Survey Needs a Cleaning Log

Health Data Lab

A survey team returns from the field with 1,211 interviews. Someone opens the export in Excel, deletes the rows that look wrong, fixes a few numbers, and saves over the file. Three months later a partner asks: "Why is Penta3 coverage 73.6%? Which records did you remove, and why?" Nobody can answer. The original data is gone.

This is the most common, and most avoidable, failure in health data management. This week's tool shows what it costs, and how to prevent it.

Open the app Cleaning-log template Code

The app runs on a free server that sleeps when idle; the first visit can take about a minute. All data in the app is synthetic.

About the Health Data Lab series

Every week I take one practice from health-sector data management (survey design, cleaning, integration, indicators, data protection) and build a small, working tool that shows why it matters for decisions. Each tool comes with synthetic data you can explore, and space to try your own.

The rule: never overwrite raw data

clean data = raw data + a cleaning log. The raw export is kept exactly as received. Every change (a corrected value, a removed duplicate, an excluded interview) is written in a log with its reason, who made it and when. The clean dataset is rebuilt from the two, so anyone can reproduce and audit every number.

Five habits make this work in practice:

  1. Keep raw untouched. Save the export, dated, in a raw/ folder and never edit it.
  2. Log every change: record, variable, old value, new value, reason, action, field verification, who, date.
  3. Flag, don't silently fix. A value that is suspicious but possible (a very short interview, a child with Penta3 but no Penta1) goes back to the field team for verification. It is not deleted.
  4. Exclude, don't delete. Records without consent or suspected fabrications are excluded from analysis but stay in the raw file, so the decision can be reviewed.
  5. Keep personal data out of analysis copies. Names and phone numbers stay in raw only (data minimisation under the Nigeria Data Protection Act 2023).

The tool: Raw → Ready

Raw → Ready is a cleaning workbench built around that rule. It comes loaded with a synthetic KoboToolbox export from a child-health household survey in a fictional state (1,211 submissions from 12 enumerators), seeded with the problems real fieldwork produces: re-sent submissions, rushed interviews, night-time entries, MUAC typed in millimetres, impossible ages, missing consent, and two enumerators filling forms from a single spot.

1. Readiness: is this data fit for decisions?

The Readiness view on arrival: not ready for analysis, with 401 issues found by 14 standard checks

Figure 1: The Readiness view on arrival: not ready for analysis, with 401 issues found by 14 standard checks

The first screen answers the manager's question in one line: not ready, ready with caveats, or ready. Below it, the record counts, open issues, and records awaiting field verification. The verdict only improves as decisions are logged.

2. What uncleaned data would have told you

The same survey computed on the raw export and after cleaning

Figure 2: The same survey computed on the raw export and after cleaning

This is the heart of the tool. Each headline indicator is calculated twice: on the raw export, and on the cleaned data. In this survey:

IndicatorRaw exportAfter cleaningWhat raw data would have said
Penta3 coverage (12–23 months)73.6%70.1%Coverage looks 3.5 points better than it is
Penta1→3 dropout16.2%18.8%The dropout problem looks 2.6 points smaller than it is
Global acute malnutrition (MUAC)11.5%12.7%1.2 points of malnutrition are hidden

Why? The fabricated interviews report every child as vaccinated and well nourished. They inflate coverage and dilute malnutrition. A programme manager reading the raw numbers would cut defaulter tracing and nutrition screening in exactly the places that need them.

3. Review flags: a decision for every issue

Review flags: issue types with the protocol's recommended action, and every flagged record

Figure 3: Review flags: issue types with the protocol's recommended action, and every flagged record

Each issue comes with the recommended action from a standard cleaning protocol: correct a clear typo (MUAC 125 → 12.5 cm), set to missing an impossible value (a child aged 120 months), exclude records without consent or suspected desk interviews, or send to the field for a call-back when a value is suspicious but possible. The data manager accepts the recommendation, chooses another action with a reason, or applies the protocol to a whole issue type in one click.

4. The cleaning log: an audit trail nobody can quietly edit

The append-only cleaning log: every change, its reason, who made it and when

Figure 4: The append-only cleaning log: every change, its reason, who made it and when

Every decision becomes a row: record, variable, old value, new value, reason, action, field verification, reviewer and time. The log is append-only: undoing a decision adds a reversal entry instead of erasing the original. The columns match the standard cleaning-log template, which you can download from this post.

5. Enumerators: fix problems while teams are still in the field

Supervision scorecard: interview length, repeated GPS points, night-time entries and MUAC rounding per enumerator

Figure 5: Supervision scorecard: interview length, repeated GPS points, night-time entries and MUAC rounding per enumerator

Data quality problems are usually people problems, and fixable while fieldwork is ongoing. The scorecard shows who needs support: two enumerators with interviews of about 6 minutes against a 25-minute team median, dozens of interviews at an identical GPS point, or MUAC readings always rounded to .0 or .5 (a measurement-training need). Each comes with a suggested follow-up. Run it daily, and share it as support, not punishment.

6. Data & privacy: proof, and protection

Data & privacy: SHA-256 fingerprint of the raw file, recognised fields, and identifier-free exports

Figure 6: Data & privacy: SHA-256 fingerprint of the raw file, recognised fields, and identifier-free exports

When an export arrives, the tool records its SHA-256 fingerprint, a unique code computed from the file's contents. If someone later questions the analysis, hash your stored copy: the same fingerprint proves it is the same raw data. Names and phone numbers are detected and never reach the analysis copy or any export, and household GPS is left out of shared files by default.

After cleaning

After every flag has a logged decision: ready with caveats, pending field verification

Figure 7: After every flag has a logged decision: ready with caveats, pending field verification

Once every flag has a logged decision, the verdict moves to ready with caveats: the remaining call-backs are reported as limitations until the field team confirms them. One click exports the audit pack (clean data, cleaning log, flags, enumerator scorecard and indicator impact in one Excel file), a shareable clean CSV, and a one-page summary for managers and partners.

Use it in your next survey

Open the app Cleaning-log template Code

The app runs on a free server that sleeps when idle; the first visit can take about a minute. All data in the app is synthetic.

Next week

Week 2: Lint your XLSForm before you deploy. Most cleaning problems start in the form. Next week's tool checks a KoboToolbox form for broken skip logic, missing choice lists and absent constraints before a single interview is collected.

All data in this post and the app is synthetic. Upload identifiable health data to any online tool only with your organisation's approval, or de-identify it first.