📘 Reference guide

DQA: checking the quality of your MEAL data

A DQA (Data Quality Assessment) checks that a figure reported to a donor really means what it claims to mean. It is not looking for blame: it walks back up the chain, from the indicator in the report to the data collected in the field, and looks at where it breaks.

The DQA, in one sentence

A Data Quality Assessment is a structured exercise checking the quality, the consistency and the traceability of reported data. You start from a figure in the report, and work back down: where does it come from, who entered it, on what basis, and do you find the same value at every level?

Worth remembering: a DQA is not a financial audit nor a hunt for culprits. It is a diagnosis of a chain. The vast majority of the gaps it finds come from an ambiguous indicator definition or a file copied by hand, not from bad faith.

The five dimensions of quality

Institutional donors, USAID first among them, structure the DQA around five dimensions. They act as a reading grid: a figure can be perfectly accurate and still unusable, if it arrives six months too late or does not measure what the indicator announces.

DimensionThe question it asksWhat breaks it
ValidityDoes the data really measure what the indicator claims to measure?An ambiguous indicator definition, read differently by different teams.
ReliabilityDoes the same method applied twice give the same result?A collection protocol that changes mid-project, or enumerators trained differently.
PrecisionIs the figure accurate, with no entry or calculation error?A manual copy between two spreadsheets, a broken formula, an undetected duplicate.
IntegrityIs the data protected from arbitrary modification?A shared file anyone can edit, with no record of who changed what.
TimelinessDoes it arrive early enough to inform a decision?A quarterly consolidation on a project whose decisions are taken weekly.

The last two are the most neglected and the most costly. Accurate but untraceable data cannot be defended in front of a donor, and correct but late data has served no one.

When to run a DQA

  • Before an important donor report, so that you do not discover the gap during the review.
  • When a donor announces one: better to have run your own first.
  • After a change of team or tool, the moment when indicator definitions quietly start to diverge.
  • When a figure is surprising: an achievement rate of 140 % is more often a definition problem than a success.

On a properly tooled project, a full DQA takes a few days. On a project living in a dozen shared spreadsheets, it takes weeks, and that is itself a finding of the DQA.

How it unfolds

The principle is the same at any scale: pick a few indicators, and walk their chain back to the source.

  • 1. Pick the indicators. Three to five is enough, chosen among those carrying the most weight in the report.
  • 2. Rebuild the chain. From the figure in the report to the consolidated file, from the consolidated file to the collection export, from the export to the original submission.
  • 3. Recalculate. Rebuild the value from the raw data, without looking at the announced figure, then compare.
  • 4. Check the definition. Ask two people to define the indicator from memory. If the definitions diverge, the data already does.
  • 5. Document and fix upstream. A gap found is fixed in the process, not only in the figure.
The test that finds the most in five minutes: ask where the raw data for the indicator is. If nobody knows, or if the answer is “in so-and-so's file”, the integrity dimension is already lost, whatever the accuracy of the figure.

What is fixed at source, not afterwards

The most useful lesson of a DQA is that a large part of what it finds should never have been possible. An absurd value blocked at entry costs a second; the same value found six months later costs a day, and is sometimes no longer fixable.

  • Entry constraints. A validation rule in the XLSForm stops an inconsistent value from getting in, with an error message written for the enumerator.
  • Testing the questionnaire before release. Filling it in yourself once reveals the question order, the skip logic, the display conditions and the real duration, before thirty enumerators leave with it.
  • Enumerator traceability. Knowing who collected what lets you compare results by enumerator and spot a collection bias that no check on values would ever see.
  • One chain, with no copying. Every manual handover from one tool to another is a place where data can change without a trace.

On preparing a robust field collection, and what gets settled before teams leave, the article speeding up data collection with KoboCollect takes the subject from the field team side. This page answers what to check; the article answers how to organise.

Common mistakes

  • Treating the DQA as a check on staff. A team that fears the DQA hides gaps instead of flagging them, and the data gets worse.
  • Checking the figure without checking the definition. Two teams correctly counting different things produce a wrong total.
  • Fixing the report without fixing the process. The same gap will be back next quarter.
  • Only running a DQA when the donor asks. That is the worst possible moment to discover a problem.
  • Confusing cleaning with quality. Removing awkward values improves the table, not the data.

Glossary

DQA
Data Quality Assessment, a structured check of the quality, consistency and traceability of data.
Validity
The fact that a piece of data really measures what the indicator claims to measure.
Reliability
The fact that the same method, applied twice, gives the same result.
Traceability
The ability to trace a reported figure back to the original data and to who produced it.
Raw data
The submission as it was collected, before any cleaning or aggregation.
Entry constraint
A validation rule preventing an inconsistent value from being recorded, with a custom error message.
Collection bias
A systematic deviation introduced by the way of collecting, for instance an enumerator rewording a question their own way.

Frequently asked questions

What is a DQA in MEAL?

A DQA (Data Quality Assessment) is a structured check of the quality, consistency and traceability of reported data. It starts from a figure in the report and walks the chain back to the data collected in the field, to verify that the value holds at every level.

What are the dimensions of data quality?

Five dimensions are generally used, notably by USAID: validity (the data measures what the indicator announces), reliability (the method gives the same result twice), precision (no entry or calculation error), integrity (the data is protected from arbitrary modification) and timeliness (it arrives in time to inform a decision).

When should you run a DQA?

Before an important donor report, when a donor announces one, after a change of team or tool, and whenever a figure is surprising. An achievement rate above 100 % is more often an indicator definition problem than a success.

How do you prevent quality problems rather than fix them?

By acting at source: validation constraints in the questionnaire that stop an inconsistent value getting in, a test of the form before release, enumerator traceability on every submission, and a chain with no manual copying from one tool to another.

Is a DQA there to sanction teams?

No, and using it that way makes it counterproductive: a team that fears the DQA hides gaps instead of flagging them. The vast majority of problems found come from an ambiguous definition or a manual copy, not from individual fault.

Data quality with Opti'

Opti Gemba treats the Data Quality Assessment as an assignment type in its own right, just like a baseline or a PDM: it has its terms of reference, its questionnaire, its collection and its report.

But the essential part happens before. The generated questionnaire carries its validation constraints and its error messages, it can be tested before release, and every submission keeps the enumerator's name as a usable variable: you can compare results by enumerator and spot a collection bias. Since the questionnaire, the collection, the analysis and the report all live in the same project, there is no manual copying between them, so nowhere the data can change without a trace.

What Opti' does not do: it does not guess that a plausible value is wrong. No tool can. What it brings is the traceability that lets a human decide, and the constraints that stop the implausible from getting in.

Make your data traceable from collection onwards

Create your free account, no card required, and generate a questionnaire that constrains entry at source.

Open Opti' →