Baseline and endline, in one sentence
The baseline, or baseline study, measures the reference situation of the target populations before the project starts acting. The endline, or final survey, measures the same thing after. The difference between the two is what the project claims to have changed.
Between the two, a midline can be run at mid-term. It is not there to conclude but to correct: if an indicator has not moved after a year, there is still time to change approach, which is no longer true at the endline.
When to run each of them
The baseline is run before activities begin, and that is the constraint most often broken. A baseline study carried out three months after distributions have started no longer measures a starting point: it already measures an effect, and the project loses its reference.
- Baseline: in the first weeks of the project, before any field activity.
- Midline: at mid-term, on a project long enough for a correction to still be possible.
- Endline: at the end of activities, early enough that participants can still be reached.
On a short project, under a year, a midline usually makes no sense: by the time you collect, analyse and decide, the project is over. A light continuous follow-up works better, such as a PDM after each distribution cycle.
The rule that decides everything: comparability
An endline only compares to a baseline if both asked the same questions, of the same kinds of people, in the same way. Every difference introduced between the two becomes an alternative explanation for the change observed, and makes the conclusion contestable.
| What must stay identical | What can change without harm |
|---|---|
| The exact wording of the indicator questions | Context questions added at the endline |
| The answer options and their order | The order of the questionnaire sections |
| The recall period (“in the last 7 days”) | The language of collection, if the translation is validated |
| The sampling method and the sampling frame | The enumerators, if training is equivalent |
| The thresholds and calculated scores, with their weighting | The collection tool, as long as the data comes out in the same format |
| The season, as far as possible | The number of supervisors |
The safest way to hold this rule is to reuse the baseline questionnaire rather than rewrite one. It is also why the baseline XLSForm file must be archived along with its data, and not only the final export.
What is measured at each end
The indicators of a baseline and an endline are those of the logical framework, at the outcome and specific objective levels. They are the ones carrying the promise made to the donor, so they are the ones that must be measured twice.
- Objective and outcome indicators: what the project claims to change. Mandatory at both points.
- Household characteristics: size, composition, sex and age of the head of household, status (displaced, host, returnee). They serve disaggregation and let you check that the two samples resemble each other.
- Composite scores: food consumption score, coping strategies index, dietary diversity score. Their weighting must be strictly the same at both points.
- Activity indicators: these belong to routine monitoring, not to the survey. Including them lengthens the questionnaire without adding anything to the comparison.
An overlong baseline questionnaire is the most widespread flaw. Every question added will have to be asked again at the endline, of the same population, in the same terms. A question nobody can say what it will demonstrate costs twice.
Sample and method
Sample size is calculated on the main indicator, not on comfort: it is what determines whether a difference observed between baseline and endline is a real change or noise. The sample size calculator gives the figure and the formula that produces it.
Two methodological points are decided at the baseline and then apply without further debate:
- Panel or independent samples. Re-interviewing the same households is statistically more powerful, but assumes you can find them two years later, which is rarely a given in a displacement context. Two independent samples drawn from the same frame are more robust in practice.
- Comparison group or not. Without a control area, a baseline and an endline show an evolution, not an impact attributable to the project. That is a limitation to write into the report, not a flaw to hide.
Qualitative work usefully completes both points: a few focus groups at the endline often explain why an indicator has not moved, which no percentage will tell you.
Common mistakes
- Running the baseline after activities have started. The reference is lost, and no analysis rebuilds it.
- Rewriting the questionnaire at the endline. A question reworded “to be clearer” breaks the comparison on that indicator.
- Changing the recall period. Going from 7 to 30 days between the two surveys is enough to reverse a trend.
- Losing the baseline file. Without the original questionnaire or the raw data, the endline measures something else.
- Comparing samples of different composition. If the endline reaches more displaced households than the baseline, the gap observed may come from that and nothing else.
- Concluding impact without a comparison group. An evolution is not an attribution.
On how to present a gap between what was planned and what is measured, without dressing it up or dramatising it, the article justifying gaps between planned and actual results takes the question from the reporting side. This page answers how to measure; the article answers how to tell the donor.
Glossary
- Baseline
- Baseline study, run before activities begin to establish the reference situation.
- Midline
- Mid-term survey, meant to correct course while there is still time.
- Endline
- Final survey, run at the end of activities to measure the change.
- Sampling frame
- The list of units (households, individuals, villages) from which the sample is drawn.
- Recall period
- The time window a question refers to, for example “in the last 7 days”.
- Comparison group
- A population not reached by the project, measured in parallel, which allows a change to be attributed to the project rather than to the context.
- Composite score
- An indicator built from several weighted questions, such as the food consumption score.
Frequently asked questions
What is the difference between a baseline and an endline?
The baseline measures the situation before project activities begin, the endline measures the same situation at the end. Comparing the two gives the change observed. Both surveys must ask the same questions of comparable populations, otherwise the difference cannot be interpreted.
Can you run an endline without a baseline?
You can run the survey, but you cannot measure change: without a starting point, an endline describes a situation without saying where it came from. Fallbacks, such as asking respondents to recall their situation two years ago, are unreliable and must be flagged as such in the report.
Should you interview the same households at baseline and endline?
It is not mandatory. Re-interviewing the same households (a panel approach) is statistically more powerful but assumes you can find them again, which is hard in a displacement context. Two independent samples drawn from the same sampling frame are more robust in practice.
What sample size for a baseline?
It is calculated on the main indicator, from the population size, the margin of error accepted and the confidence level targeted. Opti's sample size calculator gives the figure and the formula that produces it, and the same size must be targeted at the endline.
Can a baseline questionnaire be generated automatically?
Yes. Opti Gemba generates a complete XLSForm questionnaire from a plain-language description, with composite scores and their weighting, ready to import into KoboToolbox or ODK. The same questionnaire is reused at the endline, which guarantees comparability.
Running a baseline and an endline with Opti'
In Opti Gemba, one project holds several assignments: baseline, midline, endline, PDM, final evaluation. They share the project sheet, the logical framework and the indicators, which mechanically settles the comparability problem: the endline starts from the baseline questionnaire, not from a blank page.
Composite scores are built by the calculation engine, which checks that all food groups of a food consumption score are accounted for and that the weighting matches the standard. The analysis then disaggregates by sex, age, locality and status, and the narrative report exports to PowerPoint, PDF or HTML.
OptiBot
Methodological advice on choosing indicators and building the measurement plan.
Find out more →
Opti Gemba
Baseline, midline and endline assignments in one project, from questionnaire to report.
Find out more →
Opti Academy
Brush up on sampling basics before a baseline, at your own pace.
Find out more →