The principle, in one sentence
The numbers are computed, the text is written, and the two never mix. A language model is not reliable at counting: it produces plausible numbers, which is exactly the worst possible flaw in an evaluation report. It is, however, perfectly able to write a clear narrative from results it is given.
The whole method follows from that separation. Your responses are analysed by a deterministic computation engine, which produces a ledger: every value in it carries a unique identifier. The writing stage sees only that ledger, and can cite a number only by referencing it. Before the final render, every number in the text is matched back against the ledger. Anything that does not match is corrected; anything that matches nothing is flagged.
In practice: if the report states “78.4% of beneficiaries report being satisfied”, that 78.4% was computed from your responses before being written, then checked again afterwards. It was never “estimated” by the model.
The eight stages of a report
Writing is not a single call to a model. It unfolds in successive phases, each answering a specific risk of automated production. You take part in two of them, and they are not the least important ones.
| Stage | What happens | The risk it covers |
|---|---|---|
| 1. Indicators | Each indicator in your logical framework is matched to the question that measures it, and to the answers that count as “achieved”. | An indicator measured on the wrong question, or left out. |
| 2. Selection | Analysable questions are sorted and grouped into parts. Every question left out carries a reason. Those still left out are then read a second time, with the outline in view: the ones that did belong somewhere are put back in. | A useful question dropped without anyone knowing. |
| 3. Plan (you) | You review the plan before any computation: you tick, untick, rename the parts, move questions around. | Analysing something other than what matters to you. |
| 4. Reflection | The cross-tabulations worth looking at are identified, along with the hypotheses they would let you test. | A descriptive report that cross-tabulates nothing. |
| 5. Analyses | Each table is computed, then read. The results feed a ledger of identified values. | Approximate figures. |
| 6. Deeper analysis | A second pass revisits the first findings and digs into what actually stands out. | Missing what the data actually shows. |
| 7. Verification (you) | Opti' puts its readings, its apparent paradoxes and its understanding of your notes to you. You confirm, refine or reject. | The report asserting an interpretation the field contradicts. |
| 8. Writing | The report is written section by section, then its synthesis, its recommendations and its annex. | A text that repeats or contradicts itself from one section to the next. |
You follow this sequence live while the report is generated: the current stage, how many analyses are done out of the total, and the findings as they come. An interrupted generation resumes where it stopped.
What is computed, what is written
The boundary is sharp, and it is what makes the report verifiable. Here is what falls on each side.
Computed, with no model involved
- Every percentage, average, sample size and cross-tabulation.
- The annex tables, exhaustive, including questions the text does not comment on.
- Indicator results, against their target.
- The figures: their bars, their axes and their values.
- The indicative margin of error, from the sample size.
Written by the model, from those computations
- The narrative of each part, and the executive summary.
- The cross-cutting analysis, the conclusion and the recommendations.
- The figure titles, which rephrase your questions.
- The methodology note, from your own input.
The checks, one by one
Every cited number is matched back to the ledger
During analysis, each computed value receives an identifier. The writing cites those identifiers, and a final check compares the written number with the recorded value. Three outcomes: confirmed (the number matches), corrected (the number is replaced by the real value), or flagged (the number matches no value and is not shown as verified). You are given the count of all three at the end of the generation.
You can trace any number back to its source
In the report, verified numbers are underlined. Hovering over them shows where they come from: the question analysed, the cross-tabulation, the group concerned, the calculation mode, and the full table row. This traceability travels with the exported HTML document, with no connection required.
Your notes take precedence over any reading of the figures
A note you wrote on a question during analysis is treated as field truth, taking precedence over any spontaneous reading of the table. If your note explains that a group “perceived as excluded” simply does not meet the selection criteria, the report is not allowed to present it as a victim of exclusion, nor to base a recommendation on that.
The annex is exhaustive, and independent of the narrative
Every analysed question appears in the annex with all its pre-computed disaggregations, including those the body of the report never mentions. A reader who doubts a statement can go and look at the table itself. The annex passes through no model: it is a direct reflection of the computation.
The figures are legible, including to a colour-blind reader
Series colours are validated by computation rather than picked by eye: enough hue separation for all three forms of colour blindness, a minimum saturation so none reads as grey, and values printed on the marks so colour is never the only information. Beyond eight groups, the remaining ones are grouped together rather than given a colour already in use.
What the model is not allowed to do
- Invent a number. It can only cite values from the ledger, and what it cites is checked again afterwards.
- Announce a result in a figure title. A title that tells you what to see before you have looked is wrong half the time: titles rephrase the question, never the answer.
- Contradict a field note or an answer you gave during the verification phase.
- Introduce a topic absent from the data. An analysis is about the question asked and its answers, not about what a model happens to know of the sector.
- Present a hypothesis as a fact. A possible explanation is worded as one.
The limits, stated plainly
A method that does not state its limits is hiding one. Here are ours.
- The computation is exact, the interpretation remains a reading. A gap between two groups is a fact; explaining it is a hypothesis, and the report words it as one.
- Correlation is not causation. The report cross-tabulates variables, it does not establish causality, and should not be read as if it did.
- Sample size calls for caution. A gap measured on a group of twelve does not carry the weight of a gap measured on three hundred. Sample sizes are shown everywhere, precisely for that reason.
- The quality of the report depends on the quality of the questionnaire. A badly worded question produces an exact answer to the wrong question.
- Human review is not a formality. It is the last stage of the method, not an option after it.
Human review
A report written with the help of a model commits no one until a human has reviewed it. That is why every exported report carries a validation block at the end: the person who reviewed it attests to that, gives their name and role, and saves the signed document. A review is often done by several people: as many reviewers as needed can be listed.
That block is never ticked in advance. A pre-ticked box would be worth nothing, while claiming the opposite.
The saved file is a final copy: the review statement is fixed in it, with no input field and no button, and can no longer be edited there. It is an attestation of human review, carried by the document; it is not a certified electronic signature, and Opti' does not claim it is.
Your data
Sensitive content is encrypted server-side with AES-256-GCM. No collected data feeds the training of any model. The full arrangement, the processors involved and the rights attached are described on the page Ethical AI and data protection.
Report production sits inside the Opti Gemba module, which covers the MEAL assignment from the logical framework to the final report.
Frequently asked questions
Are the report's numbers produced by artificial intelligence?
No. Every number comes from a computation performed on your responses, not from the language model. The computation produces a ledger in which each value carries an identifier. The writing only cites those identifiers, and every number in the final text is then matched back against the ledger: confirmed, corrected, or flagged as unverifiable. A language model is not reliable at counting; it is reliable at writing. The method keeps the two apart.
What if the AI misreads a result?
That is exactly what the verification phase is there to prevent. Before writing a single line, Opti' puts to you the readings it is about to assert, the figures behind them and the possible alternative explanations. You confirm, correct or reject each one. Your answers become priority instructions for the writing, which can no longer contradict them. A drop in satisfaction may come from a stock-out: the AI cannot know that, your team can.
Can I edit the report after it is generated?
Yes, entirely. Every paragraph is edited directly in the document. Every figure can be renamed, switched to another chart type, recomputed on another question or another disaggregation, added or removed. The analysis plan itself is confirmed before any computation: you tick, untick, rename the parts and move questions around.
Does the report state that it was written with AI?
Yes, on its cover page, with a link to this page. And at the end of the document, a block lets the person who reviewed it attest to that and sign. A report produced with the help of a model commits no one until a human has reviewed it: saying so is part of the method.
Is my data used to train a model?
No. Sensitive content is encrypted server-side with AES-256-GCM, and no collected data feeds the training of any model. The full arrangement is set out on the page devoted to ethical AI and data.
What are the limits of the method?
They are real and we write them down. The computation is exact, the interpretation remains a reading: correlation is not causation, and a gap between two groups may come down to sample size. The report flags it when the sample is small, but it does not replace the judgement of someone who knows the field. Human review is not a formality, it is the last stage of the method.