notebook / evaluation / engineering
Our tests missed a kind of correction
A correction across two sentences exposed a model gap. Fixing the model then exposed a check that rejected the right answer.
Pom’s shorthand
- Correcting yourself across sentences exposed a gap in the old model
- The right answer was shorter than our app allowed
- A useful test follows the result all the way to what gets pasted
Written in September 2026 about the change merged on August 4, 2026 (UTC).
A person can change their mind halfway through a sentence. They can also finish the sentence, pause, and correct it in the next one. To the person speaking, both are ordinary corrections. To our cleanup model, that boundary mattered.
The v2 model handled a correction inside one sentence but could retain the superseded instruction when the correction started a new sentence. The August 4 integration record describes this failure: a speaker replaced a meeting day, but the cleaned result kept the first day.
This was a gap in the behavior we had taught and checked. More examples of the familiar shape would not, by themselves, establish that the unfamiliar shape worked. We needed to test the boundary we had missed.
Start with the kind of mistake
The useful unit was not simply “a transcript.” It was a transcript containing a particular kind of change: an instruction, followed by another sentence that replaced it.
That distinction changes what a test asks. A general check might accept a grammatical sentence about a meeting. A correction test must ask whether it kept the final intended day. The wrong answer can be short, fluent, and completely plausible.
The v3 integration targeted that gap with a model trained for cross-sentence corrections. But changing the model was only part of the work. Its answer still had to pass the app's checks before reaching the text field.
The right answer looked too short
One of those checks compared the length of the cleaned result with the original transcript. If the result was less than 30% as long, the app treated it as excessive deletion and fell back to the raw transcript.
That check has an understandable purpose: cleanup should not throw away what someone said. A correction complicates it. Removing the entire superseded sentence can be exactly the right thing to do.
In the case we measured, the correct replacement was 27.7% of the original length. The incorrect result that kept the old day was 30.8%. The length check rejected the correct answer and accepted the wrong one.
These ratios compare the length check in one measured case, not model accuracy.
Replacing the model alone would therefore have changed the failure from “the wrong instruction gets pasted” to “the uncleaned transcript gets pasted.” We had to check the whole path.
Let the answer through, without claiming the check understands it
The change lowered the length floor to 15% when the input contained a correction marker. It deliberately excluded a bare “no”: that word can be ordinary content rather than a signal to discard an earlier instruction.
There is an important limit here. A 15% floor also accepts the old, incorrect 30.8% result. The revised check does not decide which day the speaker meant. The model handles the correction; the check stops blocking its shorter answer.
Keeping those jobs separate makes the evidence easier to read. A passing length check tells us that the output cleared that check. It does not tell us that the output preserved the speaker's intent.
Test what reaches the text field
The integration was verified through model preparation, cleanup, and output acceptance. The record reports 23 of 23 end-to-end cases passing, including the five reported failures and two examples that should not be treated as corrections.
Those results describe that suite. They are not a claim that every way of changing your mind now works.
There was also an unresolved edge case: a chained triple correction described as a regression in the model handoff passed on-device with a slightly different phrasing. The implementation kept it visible as a known gap instead of treating the discrepancy as settled. A small wording difference was enough reason to keep investigating.
The lesson is to name the behavior before counting the tests. Does the correction cross a sentence boundary? Does the speaker replace an instruction or add to it? Can the app accept the shorter result? Each is a different question.
The model's answer is an intermediate result. For a dictation app, the test ends with the words that actually reach the user.