notebook / performance / experiments
Faster cleanup. Different words. We said no.
Cleaning speech in smaller pieces reduced the measured wait, but it failed the experiment's other requirement: preserve the whole-transcript output.
Pom’s shorthand
- Independent blocks reduced median tail cleanup time from 6.9 to 3.9 seconds in the replay
- None of the 21 block-by-block outputs exactly matched whole-transcript cleanup
- We recorded a no-go for this approach instead of shipping the speedup
Written in September 2026 about the experiment recorded on September 13, 2026. This was a research spike, not a shipped feature.
The tempting part of a long dictation is the time before you stop speaking. While you are saying the last sentence, the first few sentences are already there. Could cleanup work on those earlier sentences and leave only a small tail to process when you release the key?
That was the question behind our September 13 experiment. We set two requirements: reduce the wait, and preserve the output we would have obtained by cleaning the whole transcript. The second requirement matters because changing the boundaries of the input changes what the model can see.
The experiment found a shorter wait. It did not preserve the output.
What we actually measured
We replayed 21 existing dictations of at least 400 characters through the shipped v3 model. Each transcript was split at sentence boundaries into blocks of at least 200 characters, with the final block treated as the tail. This post shares only the aggregate results, not the recordings or transcript text.
The first variant cleaned each block independently and joined the results. The second supplied the previous raw block as context, then tried to align the answer with the previous cleaned block so it could keep only the new part.
Both were compared with cleaning the complete transcript at once. The replay measured the final block's cleanup time as a proxy for the work remaining at key release. It was not a live test of speech recognition and cleanup running together while a person talked.
A faster tail, a different result
Times are medians from the replay; the split variants measure tail cleanup. Hardware is not specified in the experiment note, so these numbers should not be treated as a device benchmark.
Independent blocks reduced the median from 6.9 to 3.9 seconds: about 43% less time, or a 1.77× speedup calculated from those medians. None of the 21 joined outputs exactly matched the whole-transcript result.
The original note labels that result “2.0×.” The reported medians do not yield that ratio, so the raw times are the useful comparison here. They also do not, on their own, demonstrate the original twofold speed target.
Adding previous-block context brought the median to 6.2 seconds, with three exact matches out of 21. It recovered little of the speed benefit and still failed the unchanged-output requirement.
Similar is a different promise
An exact mismatch does not automatically mean the result is wrong. The experiment note describes many small differences, including punctuation and sentence grouping. It also reports a much larger difference in the worst independent-block case.
The comparison included word-level similarity. Eight of the 21 independent-block results reached at least 0.95 similarity; the context variant reached that threshold on seven. Similarity tells us something about overlap. It does not establish that a correction kept its meaning, and it does not satisfy a requirement for identical output.
That is why “most of it looks similar” would have been a different acceptance rule. We would have needed to choose and test that tradeoff explicitly, rather than quietly substitute it after seeing the timings.
Why the boundary mattered
Our cleanup model could make different punctuation, segmentation, and filler-removal decisions when it saw the entire utterance. A sentence boundary gave us a convenient place to split the input. It did not guarantee that the model's decisions were independent on either side.
The context variant tried to restore some of the missing information. But aligning its answer with a previously cleaned block introduced another failure point, and processing more context cost time. More surrounding text did not automatically make the joined result equivalent.
There was a separate integration question too. Live drafts were not guaranteed to remain a stable prefix of the final transcript. Even a successful replay would still have needed a reliable rule for deciding which words were settled enough to clean.
A useful result can be a no-go
We recorded the experiment as a no-go for this approach. We kept the spike as an experiment rather than shipping it as a feature. That distinction is the outcome of this story.
The note left possible reasons to revisit it: a model trained specifically for block-shaped inputs, or an explicit product choice for very long dictations. Those were future directions, not demonstrated fixes.
For this experiment, the decisive result was simple. We found a way to do less work at the end, but we had not shown that the result stayed the same. Keeping that requirement intact was more useful than shipping a faster number.