notebook / performance / engineering

The optimization we stopped questioning

A comment said prompt caching could not work. Measuring the repeated work showed why we needed to question it.

Pom’s shorthand

  1. Re-reading an unchanged prompt took most of the measured cleanup time
  2. The blockers were in our caching code, not the model architecture
  3. Faster only counted after cached and uncached runs produced identical text

Written in September 2026 about the change merged on August 4, 2026 (UTC).

After you stop speaking, waiting for your words interrupts the next thing you wanted to do. In one of our cleanup measurements on an M1, that wait included work we had already done: reading the same instructions again.

The instructions told the model how to clean a transcript. They did not change between dictations. Yet we were paying to process them every time. In the measured request, reading the prompt took about 1.09 seconds of a roughly 1.48-second cleanup run.

We had a reason for leaving it that way. A comment explained why reusing that work could not work with the model. The observations behind the comment were accurate. The conclusion was wrong.

The work that kept happening

A cleanup model first reads its instructions and the transcript, then writes an answer. The first stage is called prefill. A prefix cache saves the model's state after reading the unchanged beginning, so the next request can start from there.

Our frozen instructions accounted for about 256 of the roughly 279 tokens in a typical request. Tokens are the pieces of text the model reads. Most of this request was therefore the same material, in the same order, as the previous request.

That was worth measuring before changing anything else. Clearing the memory cache looked wasteful, so I tried removing it. That made the run slower. Clearing the cache, rendering the prompt, and tokenizing it together took about seven milliseconds; they were not where the missing second had gone.

“Our code cannot do this” became “this cannot be done”

The model mixed two kinds of layers. Some tracked how many tokens they had processed. Others maintained a running state but reported an offset of zero. Those latter layers also could not undo an extra step by trimming their state.

Our caching code made two assumptions that did not fit that combination. It checked only the first layer's offset, which was always zero here. It also used an iterator that generated an extra token while preparing the cache. That extra step would then need to be undone—the very operation the running-state layers could not perform.

Those were real obstacles in our implementation. They were not proof that the model could never reuse a prompt.

The fix was to read the prefix with a plain forward pass, without generating an extra token. Nothing overshot, so nothing needed trimming. We also checked every layer, allowing zero for the running-state layers and requiring at least one counting layer to reach the prefix length.

The useful distinction was small: a model limitation had exposed a mismatch in our recipe. Changing the recipe removed the obstacle.

The measured result—and its limits

Two bars starting at zero compare median cleanup time on an M1: 1.48 seconds before prefix reuse and 0.58 seconds after. Lower is better.

The 23-case end-to-end cleanup suite reported p50 falling from 1.48 seconds to 0.58 seconds, a 2.55× speedup. This measures the cleanup path on the tested M1 setup, not speech recognition, insertion, or every user's total wait.

That timing result was encouraging. It was not enough to accept the change.

A bad cache need not crash. It can produce fluent text that differs from the uncached answer. A test that asks only whether the output looks reasonable can miss precisely the change we need to catch.

So we ran the same 13 transcripts both ways: once without prefix reuse, once with it. All 13 outputs were identical. The test also checked that the cache actually existed, so it could not accidentally compare two uncached runs and declare success.

That separate comparison measured medians of 1.61 seconds and 0.70 seconds. Those are different runs from the 23-case suite above. They support the equivalence check; they should not be mixed into a single benchmark. Nor do 13 matching transcripts prove that every possible input will match.

Even the green test run needed a question

There was one more assumption hiding in the verification. The first full Swift run reported 519 tests and no failures. It had never compiled the new test file: the local Xcode project was stale.

After regenerating the project, the suite reported 520 tests. The model-dependent comparison was separately exercised for the 13-transcript check. “The suite passed” had not, by itself, established that the new check had run.

The lesson I want to keep is not “ignore comments.” It is to preserve the difference between an observation and the explanation attached to it. “This layer reports zero” is a useful fact. “Caching is impossible” closes a question that still needs an experiment.

For this change, the experiment had two parts: measure the work we were repeating, then compare the output with and without the shortcut. We needed both before the speedup meant anything.

was this useful? stamp it.