Where: the Plan-with-Kai feature and its repository-context step, in production for every customer.
Symptom: none. Nothing was broken. That is what makes this the most important entry in the diary.
What it actually was
Some weeks ago, Kai genuinely could not read the contents of a connected repository. It saw only names, owners and whatever annotations a human had typed. So an engineer wrote an entirely sensible workaround for Plan-with-Kai: before each turn of the conversation, fetch the README and top-level file listing live from GitHub for up to six repositories, and paste the result into the prompt.
They documented it carefully, in both files. The comment said, in effect, that Kai’s own retrieval only saw pre-annotated metadata, and that the context was fetched fresh each turn, not persisted in the thread, so it could not go stale mid-conversation. Correct decision. Correct reasoning. Correctly written down.
Two weeks later, a different change taught Kai to read repository contents directly from its own index, the very capability whose absence justified the workaround.
Nobody removed the workaround. Every turn of every planning conversation still mints a token, makes up to a dozen live API calls, and injects up to six READMEs (3,000 characters each) plus six file trees into the prompt. That is roughly 20,000 characters, about 5,000 tokens, re-fetched and re-sent on every turn, explicitly uncached, to supply context the brain may already hold.
The reason nobody removed it is the entire point: the comment explaining why it was necessary is still there too, and it is still perfectly persuasive. It is just no longer true. Any engineer, or any AI agent, reading that file would conclude the fetch is load-bearing, because the written record says so.
There is a final twist. A design document written in between had explicitly warned, “Do not fix this a second, divergent way.” That is now exactly what exists, and it is recorded in neither document.
- Time hidden: ongoing at the time of writing. Found by auditing, not by anything failing.
- What it cost: unknown, and that is the interesting part. Nobody had ever measured it, because nothing was broken. On a six-repository customer it is roughly 5,000 tokens and a dozen network calls per conversational turn, on every planning conversation since.
The lesson
A workaround outlives the constraint that justified it, and its own justification is what protects it. Every workaround needs an expiry condition written next to it, “delete this when X becomes true”, not just an explanation.
What we did about it
This exact code path becomes the first internal customer of the Context Gateway. We measure what the turn costs today, prove the index can serve it, remove what is redundant, and publish the difference as our first token receipt. If our own brain cannot beat twelve GitHub calls and 5,000 tokens a turn, we would rather find out before a customer does.



