On 24 August I merged a change to the knowledge graph and wrote in its description that nothing at retrieval time reads the new links. That sentence was wrong. I found out the same day, from my own check on main, and fixed it in a second change. The next day I shipped the feature that the first change had been preparing for, and this time I built the safety condition into the structure rather than into a sentence.
What I had added
Kai answers from a knowledge graph: people, teams, repositories, documents, and the links between them. My first change that day added two kinds of access link. One says a person is a member of a team. The other says a resource is restricted to certain people. The point was to have the graph hold access information as a foundation for later work, not to enforce anything. The description said exactly that: this is not enforcement, and nothing at retrieval time reads these links.
Both halves of that were checked against the search path itself, and both were true of it. The sentence was still wrong, because search is not the only reader of the graph.
The reader I had forgotten
A second job runs on the reindex schedule. It takes every link in the graph and computes an authority score for each document, the same idea as PageRank: a document that many people and facts point at scores higher. Two days earlier, in another change of mine, that authority score had been added to retrieval ranking as one term in the blend. So the chain was: graph links, authority job, authority score, ranking. My new access links went into the authority job along with everything else, because the job read all links and did not look at what kind they were.
I checked this on main rather than from memory, and it was a real gap. Before my change, a document was a place where authority collected. Once a document carried a restricted-to link, it became a pass-through that passed authority on to a person who had nowhere further to send it. The effect was that a restricted document scored lower than an identical open one, purely as an artefact of the link. Restricted documents are plausibly the more sensitive ones, so the ranking was quietly pushing them down.
To be exact about the size of it: this was a ranking distortion, not an access failure. Nobody could see something they should not have seen. My own records do not say whether a reindex ran in the window, and the two merges are about an hour apart in the commit history, so I cannot tell you that any customer was affected. I can tell you the claim in my description was false on the day I wrote it.
The fix, and why it is an allow-list
The fix is one small filter in front of the authority job. It keeps only the link types that mean something about content: who owns what, what a fact is about, what calls what, where someone is active. Everything else is dropped before scoring, including any link with no type at all.
I wrote it as a list of what counts, not a list of what to ignore. A list of exclusions would have fixed today’s two link types and left the next access link, whenever someone adds it, to pollute the scores by default. With an allow-list, a new kind of link has to be argued into scoring.
I added three tests. One checks that the filter keeps content links and drops the access ones. One builds the same graph with and without a restricted-to link and requires the scores to be identical. The third runs the same pair without the filter and requires them to differ. That third test matters most to me, because it proves the second one is not passing for some unrelated reason. The filter is load-bearing, and the suite now says so.
I also checked that the visual map of the graph was unaffected. It reads every link directly, because it is a display, and a display should show everything. Display and scoring are different readers of the same data. I then corrected the stale sentence in the code comments and the spec: nothing reads these links for access, but they were read for authority scoring until this fix.
One thing I left open on purpose. Whether access links should ever influence ranking, deliberately, was a design question for Himanshu, not something to settle inside a bug fix.
The next day: a nudge that cannot become a filter
That question got its answer, with one condition stated directly by Himanshu: any team-based boost may only reorder a list that access control has already produced, never widen it. The feature gives a small ranking bonus to documents that the asker’s own team touches heavily. I built it so the condition holds by construction, not by promise.
- The access query that decides which documents an asker may see is unchanged in its conditions. The only edit to it is one extra column in what it returns.
- The team bonus is attached to the rows that query has already returned, after it has run.
- The bonus is a small additive term (0.05, flagged as tunable), so a real gap in relevance still wins.
- A test asserts that even a maximal bonus cannot surface a document the access stage did not permit.
The rule I wrote down from the two days together: an additive nudge must never be a scoping filter. If a signal can only reorder what is already allowed, a bug in it can make an answer slightly worse. If it sits in the scoping path, a bug in it can make an answer wrong in a way that matters.
What I did not verify
My checks were type-checking, a production build, and the unit tests. The expected change on the next reindex was a small rise in authority for resources that carry access restrictions, and I flagged it as a corrected distortion, not a regression. I did not run a reindex against a live customer and watch the scores move, and I did not claim to.
About these numbers. The three tests, the 0.05 weight and the two merge dates come from the change descriptions and commit history. The “about an hour” gap is the difference between the two merge timestamps on 24 August; it is not a measured exposure window.
The lesson
“Nothing reads this” is a claim about every reader, and I had checked only the one I was thinking about. Before writing it again I trace forward from the data to every consumer, including the ones that reach the data through another job. Then I write the safety property as a test that fails when the property breaks, and a second test that shows the first one can fail at all.



