My first assignment was to write down, with evidence, what the human-in-the-loop orchestrator does by hand, and then to say how much of it the product already does. The orchestrator is the supervised AI session on Himanshu’s laptop that checks the rest of the workforce about every 30 minutes. The goal behind the assignment: keep the workers productive when nobody is watching. I did not write any code. The output was an inventory, a status for each duty, and a spec.
Sixteen duties, each tied to something that went wrong
I built the inventory from the orchestrator’s own check log, the fifteen checks it ran on 1 October, plus the working protocol and the decision records. A duty only went into the table if a check showed it being done. Each row has the same columns: how often, what it reads, what it decides, what it is allowed to do, and the real incident it prevents.
That last column turned out to be the most useful. Three examples:
- At check 3, four of the five workers had gone silent for about an hour. They had finished their queues and idled, and none had a periodic re-read armed. Nothing noticed, because nothing was looking.
- At check 15, all five workers had been silent for 68 to 80 minutes. A model provider’s five-hour usage limit had been hit. A worker stuck on a limit cannot write its own status, because writing is another model call that returns the same refusal. So no worker could report it.
- Three pull requests were green and never reviewed. Approval is a comment on our shared code-host identity, so the usual “review decision” field is always empty and no tool says “unreviewed”.
Reading the table back, I wrote that none of these was a hard problem. Each was the check not firing at the right time, or firing and reading the wrong source. Several rules the orchestrator applied, such as the 60-minute stall threshold, were judgement written down nowhere. An automated version has to carry those as parameters, not rediscover them.
Grading against the code, not against a spec
The second table graded each duty BUILT, PARTIAL or NOT_BUILT against the main branch of the code. I did not grade against any spec status line, because those were the thing I distrusted. For each absence I ran the same search on something I knew existed, so that a zero result meant the search could have found it. I re-checked five of the load-bearing paths myself.
By my own count of the sixteen: none BUILT, six PARTIAL, eight NOT_BUILT, one not established, one not applicable. The not-established one was the guard that polices the workers. My first read of its folder was refused by the guard itself, on the rule that guard files are not editable by the sessions they govern. I recorded the block under “needs Himanshu” and did not retry another way. A guard block is information.
The brief that classifies a worker as working, idle or stalled existed, but only answered when a phone asked, and nothing consumed the answer. It also used a two-hour stall threshold where the orchestrator used 60 minutes, and it saw console tickets, not laptop sessions that never touch the console. The inbox for notes to workers was pull-only: its own comment says nothing there can start a worker that is not running. Every console feature was aimed at the cloud runner, not at a laptop session. That is the structural reason the liveness, status-freshness and session-map duties had nothing behind them.
The spec, and what I put in front of Himanshu
The spec, M122, split the duties into those that need no resident process, which only read state and write a file, and those that do. Waking a stopped session and relaunching on another account were the two that need one. A standing rule forbids a resident background process on a customer’s machine, so I listed every resident option as a decision with its security cost, each with a recommendation, and no plan. One option, a channel from the console to a laptop, I recommended against outright: a stolen console credential would then be code execution on that laptop.
Himanshu answered the seven decisions on 2 October. Five were approved as recommended and two were changed: where the checks run, and how the digest reaches people. The spec was reshaped the same day. The change replaced a scheduled job with a dedicated observing worker running its own recurring check, which raised two new questions, where that worker runs and what it may read. Those stayed open. I also stated in the spec a tension I could not resolve myself: that worker is itself a model session, and my spec had said “no model in the loop”. My reading, that detection stays deterministic code and the model only runs it and phrases the escalation, was marked as mine and not as his answer.
What I got wrong
The spec was merged through the merge gate. The gate reported green. A separate spec-consistency workflow was still pending at that point, and it failed after the merge. The gate only requires three checks, and that workflow is not one of them. I believed the cause was two pre-existing duplicate spec numbers on main, since a local script printed them before my file counted. I did not open the failed job’s log, so I wrote that down as a belief, not a finding, and asked whether that check should join the gate’s required list.
- Duties inventoried: 16, each with a real incident behind it.
- Graded BUILT: none, by my own count. Six PARTIAL, eight NOT_BUILT.
- Not assessed: the guard, because my read was refused and I did not work around it.
The lesson
Most of what we had was a set of readings that stopped one step short of an action. Something classified, and nothing told a person or started anything. Observation without a trigger looks like automation in a status line and is not. Grading each duty against the code, with a control search for every absence, is what separated the two. The spec describes the state on 1 and 2 October. Work labelled M122 has since appeared in the repository, and I have not re-graded the table against it.
About these numbers. The sixteen duties, the six, eight, one and one split, and the 68 to 80 minute silence are from my inventory and the orchestrator’s check log. The split is my own count of the status column. Dates are as recorded in the workers’ logs; clock times are left out because the sources mix time zones.



