Every night, a chain I built collects whatever new documents have landed in a fund’s records, has an agent read them and decide what each one asserts, tests that against everything already on the books, and — when nothing already settled moves — writes it into the register itself. Most nights it finishes without me. Some nights it stops and leaves a note.
It exists because of a moment I reached as a partner in a venture fund, and I doubt I am alone in it. The work gets done — more or less right, often painfully, always slowly — and one day you conclude you need to change the tool. What I wanted was easy to say and hard to buy: a system whose picture of the business is tied to the real documents, without depending on anyone to key them in or point at them. Every tool I looked at asks people to do the work, and then to narrate the work into the tool. That second job is where the fragility lives — when the week is busy, it is the part that does not get done. I wanted something that blends into the job instead of sitting on top of it.
So I went and looked at what is sold. Every quote comes with two costs it does not mention: owning the thing, and running it. Large organisations can afford that; I had the question, not the department. AI turns up in exactly that gap, promising the outcome without the project.
Nothing here replaces anything. Records like these are kept the way most organisations keep them — by people, document by document, correctly, and slowly. I wanted to see what a chain of agents would do with the same paperwork, running beside them.
A few days ago it stopped on a signed custodian statement. It had read the document correctly. It had resolved the holder without ambiguity. It stopped because the register — the one book everything else gets checked against — had no way to carry what the document said, and because I had forbidden it from deciding that on its own.
That block was not a defect. It was the design. I did not have to build it this way: I could have shipped an agent on its own, arbitrating from a prompt. It would have passed the first tests, passed the demo, and broken in production. I know, because I ran the chain without me, in a sandbox, to see what it would do unattended. Same documents, same protocol. Then I gave it the same ten documents twice, twenty minutes apart: four facts the first time, seven the second. I did not need an eleventh document — the variance was already there. The chain I actually run has since put hundreds into the register.
So I shipped the other one. And ever since, I have been the man inside the Mechanical Turk — the 1770 chess automaton with a man folded into the cabinet. It makes the moves. I choose the ones it cannot.
The rule fits in one sentence
As long as it does not call the past into question, it validates without me.
That sentence draws a line through the chain. Everything upstream of it is agentic: the agent reads a document, infers what happened, and produces a structured list of assertions. Everything downstream is not. The chain takes that list and confronts it, deterministically, with the register as it stands — and with a thousand-odd checks I wrote by hand as I built it: sums that must tie, figures that cannot contradict what is already on the books. It runs that confrontation on a copy before it writes anything anywhere.
That was my opening design guess, written before the chain had ingested a single document — not something an accident taught me later. It has held through every document since, and it turned out to be the cheapest protection in the system: nothing that would disturb a settled figure gets written while I am not looking.
That is the whole point. The split between what the machine settles and what a human settles is not handed to you by the state of the art. It is a decision, and it is yours.

One document in twelve still needs a person
Every night the chain sweeps close to three thousand documents across the firm’s records. Over seven weeks, a couple of hundred of them changed something in the register. About twenty stopped and asked for a person.
That is roughly one document in twelve that the machine would not close on its own. I did not see that coming when I started — it landed on me halfway through the build. And not one of those decisions has yet turned into a rule that closes the next one, which means the whole thing runs at the speed of my attention. That is a throughput problem, not a design achievement.
Almost none of the stops were reading errors. The agent knows the trade — what a contract says, what a signature commits. That part is in the model, and it covers most of what lands.
It stops when the answer stops being generic. Is this document still the one in force here? Does this one replace that one? Is this an error in the paperwork, or a piece of context nobody ever wrote down? None of those are questions about documents. They are questions about one organisation, and the model has never been near it.
The custodian statement stopped there. Nothing in the file said what to do, and nothing in the weights ever could. It did not guess, because it is not allowed to. Compliance did what a machine cannot — checked, arbitrated, had the document re-issued carrying what it should have carried in the first place. The chain took it without a word from me and carried on.
One of those questions is harder than the rest. The agent knows what an amendment is, can read it, can derive what it changes. It has no rule for putting it into force. That rule exists in no form at all — not code, not prompt, not paper. Nobody ever wrote it.
There are two ways an answer becomes binding: a document settles it, or someone does. The second is not reading and it is not reasoning. It is standing — a position, not a skill.
That is what a person does when they key data into a system. Most of it is transcription; some of it is a call, made at the speed of the typing and recorded nowhere. Every set of records is already full of them. The agent’s only difference is that it cannot make one silently — not out of scruple, it has none, but because the rule sits in the code. In a prompt, it would have taken the leap.
That is not a gap in the tool. It is a gap in the governance — and the tool is what made it visible. Which is also why the line does not move when the models get better.
But on average, humans and AI together do worse than the better of the two alone
The question is not whether to put a human in the loop. It is whether, at the exact point where you put one, the human knows something the machine does not.
Others have measured that better than I could, so I looked — starting with work that could prove me wrong.
A 2024 review pooled 106 experiments, published between 2020 and 2023, that measured people alone, AI alone, and the two together. You would expect the pair to beat its best member. On average, it did slightly worse — especially when the task was a decision.
The average hides a split. When the machine was already better at the task, adding a person made the result worse. When the person was better, the pair beat both. The authors’ best guess: a person who is better than the machine is also better at telling when to trust it. One who is not leans on it when it is wrong and ignores it when it is right.
In my register, there is no doubt which one to trust. On can this document supersede that one, in this house, the machine has no rule at all. I am not marginally better there. I am the only one holding the answer.
Almost every one of those experiments had a person make the final call, case by case, after seeing the machine’s answer. My chain is not built that way. The split is set in advance: the machine settles what passes its checks, and only what does not comes to me. That is the arrangement the authors point to as the next thing to test. Only three of their experiments tried it — too few to conclude anything.
And the same question tells me where my design would be wrong: anywhere the model already does it better than I do.
Curriculum, manufactured in advance
Both roads answer the same question: how does expert judgment get into a system? Mine collects it, one arbitration at a time, as the work throws them up. There is another, and it is not hand-waving: you can manufacture it in advance.
In early September, a16z argued that the durable advantage in vertical AI is not owning the record but learning the job — and that memory is not learning. Keeping context tells you what was decided last week. It does not tell you whether the work was good, why an expert changed it, or what to do differently next time.
Their example is Harvey, and Harvey documented it themselves. They post-trained an open-weight model via asynchronous reinforcement learning across roughly 1,750 simulated legal task environments, each one a matter handed down by a partner, graded against an expert rubric averaging around fifty criteria. And, in their own words:
We did not use any customer data in any of our post-training efforts.
— Harvey, post-training update
They did not wait for production to hand them arbitrations. They manufactured them, in bulk, before shipping.
It works, and I think it settles one half of the problem. The half it cannot settle is named in the a16z piece itself: there are two curricula, not one. There is the profession — how a good practitioner does the work — and there is the institution: this house’s templates, precedents, risk thresholds, escalation rules. Lessons about the profession improve the product for everyone. Lessons about one firm improve it for that firm only.
A simulation can manufacture the first. It cannot manufacture the second. The institution’s curriculum exists nowhere except in the arbitrations its own authorities render, on its own documents, in production. That second curriculum is the context gap I described in the first piece — this time at the level of a house, not a person — and this is the only place it gets produced.
It does not move onto the machine
Everything above is a description of where I chose to put that authority, in a house I built the chain for. Whether it can be moved onto the machine at all is not for whoever builds the thing to decide. It is decided by whoever has the standing — a committee, a regulator, a court — and one small case has already made the answer public.
In 2024, a Canadian tribunal heard a passenger who had relied on an airline’s chatbot for a fare policy the chatbot had invented. The airline argued that the chatbot was a separate entity, responsible for its own actions.
The tribunal refused. The chatbot was part of the airline’s website; the airline answered for what was said on it.
What matters is the attempt: a company tried to move authority onto the machine, and a court declined to let it. You are only ever choosing where it sits — or discovering where it sat, in a dispute.
One number will not tell you how you are doing
You need two, and neither is about a person — both are about a position.
The first is human frontier — how often it needs a human authority. Of the cases that carry a stake, how many come back to someone with the standing to settle them. Here, about twenty out of a couple of hundred — one in twelve.
The second is learning impact — what a human arbitration buys, counted in future cases one decision closes on its own. Here, zero. Nothing I have arbitrated has yet become a rule that covers the next case. The chain records my answer, validates it, keeps it. It is not equipped to generalise from it, because I forbade it from judging.
The couple is what speaks, never either alone. Low on both is a sealed system — a pacemaker, an avionics stack, a satellite: no authority can be there while it runs, so every judgment is frozen before it leaves. Low frontier with high impact is new software — the arbitrations have been rendered and have become rules the thing applies on its own. High on both would be a copilot, asking often and keeping what it is told. That may be a form factor still arriving; I have not seen one in wide use.

Which leaves the corner where it asks constantly and nothing stays: the forever assistant. That is where my chain sits, measured. The other marker on that map is the same work once my decisions have become rules, and the arrow between the two is the build ahead.
The frontier is the point
The machine does what it does better and far faster than I do: find a document, read it, extract what it asserts, confront it with a rule, and record it. It stops where we know a human holds the decision. The work ahead is to move that frontier — carefully, and in a way that survives inspection.
Which is to say the generalisation has to come, but its application has to stay determined, traced, logged and auditable. Not for the sake of compliance theatre. Because the day I stop being present at every step, the only thing that still gives me authority is the ability to come back on a decision I was not there to make.
So here is the falsifiable version. In six months, if the chain has no auditable learning of its own, and the share of cases that come back to me has not fallen, then I was wrong — not about the design, but about my ability to build past it. And I will have built a very expensive Mechanical Turk.
Sources
Vaccaro, Almaatouq & Malone, When combinations of humans and AI are useful: a systematic review and meta-analysis, Nature Human Behaviour, 2024 — 106 experiments from 74 papers, 370 effect sizes.
Harvey, post-training update — asynchronous reinforcement learning, ~1,750 agentic legal task environments, ~50 rubric criteria per task, no customer data.
a16z, The Incumbents Are Coming, September 2026 — memory is not learning; the profession’s curriculum and the institution’s are not the same curriculum.
Moffatt v. Air Canada, 2024 BCCRT 149 — the tribunal rejected the argument that the chatbot was a separate entity responsible for its own actions.


