How we work
An agent team is an institution. Design it like one.
Most agentic AI work stalls between a working demo and something that survives real users. The model is rarely what failed. What failed is that nobody designed the institution the model works inside.
Institutions hold up when three things are true of them. Below are those three conditions, what breaks without each, and how we satisfy them today. The conditions are the part we would argue for. The implementation is ours — yours will look different, and working that out is most of an engagement.
What I do know is that good intentions without structural support have a shelf life, and the shelf is shorter than we think.
No Key
-
Condition 01
The record has to outlive the session
Whatever the team knows must sit somewhere no agent can quietly revise and any human can read. Without it, long sessions drift: what was true at the start stops being true, and nobody can point at the moment it changed.
How we do it: intent, decisions and disputes live in git. Every seat is disposable — kill it mid-task and boot a successor from a context pack, which is a drawing rather than an essay. The institution remembers so no single model has to. A different shop could satisfy this with a different store; the requirement is that state lives outside the agents and stays readable without us.

Cut into the record. 
Or traced in dust, and gone by morning. -
Condition 02
Whoever judges cannot be whoever built
Independence has to be structural, not asked for politely. A reviewer that shares the builder's context shares the builder's blind spots, and will wave through exactly the things worth catching.
Consensus is not safety. Consensus among observers who share a cognitive environment is just correlated error.
The Seismograph Is Not EnoughHow we do it: the two seats are held by models from rival labs, so no vendor grades its own homework. That is one way to buy independence. Separate orgs, separate context windows, or a human reviewer with real authority are others. The condition is separation; the mechanism is a design choice.

A seat starts with its rules installed. 
Or starts with nothing, and improvises. -
Condition 03
Something has to be trying to break it
Verification that is nobody's job does not happen. Left implicit, review becomes approval, and defects are found by users instead.
Agreement is cheap. Disagreement is expensive to manufacture. Therefore, build a system that treats disagreement as the primary safety signal.
The Seismograph Is Not EnoughHow we do it: whatever one vendor's AI builds, the other vendor's AI attacks — fresh context, one mandate. Reviewers can and do overturn the arbiter above them, which is the part that makes it real rather than ceremonial. What the adversary is matters less than that it exists, is mandated, and can win.

Checked against the drawing. 
Or laid by eye, and cracked by Friday.
The honest part
Human headcount: one. Never out of the loop. I review every new shape of the product, I watch the governance run, and when it cracks I upgrade the rules. That is where the human hours go. The AI writes the code. The human writes the constitution — and rewrites it, which is the part people underestimate.
PLACEHOLDER — the interactive review demo (a real snippet with a real defect, and a “ship it?” prompt) goes here once the copy is settled.