HERMOD SYSTEMS

How we work

An agent team is an institution. Design it like one.

Most agentic AI work stalls between a working demo and something that survives real users. The model is rarely what failed. What failed is that nobody designed the institution the model works inside.

Institutions hold up when three things are true of them. Below are those three conditions, what breaks without each, and how we satisfy them today. The conditions are the part we would argue for. The implementation is ours — yours will look different, and working that out is most of an engagement.

What I do know is that good intentions without structural support have a shelf life, and the shelf is shorter than we think.

No Key
  1. Condition 01

    The record has to outlive the session

    Whatever the team knows must sit somewhere no agent can quietly revise and any human can read. Without it, long sessions drift: what was true at the start stops being true, and nobody can point at the moment it changed.

    How we do it: intent, decisions and disputes live in git. Every seat is disposable — kill it mid-task and boot a successor from a context pack, which is a drawing rather than an essay. The institution remembers so no single model has to. A different shop could satisfy this with a different store; the requirement is that state lives outside the agents and stays readable without us.

    A human hand cuts deep lines into a thick clay slab.
    Cut into the record.
    The same three lines traced in loose dust, already broken up and trodden through.
    Or traced in dust, and gone by morning.
  2. Condition 02

    Whoever judges cannot be whoever built

    Independence has to be structural, not asked for politely. A reviewer that shares the builder's context shares the builder's blind spots, and will wave through exactly the things worth catching.

    Consensus is not safety. Consensus among observers who share a cognitive environment is just correlated error.

    The Seismograph Is Not Enough

    How we do it: the two seats are held by models from rival labs, so no vendor grades its own homework. That is one way to buy independence. Separate orgs, separate context windows, or a human reviewer with real authority are others. The condition is separation; the mechanism is a design choice.

    A hand slides an incised clay tablet into the slot in a robot’s head.
    A seat starts with its rules installed.
    An empty clay slot clogged with dust, a crumpled unreadable wad on the ground below.
    Or starts with nothing, and improvises.
  3. Condition 03

    Something has to be trying to break it

    Verification that is nobody's job does not happen. Left implicit, review becomes approval, and defects are found by users instead.

    Agreement is cheap. Disagreement is expensive to manufacture. Therefore, build a system that treats disagreement as the primary safety signal.

    The Seismograph Is Not Enough

    How we do it: whatever one vendor's AI builds, the other vendor's AI attacks — fresh context, one mandate. Reviewers can and do overturn the arbiter above them, which is the part that makes it real rather than ceremonial. What the adversary is matters less than that it exists, is mandated, and can win.

    A clay robot checks a course of bricks with a set square and plumb line.
    Checked against the drawing.
    A clay robot presses a brick onto a crooked, cracked course by eye with no tool.
    Or laid by eye, and cracked by Friday.

The honest part

Human headcount: one. Never out of the loop. I review every new shape of the product, I watch the governance run, and when it cracks I upgrade the rules. That is where the human hours go. The AI writes the code. The human writes the constitution — and rewrites it, which is the part people underestimate.

PLACEHOLDER — the interactive review demo (a real snippet with a real defect, and a “ship it?” prompt) goes here once the copy is settled.

Start a conversationRead the diaries