How I grade my agents
Different work earns different graders. Borrowed scoreboards where a standard already exists, plain-English rubrics where none does, probability bets on the calls that carry weight, and a running log of every time I overrule the machine.
I run a fleet of agents now, and once someone sees it working, the question is always the same: how do you know any of it is good? Fair question. An agent will hand you confident work all day. Confident and good are different claims, and only one of them is free.
So I grade them. Not with one big score, because the work is not one kind of thing. Different work earns different graders, and knowing which grader fits is most of the craft. Here are the four I actually run.
Borrow the scoreboard when reality keeps one
Some work gets graded by the world whether you like it or not. A proposal wins or loses. An outreach note gets a reply or gets silence. For that work I do not invent a rubric. I capture the scoreboard: a library of past proposals labeled with what actually happened, an outreach log that knows who answered. New work gets scored against the winners before it ships.
And the outcomes feed back in. In my discovery loop, a find that turned into a reply or a meeting counts as a double-weight good grade on the next run, and silence counts as a mild bad one. The agent learns what a good catch looks like from what happened, not from what looked exciting at seven in the morning.
Write the eval in plain English
The most useful eval I run is one paragraph long. Every hunt my discovery agent runs is a small config file, and the line that matters is plain English: what counts as signal for this hunt, and what does not. That paragraph is the whole eval. Every raw find gets judged against it, and the junk dies there. An empty day is a valid result, because a feed that pads is a feed you stop trusting.
People imagine evals as infrastructure. Start smaller. If you can write down what good means in a paragraph, you can grade against it today.
Make the agent put a number on it
The newest layer is my favorite. Any committed action in my system carries a prediction: what the agent expects to happen, with an explicit probability, inside a time window. When the window closes, the outcome gets recorded next to the bet. Every month, the predictions roll up against reality.
That is calibration, and it changes how trust works. An agent that says seventy percent and lands seventy percent gets more rope. An agent that says ninety and lands forty gets its scope cut. I stopped asking whether an agent sounds right and started tracking whether its bets pay.
Keep the human grade, but write it down
Some calls only a human can make: is it in my voice, is it the best of three, would I put my name on it. Those run through a gate with four buttons, not two. Reject. Approve. Approve and benchmark, when it beats the current best and the system should learn from it. Approve with a note, when I almost killed it for a reason worth remembering.
And every time I overrule the machine’s pick, the override lands in a log. That log is the cheapest, highest-signal training data the whole system produces. It is taste, written down, and it compounds into the graders instead of living only in my head.
Where this is going
The piece in build now is the audit ritual. Once a week, the system deals me a short queue of the calls worth a second human look: the ones where graders disagreed, the borderline calls, the cases unlike anything graded before, and a few controls to keep me honest. A dozen items, once a week, and the grading system itself gets graded.
None of this arrived as infrastructure. Each grader started as one paragraph, one labeled library, one log file. The machine underneath is The Loop Is the Product; this is how it keeps score. And if you run agents and cannot say which ones deserve more rope, start grading. The fleet gets honest fast.
Confidence is free. Grades are earned. Give the rope to the agents that hit their numbers.
