Research
Research
The evidence behind the claims made everywhere else on this site: how the agent is measured, what it is measured against and eventually what it scores. The method is published first, on purpose. No model matrix has been run yet, so there are no results here.
Not written yet
The first model matrix
Strict pass rate and cost per passing task, per tier and per model, from the first live run of the suite. Records and the generated leaderboard published alongside, so the table can be recomputed rather than trusted.
Drift measurement
Putting a number on the problem the tool exists to solve: how far a schematic and its documents diverge over the life of a design, measured across open hardware repositories.
The blog carries the argument and the repo carries the code. Both are the best place to push back on any of it.