copperhead

Research

Research

The evidence behind the claims made everywhere else on this site: how the agent is measured, what it is measured against and eventually what it scores. The method is published first, on purpose. No model matrix has been run yet, so there are no results here.

Not written yet

  • The first model matrix

    Strict pass rate and cost per passing task, per tier and per model, from the first live run of the suite. Records and the generated leaderboard published alongside, so the table can be recomputed rather than trusted.

  • Drift measurement

    Putting a number on the problem the tool exists to solve: how far a schematic and its documents diverge over the life of a design, measured across open hardware repositories.

The blog carries the argument and the repo carries the code. Both are the best place to push back on any of it.