Research
Research
Nothing here yet. This is where the evidence will go: the benchmarks, the methodology behind them, and the raw results, so that the claims made everywhere else on this site can be checked.
Benchmarks
How the agent scores on real boards rather than toy ones: pass rate on ERC and DRC, how often a change lands without human repair, and what a run costs in tokens.
Drift measurement
Putting a number on the problem the tool exists to solve: how far a schematic and its documents diverge over the life of a design, measured across open hardware repositories.
Evaluation method
The harness, the fixtures, and the scoring, published so the numbers above can be reproduced and argued with.
In the meantime the blog carries the argument and the repo carries the code. Both are the best place to push back on any of it.