The method is public before the numbers are.
Grace makes two claims — fewer frontier calls, so a fixed allowance carries more work, and changes that are proven rather than asserted. Neither is worth anything without a measurement anyone can repeat. This page is that measurement's method. The results are not in yet, and this page will say so until they are.
This one has not been measured yet. The method is published before the result —see the benchmark page.
This one has not been measured yet. The method is published before the result —see the benchmark page.
This one has not been measured yet. The method is published before the result —see the benchmark page.
The corpus
18 tasks against real repositories, fixed in advance and versioned. A task is a written instruction plus the repository state it starts from, so the same run can be replayed later against the same starting point.
The corpus deliberately includes tasks the local model is expected to fail. A benchmark made only of tasks that succeed measures nothing.
What counts as a passed run
A run passes when the change satisfies the task's assertions and the executed checks confirm it. A model statement never counts. Neither does a run whose effect could not be observed — that is recorded as not verified, not as a pass.
Why scoring is separate from the product
The agent's own verification decides what Grace does next. It must not decide whether a benchmark run counted. Those are two different judgements and they are kept apart on purpose: the assertions used for scoring are written with the task, are not visible to the agent, and are evaluated after the run finishes.
What is recorded per run
- Number of frontier calls, and which step triggered each one
- How many tasks finish on a fixed frontier allowance — the same allowance, once with Grace and once without
- Tokens and cost, split by local and frontier
- The verification state the run ended in
- Which checks executed, and which were skipped because the project does not declare them
- The corpus fingerprint and the configuration fingerprint of the run
What would make a result invalid
- A corpus changed between runs without a new fingerprint
- Assertions weakened after seeing a failure
- Runs excluded from the average without saying which and why
- A comparison against a different task set than the one published here
Results
No results published yet. The runs are in progress. When they finish, this table fills from a single data file — the method above will not be edited to fit the outcome.
Table structure is in place: Task · Difficulty · Finished locally · Escalated to frontier · Verification · Tasks per allowance.
Want the numbers when they land? Leave an address.