Cognition has released FrontierCode 1.1, a refined version of its eval designed to measure code quality. The new version includes improvements to fair internet use, grading criteria, and model scores. According to the Cognition blog, the refined methodology aims to eliminate unfair internet use while preserving the realism that internet access provides.
What Shipped
The FrontierCode 1.1 release includes several key improvements. The methodology for fair internet use has been refined to capture the nuance between legitimate internet use and unfair use. The grading criteria have been audited, and 75 overly strict blockers have been demoted to non-blocker status. New model scores have been released for Sonnet 5 and updated scores for Fable 5. The results of the FrontierCode 1.1 Main eval show that the relative performances of the models did not substantially change compared to the previous version.
Implications for Builders
The improved methodology for fair internet use is expected to reduce the occurrence of unfair internet use, while still allowing agents to look up documentation and other relevant information. The refined grading criteria are also expected to reduce the occurrence of false negatives in grading.
Caveats
The deprecation of the FrontierCode Diamond set may affect the comparability of results between the old and new versions.
Sources
- Cognition Blog: FrontierCode 1.1 *Fan coverage from Devin Central — not an official Cognition announcement. Devin is a trademark of Cognition.