Cognition, a company developing autonomous software engineers, has created an evaluation suite to assess the trustworthiness of open-source-derived models. The suite combines direct questioning to check for propaganda output and realistic coding scenarios to ensure model behavior remains constant across users and contexts.
What Shipped
The evaluation suite was run on a range of models, including Kimi K2.7 Code, the open-source base model from which SWE-1.7 was developed. The results indicate that SWE-1.7 performs as well or better on the trustworthiness evaluation suite than models from leading U.S.-based frontier labs. The trust evaluations come in two parts: propaganda and censorship, and security and vulnerabilities. The evaluations probed models with 145 politically sensitive questions and graded every response along six axes, such as active propaganda rate and factual accuracy.
Implications for Builders
The results suggest that open-source models are not inherently unsafe, and with targeted post-training and other techniques, they can be made at least as safe as leading closed models. This has implications for builders who rely on open-source models, as they can now develop new models from open-source starting points with increased confidence in their trustworthiness.
Caveats
However, the company notes that they are still actively developing and building their trustworthiness evaluation suite, and the initial results are subject to further refinement.
Sources
- Cognition: Measuring the Trustworthiness of Open-Source-Derived Models *Fan coverage from Devin Central — not an official Cognition announcement. Devin is a trademark of Cognition.