Editorial cover for ai story

Cognition SWE-2 Posts 50% on In-House Coding Benchmark, Trails Rivals on Terminal-Bench 4

Cognition released SWE-2 on September 10, 2026, positioning the new model as the engine behind its Devin coding agent and publishing benchmark results that place it within striking distance of frontier competitors on the company’s own evaluation suite. The headline number is 50.0% on the Cognition SWE-2 frontier coding benchmark, dubbed FrontierCode 1.1 Main, putting the model 0.9 points behind Anthropic’s Claude Fable 5.1 at 50.9% and 3.3 points behind OpenAI’s GPT-6 Astra at 53.3%. Cognition also claims SWE-2 reaches that score at 64% less cost than Fable 5.1, a figure the company is leaning on heavily in launch coverage.

Cognition SWE-2 frontier coding benchmark: The Model and the Numbers

SWE-2 is a post-trained variant of Moonshot AI’s open-source Kimi K3, a mixture-of-experts model with 2.8 trillion parameters. Cognition handled the post-training pipeline in-house, and the company says it scaled reinforcement learning to the multi-trillion-parameter regime for what it claims is the first time. The training recipe includes a cost-penalised reward function designed to reward not just correctness but efficiency of token and tool use. Compared with the previous SWE-1.7 release, Cognition reports SWE-2 uses 58% fewer turns and 81% less cost to complete equivalent coding tasks.

The FrontierCode 1.1 Main benchmark is built around real-world software engineering tasks drawn from open-source repositories. Cognition says it co-developed the suite with more than 20 open-source maintainers and that the evaluation covers bug fixing, feature implementation, code review, and multi-file refactoring. On that benchmark, SWE-2’s 50.0% places it firmly in frontier-adjacent territory, just outside the top tier.

The Caveat: A Benchmark Cognition Wrote

The catch, flagged by multiple early analysts, is that FrontierCode 1.1 Main is Cognition’s own benchmark. When the model is measured on Terminal-Bench 4, an independent third-party coding evaluation, the picture changes sharply. SWE-2 scores 27.3% on Terminal-Bench 4 versus 55.8% for Claude Fable 5.1 and 57.9% for GPT-6 Astra. That is a gap of roughly 28 percentage points on a level playing field. The dispersion between the two evaluations is large enough that the FrontierCode 1.1 Main numbers should be read as a directional signal rather than a definitive ranking.

Artificial Analysis, one of the more cited third-party evaluators of frontier coding models, has not yet added SWE-2 to its index. No independent lab has published a head-to-head result against the September 10 release as of press time, which means the Terminal-Bench 4 delta is currently the cleanest external reference point available. Coverage from launch day has generally framed SWE-2 as near-frontier performance at a fraction of the cost, but on a benchmark Cognition itself designed.

Availability and Access

SWE-2 ships exclusively inside the Devin product family. The model is live in Devin Desktop, the Devin command-line interface, Devin Web, and Devin Fusion, the orchestration layer that routes tasks across multiple agents. There is no public API, no downloadable weights, and no self-serve tier for external developers. Enterprises wanting to benchmark SWE-2 against their own codebases need to do so through one of the Devin interfaces, which limits the kind of apples-to-apples testing that produced the FrontierCode 1.1 Main numbers in the first place.

Pricing for Devin access was not detailed in the launch post beyond the cost-comparison claim against Fable 5.1. Cognition has historically sold Devin seats on a per-user or per-task basis, and the 64% cost reduction figure refers to internal task-completion cost rather than sticker price. Anyone evaluating SWE-2 against Claude Fable 5.1 or GPT-6 Astra on raw economics will need to wait for clearer per-task pricing from Cognition or for independent cost studies.

What the Launch Tells Us About the Frontier

The SWE-2 release is consistent with a pattern visible across 2026: post-trained open-weight bases are closing the gap with closed frontier models on narrow task suites. Kimi K3 from Moonshot AI is one of several open bases, including Llama 4 Behemoth and DeepSeek V4, that smaller labs are now post-training for coding-specific workloads. Cognition’s contribution is the cost-penalised reward and the RL scaling claim, both of which would be technically meaningful if reproduced independently. The FrontierCode 1.1 Main result is best read as evidence that the post-training pipeline works on its own evaluation, not as a standalone ranking claim.

The bigger story may be the Terminal-Bench 4 gap. A 28-point spread between an in-house suite and a third-party benchmark is large enough to suggest that FrontierCode 1.1 Main over-represents the kinds of tasks SWE-2 was optimised on, or that Terminal-Bench 4 emphasises something SWE-2’s training did not target. Either way, the Cognition SWE-2 frontier coding benchmark result is paired with enough asterisks that buyers should treat the 50.0% figure as one data point among several, not as a verdict. Until Artificial Analysis, SEAL, or another independent evaluator adds SWE-2 to its leaderboard, the Terminal-Bench 4 delta is the number to watch.

Source: https://www.orcarouter.ai/blog/cognition-swe-2-release

Leave a Comment

Your email address will not be published. Required fields are marked *