MODEL

SWE-2

modeltopic-note

Overview

SWE-2 is Cognition’s autonomous-coding model, post-trained from Moonshot AI’s Kimi K3 (2.8T) and shipped exclusively inside Devin ($20/mo Pro+ tier) with no open weights and no standalone API. Positioned by Cognition against Claude Fable 5.1 and GPT-6 Astra on FrontierCode 1.1 at ~70% lower inference cost, but trails badly on the harder Terminal-Bench 4 agentic-tool-use benchmark — a coding-model cost play rather than a frontier-agent challenger.

Timeline

  • 2026-09-11-AI-Digest — Cognition released SWE-2 on Sep 10 — a post-train from Kimi K3 (2.8T), Devin-exclusive at the $20/mo Pro+ tier, no open weights, no standalone API. Cognition-reported benchmarks put SWE-2 at 50.0% on FrontierCode 1.1 (versus Fable 5.1‘s 50.9% and GPT-6 Astra‘s 53.3%) at ~64–70% lower inference cost, but on Terminal-Bench 4 it lands 27.3% against Fable 5.1’s 55.8% and Astra’s 57.9% — a two-times gap on the harder agentic-tool-use bench. Not Cognition’s first model (SWE-1.7 exists) — the “first Cognition-branded model” framing some outlets ran is wrong. Carry as coding-cost play, not architectural novelty.

Key Developments

  1. SWE-2 launch (Sep 2026): Post-train of Kimi K3, Devin-exclusive, no open weights. Same post-train-of-open-base pattern as Cursor Composer 2.5 and Codex — application lab going frontier-model-shaped via the open-base recipe rather than a from-scratch training run. Benchmark-selection carries the story: matches Fable 5.1 on FrontierCode 1.1 within a point but trails 2× on Terminal-Bench 4.