Cognition SWE-2 Explained: Frontier Code Benchmark Results and Cost Trade-offs
The headline circulating since 10 September 2026 is tidy: a coding model that scores 50 percent on a hard benchmark and costs roughly two-thirds to 70 percent less than the frontier. Cognition SWE-2, the in-house model behind the Devin coding agent, is being passed around in exactly those terms. Read Cognition’s own announcement and the story gets more interesting and less comfortable. The 50 percent is a score on a benchmark Cognition itself created. The cost figure is a task-level curve, not a published price. And on one harder terminal benchmark, the model scores less than half of its rivals.
That matters now because coding agents are being bought on exactly this kind of chart. A procurement decision that rests on a vendor-authored benchmark, a vendor-defined cost metric, and a vendor-selected set of comparison rows is a decision made on very little independent evidence. This post separates what is stated by Cognition from what has been independently checked (almost nothing so far), explains why coding benchmarks mislead even when everyone is honest, and ends with a concrete method for building an in-house evaluation that tells you whether SWE-2 is cheaper for your codebase.
What this covers: the SWE-2 lineage and reported architecture, the exact benchmark and cost claims with their caveats, how Cognition’s cost-penalized reinforcement learning works, how FrontierCode, DeepSWE and Terminal-Bench differ, contamination and harness effects, a decision matrix against the peers Cognition named, and a step-by-step in-house evaluation plan.
Context and Background
Cognition is the company behind Devin, the autonomous software engineering agent that launched in 2024, and its own SWE model family, which sits inside the Devin product. SWE-2 is the successor to SWE-1.7, and its distinguishing feature is not raw capability but the efficiency framing in the title of the announcement, “Pushing the Pareto Frontier”. A Pareto frontier here means the curve of best achievable score for each level of cost: a model is on the frontier if no alternative is both cheaper and better. The claim is that SWE-2 moves that curve, delivering near-frontier scores at a fraction of the spend.
For practitioners, the cost axis has become as important as the accuracy axis. A coding agent is not one model call. It is a loop of dozens or hundreds of steps, each of which re-reads context, calls tools, edits files and runs tests. Cost per resolved task is the product of tokens per step, steps per task, price per token, and the retry rate. A model that is slightly less accurate but takes far fewer steps can win on total cost, and that is the specific bet SWE-2 makes. Our earlier look at how agent benchmarks such as SWE-bench, GAIA and tau-bench are built and where they leak covers the general mechanics; this post applies them to one concrete release.
The credibility problem is also general. In February 2026 OpenAI published an analysis explaining why it no longer evaluates on SWE-bench Verified: it reported that 59.4 percent of a 138-problem audited subset contained material issues in test design, and that every frontier model it tested showed signs of contamination, meaning the model could reproduce the original human-written fix or verbatim problem details from training data. OpenAI recommended reporting SWE-bench Pro instead. You can read the full write-up in OpenAI’s “Why we no longer evaluate SWE-bench Verified”. The lesson is not that benchmarks are useless. It is that every public coding benchmark has a half-life, and that a vendor-built benchmark has an additional property worth stating plainly: the author is also a competitor.
Finally, a note on sourcing. Cognition’s announcement at cognition.com/blog/swe-2 is the primary source for everything below. Secondary coverage, including a few independent write-ups that re-tabulate the same numbers, adds arithmetic and commentary but not new measurements. Where a figure is derived by a third party rather than published by Cognition, this post says so.
What Cognition SWE-2 Is, Reportedly
In short: Cognition SWE-2 is a coding-agent model released on 10 September 2026, built by applying large-scale reinforcement learning to Kimi K3, a 2.8-trillion-parameter base model. Cognition reports 50.0 percent on FrontierCode 1.1 Main, within one point of Fable 5.1, at 64 percent lower cost. It ships in Devin; no open weights, API or model card are published.

Figure 1: How Cognition describes SWE-2, from a Kimi K3 base through cost-penalized reinforcement learning to evaluation on self-authored benchmarks.
Figure 1 summarizes the pipeline as Cognition describes it. A very large mixture-of-experts base model is post-trained with reinforcement learning across a tripled set of coding environments, with a reward that subtracts a cost term from a binary success signal. The resulting model is exposed at several reasoning-effort levels, evaluated on benchmarks that Cognition runs, and shipped inside Devin. Every box after the base model is controlled by one company, which is the single most important structural fact for a reader deciding how much to trust the results.
Lineage: from SWE-1.7 to SWE-2
Cognition’s comparison to its own predecessor is the least contested claim in the announcement, because it compares like with like on the same harness. On FrontierCode 1.1 Main, Cognition reports SWE-1.7 at 42.0 percent and SWE-2 at 50.0 percent, an 8-point gain. It further reports that SWE-2 at medium effort scores higher than SWE-1.7 while taking 58 percent fewer turns and costing 81 percent less on average, with mean steps per run falling from 127 to 53 and the median step of the first edit moving from 48 to 18.
Those last two numbers are the most informative in the post, because they describe behaviour rather than outcomes. A model that makes its first edit at step 18 instead of step 48 is spending less time exploring and more time acting. That is consistent with a reward that penalizes cost: an agent that wanders through a repository burns tokens, and the training signal teaches it not to. Whether that behaviour generalizes to your repository, with its own build system and conventions, is the open question.
The base model and what is reported about it
Cognition states the base is Kimi K3 with 2.8 trillion parameters, and secondary coverage reports it as a mixture-of-experts design with roughly 104 billion active parameters per token. The 104 billion active-parameter figure comes from that secondary coverage, not from a Cognition model card, so treat it as reported rather than confirmed. Cognition describes the training as reinforcement learning “scaled to the multi-trillion-parameter regime for the first time”, which is a claim about its own engineering, not an independently verifiable fact.
What is not disclosed matters as much. There is no published context window, no tokenizer detail, no per-token price, no model card, and no standalone API at launch. Availability is Devin Desktop and the Devin command-line interface from day one, with Devin Web and Fusion rolling out afterwards. That means the only way to try SWE-2 is through Cognition’s own harness, which has consequences for any independent comparison, discussed below.
How cost-penalized reinforcement learning works
The technical contribution Cognition highlights is a reward function that bakes cost into training. In the post’s notation, the reward for a rollout is R = S minus a cost term, where S is 1 if the task succeeded and 0 otherwise, C is the cost of the rollout, and a coefficient set per effort level scales how much cost matters. The coefficient is tuned, Cognition says, to match the slope of the Pareto curve of the base model. In plain terms: at each effort level the model is told how many points of success probability one unit of cost is worth, using the base model’s own trade-off curve as the exchange rate.
This matters because the usual alternative is to train for success only and then control cost afterwards, by capping steps or truncating context. A success-only reward produces agents that keep trying: if another hundred steps raise the chance of a pass by one percent, the reward favours taking them. A cost-penalized reward makes that behaviour unprofitable. A single training run covers all effort levels rather than training separate models, which is why one model can offer low, medium and high settings along one curve.
Cognition also describes a length-weighted baseline for variance reduction, computed as the sum of each rollout’s reward times its length divided by the sum of the lengths. Policy-gradient methods subtract a baseline from the reward to reduce noise; when rollouts vary enormously in length, as agentic coding runs do, a plain average over-weights short runs. Weighting by length is a sensible fix, and Cognition reports it stabilizes training. It is a training detail, not a result you can verify from outside.
The Benchmark Claims, Line by Line
Cognition’s announcement reports six comparison rows. The table below reproduces them as reported by Cognition and re-tabulated by third-party coverage. All figures are Cognition’s own runs. None has been independently replicated as of this writing.
| Benchmark | SWE-2 | Fable 5.1 | GPT-6 Astra | SWE-1.7 |
|---|---|---|---|---|
| FrontierCode 1.1 Main | 50.0% | 50.9% | 53.3% | 42.0% |
| DeepSWE 1.1 | 73.0% | 67.4% | 74.1% | not reported here |
| Terminal-Bench 2.1 | 92.8% | 91.4% | 89.9% | not reported here |
| Terminal-Bench 4 | 27.3% | 55.8% | 57.9% | not reported here |
Two cautions about the table. First, the Fable 5.1 and GPT-6 Astra figures for DeepSWE and Terminal-Bench come from a secondary write-up that tabulated Cognition’s chart; Cognition’s page states the FrontierCode, Terminal-Bench 2.1 and 4, and GPT-6 Astra DeepSWE comparisons in text; the remaining cells should be checked against the original charts before you quote them. Second, Cognition notes that its charts use each model’s best score across reasoning-effort settings and omit the highest-effort variants of the Fable models for readability, which affects how the comparison looks.

Figure 2: The shape of the reported results. SWE-2 sits within a point or two of the frontier on three rows and far behind on the harder terminal benchmark.
Reading the pattern
Three of the four rows say the same thing: SWE-2 is at or within about three points of the best rival. On Terminal-Bench 2.1 it is reported ahead, though secondary coverage notes that suite is close to saturated across vendors, which means differences near the ceiling say little. The fourth row, Terminal-Bench 4, says something quite different. At 27.3 percent against 55.8 and 57.9 percent, SWE-2 scores less than half of either rival. One independent write-up pointed out that the headline of the announcement leaves this row out.
The honest reading is that SWE-2 looks strong on the kind of work FrontierCode measures, which is producing a mergeable patch for a single issue in a repository, and weak on long-horizon terminal tasks where an agent must plan across many dependent operations. That is a plausible consequence of the training target. Cost-penalized reinforcement learning on short-to-medium repository tasks rewards finishing quickly; it would not obviously help on tasks where the right strategy is to spend more.
The cost claim, and three different numbers
The cost headline appears in three forms, and they are not the same claim. The announcement itself says SWE-2 is within one point of Fable 5.1 while being 64 percent cheaper. A social-media post from Cognition reportedly said “up to 70 percent lower cost”, which is a ceiling across some set of settings, not a typical figure. And the 81 percent figure refers to a different comparison entirely: SWE-2 medium against SWE-1.7, on average. If you see “70 percent cheaper than the frontier” in a newsletter, the primary source supports 64 percent on one benchmark, with 70 percent as an upper bound in a promotional post.
Cognition states cost is computed at list pricing, including public discounts. It does not publish a per-token price for SWE-2 or a per-task cost. A secondary write-up used a Fable 5.1 medium reference of 3.28 dollars per task and back-calculated roughly 1.18 dollars per task for SWE-2. Treat that as derived, not published: it applies a vendor percentage to a reference figure and inherits every assumption in both.
An illustrative calculation shows why the metric deserves scrutiny. Suppose, purely hypothetically, a rival model resolves 51 of 100 tasks at 3.28 dollars each, and SWE-2 resolves 50 of 100 at 1.18 dollars each. The cost per resolved task is 3.28 divided by 0.51, about 6.43 dollars, versus 1.18 divided by 0.50, about 2.36 dollars. That is a 63 percent reduction in cost per resolved task, consistent with the headline. But the picture changes if failed attempts need human cleanup: if each failure costs 20 dollars of engineer time, the rival’s expected cost per task is 3.28 plus 0.49 times 20, or 13.08 dollars, and SWE-2’s is 1.18 plus 0.50 times 20, or 11.18 dollars. The saving falls to about 15 percent. These numbers are illustrative only, but the structure is real: when human review and rework dominate, model cost is a small term, and one point of accuracy is worth more than it looks.
How These Benchmarks Work, and How They Mislead
To judge the numbers you need to know what each benchmark measures and who controls it. The three families in the announcement differ in design, and the differences explain the pattern in the results.
FrontierCode: mergeability, authored by the vendor
FrontierCode asks whether an agent’s patch for a real open-source issue is good enough to merge, not merely whether it passes tests. According to the benchmark’s published description, each task pairs a repository with one issue; the agent works autonomously in a container; and the patch is graded against held-out tests plus maintainer-written rubrics covering behavioural correctness, regression safety, build and style cleanliness, and adherence to project conventions. Scores are aggregated as mean over five runs against weighted rubrics, and a failing blocker criterion zeroes the task. Version 1.1 has a Main subset of 100 hardest tasks nested inside an Extended set of 150.
Cognition developed FrontierCode. Independent commentary reports it was built with more than 20 open-source maintainers across 36 repositories, with each task reportedly representing over 40 hours of expert work, and that the tasks are not public. Those are strengths: rubric grading is closer to code review than a pass-or-fail test, and private tasks resist training-set leakage. They are also the source of the main caveat. A reader cannot inspect the tasks, the rubric grading involves judgment, and the organisation that built the exam is the same one reporting the score.
None of this implies bad faith. It describes a structure in which the benchmark was designed, in good faith, around the strengths of a product line the author sells. Any author’s benchmark carries that risk, which is why independent replication is the thing to wait for.
DeepSWE and Terminal-Bench
DeepSWE 1.1 is described in secondary coverage as targeting deeper repository work beyond surface-level changes. I could not find a primary methodology document for it, so I will not characterize its scoring beyond that. Terminal-Bench is a suite of command-line agent tasks, where the agent must operate a shell to accomplish goals such as building software, configuring services or debugging an environment. Version 2.1 is reported to be near saturation across vendors, while version 4 is described as harder and longer-horizon. The gap between SWE-2’s 92.8 percent and 27.3 percent on those two is the largest internal inconsistency in the announcement, and it is worth taking seriously: it suggests the model’s strength is task-length specific.

Figure 3: Where a reported coding score can drift away from real-world capability, from task selection through harness to grading.
Four ways a coding score drifts from reality
Figure 3 traces the path from a task in a benchmark to the number on a chart, with four points where distortion enters.
Contamination. If a model has seen a task or its fix during training, the score measures memorization. OpenAI’s SWE-bench Verified audit is the clearest public example: all frontier models it tested reproduced original human-written fixes or verbatim problem details. Private, held-out tasks like FrontierCode’s reduce this risk, and FrontierCode 1.1 reportedly zeroes runs flagged for consulting solution-bearing sources on the internet. But contamination is hard to rule out for any model trained on web-scale data, and an outside party cannot audit a private set.
Harness effects. An agent score is a model-plus-harness score. The harness decides which tools exist, how context is pruned, how many retries are allowed, and how test output is shown. The same weights can move by several points between harnesses. SWE-2 is available only inside Devin, so its FrontierCode result is a Devin-harness result. Rival models in Cognition’s chart may have been run in Cognition’s harness as well, or in their vendors’ own; the announcement’s methodology matters here and a reader should check it.
Best-of effort selection. Cognition reports each model at its best-performing reasoning-effort setting. That is a legitimate way to draw an upper envelope, but it means the chart shows peak capability, not what you get at a default setting, and cost at the peak setting is typically far above cost at medium. A cost claim that mixes a medium-effort price with a peak-effort score would overstate the benefit, so insist on seeing which setting generated which point.
Grader variance. Rubric grading that involves model or human judgment adds noise. With a Main subset of only 100 tasks, a difference of one point is one task. Cognition’s 50.0 versus 50.9 is not a statistically meaningful gap on a sample that small, which cuts both ways: it supports “roughly equal” but it also means the 64 percent cost advantage is the real claim, because the accuracy difference is within noise.
Why aggregate scores hide the distribution
A mean over 100 tasks hides which tasks a model fails. Two models with the same 50 percent can have almost disjoint solved sets: one strong on small bug fixes in popular frameworks, the other on cross-file refactors in unusual build systems. For an engineering organisation that matters more than the mean, because your workload is a particular mix of task types. A benchmark average is a prior; your own task distribution is the likelihood. It is the reason an in-house evaluation, described later, is not optional for a team spending real money on a coding agent.
Cost Per Resolved Task: The Metric That Matters
Cost claims in this category deserve a dedicated look, because the unit changes the answer. There are at least four plausible denominators, and vendors choose the one that flatters.

Figure 4: Components of cost per resolved task, from per-step tokens through retries to the human review that follows an agent’s patch.
Decomposing the number
Per-task cost is tokens per step times price per token times steps, summed over attempts. SWE-2’s reported edge comes from fewer steps, 53 versus 127 on average against its predecessor, and not necessarily from cheaper tokens; Cognition does not publish a per-token rate. A step-count reduction is a durable advantage because it compounds: fewer steps means less accumulated context re-read at each later step, so cost falls faster than linearly. Agent context grows with every tool result, and a step late in a long run costs far more than an early one. Cutting the run from 127 to 53 steps therefore saves more than the 58 percent the step count alone suggests, which is consistent with the reported 81 percent saving versus SWE-1.7.
Where the savings can vanish
First, retries. If a cheaper model fails more often and you rerun it, expected cost per resolved task rises. Second, escalation: if a team routes hard tasks to a stronger model after SWE-2 fails, the cheap attempt becomes a pure add-on cost for those tasks. Third, review burden. A patch that is merely test-passing but not mergeable costs a human reviewer time, and that time dwarfs inference cost, as the earlier illustrative calculation showed. Fourth, pricing form. SWE-2 is, at launch, offered inside Devin subscriptions rather than per token, and reportedly free for Pro, Max and Teams subscribers for a promotional month ending around mid-October 2026. For subscribers the marginal inference cost is zero during that window, so the 64 percent figure is about Cognition’s cost to serve and list-price equivalence, not your invoice.
The relationship between model price and total cost of ownership echoes a theme in our analysis of small versus large language models on agentic tasks: the cheapest model per token is rarely the cheapest per completed job, and the right comparison is always end to end.
How SWE-2 Compares to the Peers Cognition Named
Cognition’s comparison set is Fable 5.1 and GPT-6 Astra. I have not independently verified specifications or prices for either, so the matrix below uses only what Cognition’s chart reports for each, plus qualitative notes on what is unknown. Treat it as a decision aid for what to test, not a verdict.
| Use case | SWE-2 | Fable 5.1 | GPT-6 Astra |
|---|---|---|---|
| Single-issue patch in a repo (FrontierCode style) | 50.0%, reported cheapest | 50.9%, reported about 3x SWE-2 cost | 53.3%, reported about 4x SWE-2 cost |
| Deep multi-file repo work (DeepSWE) | 73.0% | 67.4% | 74.1% |
| Saturated terminal tasks (Terminal-Bench 2.1) | 92.8% | 91.4% | 89.9% |
| Long-horizon terminal tasks (Terminal-Bench 4) | 27.3% | 55.8% | 57.9% |
| Access outside the vendor’s product | None at launch | Not covered in this post | Not covered in this post |
The cost multiples in the first row are derived from Cognition’s statements that SWE-2 is 64 percent cheaper than Fable 5.1, which implies Fable at about 2.8 times SWE-2’s cost, and about a quarter of GPT-6 Astra’s cost, which implies roughly four times. They apply to the FrontierCode comparison and nothing else.
Read as a routing guide: SWE-2 is a reasonable candidate for high-volume, well-scoped issue resolution where cost per ticket dominates, and a poor candidate, on current evidence, for long-horizon agentic terminal work. A sensible architecture for a team that adopts it is a router that sends scoped patches to the cheap model and escalates long-horizon or repeatedly failing tasks to a stronger one, a pattern covered in our piece on agent frameworks and their cost trade-offs.
Building an In-House Evaluation
If you take one thing from this post, take this: no published number substitutes for running candidate models on your own repositories. A small in-house evaluation costs a few hundred dollars and a week of engineering time, and it answers the only question that matters, which is whether a given model resolves your tasks at an acceptable cost. The structure below follows the same trajectory-based thinking as our production architecture for agent evaluation harnesses.
Step 1: Build a task set from your own history
Mine closed issues and merged pull requests from the last six to twelve months. For each, record the issue text, the base commit, and the human-merged diff as a reference. Select 40 to 100 tasks that span your real distribution: bug fixes, small features, dependency bumps, test additions, and a handful of cross-cutting changes. Keep them private and never paste them into a prompt that a vendor could log for training. Hold back a third as a final validation set you do not tune against.
Step 2: Define graders before you run anything
Use layered grading. Layer one is mechanical: the build passes, the existing test suite passes, and new or updated tests from the reference merge pass. Layer two is a rubric: does the diff touch only relevant files, follow the repository’s conventions, avoid new lint violations, and include tests? Layer three is a blind human review of a sample, scoring would-I-merge on a three-point scale. Where you use a language model as judge for layer two, calibrate it against your human scores first; our post on LLM-as-judge pipelines covers the bias and calibration issues.
Step 3: Run each model in a controlled harness
Fix the harness, tool set, time limit and step cap, and vary only the model and effort level. Where a model is only available inside its vendor’s product, as SWE-2 is, record that as an unavoidable confound and run the others in the same product if possible. Run each task at least three times; agent runs are stochastic, and a single run per task conflates luck with capability.
Step 4: Record cost at the level you will pay
Log prompt and completion tokens per step, wall-clock time, tool calls, retries and the final outcome. Compute cost from your actual billing terms, which may be a subscription seat and not a per-token rate. Then add reviewer time: measure minutes a human spends on each patch, multiply by a loaded hourly rate, and include it in cost per resolved task.
import statistics as st
def summarize(runs, reviewer_rate_per_min=1.2):
"""runs: list of dicts with keys
resolved (bool), mergeable (bool), model_cost (float USD),
review_minutes (float). Rate is an illustrative loaded cost."""
n = len(runs)
resolved = [r for r in runs if r["resolved"]]
mergeable = [r for r in runs if r["mergeable"]]
total_cost = sum(r["model_cost"] + r["review_minutes"] * reviewer_rate_per_min
for r in runs)
return {
"resolve_rate": len(resolved) / n,
"merge_rate": len(mergeable) / n,
"cost_per_run": total_cost / n,
"cost_per_mergeable": total_cost / max(1, len(mergeable)),
"median_model_cost": st.median(r["model_cost"] for r in runs),
}
The key output is cost per mergeable patch, which folds model price, failure rate and review burden into one number you can compare across models and effort levels. Plot it against merge rate to draw your own Pareto frontier, which is the only frontier that applies to your work.
Step 5: Quantify uncertainty
With 60 tasks, a measured 50 percent merge rate has a 95 percent confidence interval of roughly plus or minus 13 points. Two models separated by a point or two are indistinguishable. Use a paired comparison, counting tasks where model A succeeds and B fails and vice versa, and apply a sign test or bootstrap. If the interval for the cost difference excludes zero while the accuracy difference does not, you have a defensible cost-based decision, which is exactly the shape of claim Cognition is making.
Trade-offs, Gotchas, and What Goes Wrong
Vendor lock-in through the harness. Because SWE-2 runs only inside Devin at launch, adopting it means adopting Devin’s context management, tool permissions and billing. If Cognition changes the harness, your results change with it, and there is no way to hold the model constant and swap the product. Open-weight or API-accessible models avoid that coupling.
Narrow optimization. A reward that trades success against cost is tuned to a distribution of training environments. Cognition says it tripled the number of reinforcement-learning environments, but their composition is undisclosed. Behaviour outside that distribution, such as obscure languages, monorepos with custom build tooling or tasks needing long exploration, is exactly where the Terminal-Bench 4 result warns you to expect weakness.
Premature commitment. A model trained to act early, with its median first edit at step 18, can commit to a wrong diagnosis. Cheap, fast agents are most dangerous when they are confidently wrong and a tired reviewer approves the patch. Make tests, not impressions, the gate.
Security surface. Autonomous coding agents with shell access are an attractive target for indirect prompt injection through issue text, dependency READMEs or retrieved files. Nothing in the announcement addresses the safety properties of SWE-2, and a cost-optimised model that takes fewer steps may also take fewer verification steps. Our write-up on prompt injection in agentic systems lists the controls that apply regardless of model: sandboxing, least-privilege tokens, network egress limits and human approval for destructive actions.
Promotion pricing. The free-for-a-month offer ends around mid-October 2026, according to secondary coverage. Any cost model built during the promotion should be recomputed at the post-promotion terms, which Cognition had not clearly published for SWE-2 when this was written.
Benchmark half-life. Even a private benchmark ages as vendors learn its shape. If FrontierCode becomes the standard target, model builders will optimize for it, and the score will drift upward faster than real capability. Rotate your in-house tasks.
Practical Recommendations
Treat SWE-2 as a promising, unverified option, with the cost claim more credible than the capability claim. The step-count and first-edit data describe a real behavioural change against its own predecessor, and a free promotional window makes a trial cheap. What the evidence does not support is a switch of production workflows on the strength of vendor-authored charts.
A pragmatic path is to trial it on a bounded class of work, such as dependency updates, small bug fixes and test generation, while keeping long-horizon tasks on a stronger model. Compare it against your current default on the same private task set and judge on cost per mergeable patch, not on benchmark rank. Keep an eye out for independent evaluations: a replication by a party without a stake in the result would change the confidence level of everything above.
- Quote 64 percent, not 70 percent, and say it is a Cognition-reported figure on one benchmark.
- Note that FrontierCode is authored by Cognition and its tasks are private.
- Do not generalize from FrontierCode to long-horizon work given the Terminal-Bench 4 result.
- Build a 40 to 100 task private set from your own merged pull requests.
- Run at least three attempts per task and report confidence intervals.
- Measure cost per mergeable patch including reviewer time.
- Pin the harness and record effort level for every run.
- Recompute costs when the promotional period ends.
- Sandbox the agent and require approval for destructive commands.
Frequently Asked Questions
What is Cognition SWE-2?
Cognition SWE-2 is a coding-agent model released on 10 September 2026 and offered inside Cognition’s Devin product. Cognition says it was built by applying reinforcement learning to the Kimi K3 base model, a 2.8-trillion-parameter system. It is reported as scoring 50.0 percent on FrontierCode 1.1 Main while costing 64 percent less than Fable 5.1. There are no open weights, no standalone API and no model card at launch.
Is the 70 percent cost saving true?
Not as a general statement. Cognition’s announcement says SWE-2 is within one point of Fable 5.1 while being 64 percent cheaper, based on list pricing including public discounts. A social post reportedly said up to 70 percent, which is an upper bound. A separate 81 percent figure compares SWE-2 medium with SWE-1.7, not with a rival. No per-token or per-task price has been published, and no independent evaluation had verified any of it when this was written.
Is FrontierCode an independent benchmark?
No. FrontierCode was developed by Cognition, reportedly with more than 20 open-source maintainers across 36 repositories, and its tasks are private. That design resists training leakage and grades patches on mergeability rather than test passing, which are real strengths. But the vendor that sells the model also authors the exam, so results should be treated as vendor-reported until a third party replicates them on the same or an independent task set.
Why does SWE-2 score so low on Terminal-Bench 4?
Cognition has not published an explanation. Its reported score is 27.3 percent against 55.8 and 57.9 percent for the two rivals. A plausible but unconfirmed reading is that training for cost-penalized success on repository tasks favours short, decisive runs, which would not help on long-horizon terminal work where the right strategy is to spend more steps. Test that hypothesis on your own tasks before relying on it.
How is SWE-bench Verified different, and should I still trust it?
SWE-bench Verified is a public benchmark of real GitHub issues validated by human annotators. In February 2026 OpenAI reported that 59.4 percent of a 138-problem audited subset had material test-design flaws and that frontier models showed contamination, and it recommended SWE-bench Pro instead. Public scores can still show broad trends, but as a precise ranking tool for new models they are unreliable, and private task sets are better.
How do I decide whether SWE-2 is cheaper for my team?
Run it against your current model on 40 to 100 private tasks drawn from your own merged pull requests, three attempts each, in a fixed harness. Compute cost per mergeable patch, including token or subscription cost, retries and reviewer time, and compare with a paired test. If the accuracy gap is within noise and the cost gap is outside it, the cheaper model wins for that task class. Revisit after the promotional pricing ends.
Further Reading
- AI agent benchmarks: SWE-bench, GAIA and tau-bench for how agent benchmarks are built and where they leak.
- AI agent frameworks benchmark: LangGraph, OpenAI and Google ADK for harness and orchestration trade-offs.
- Small versus large LLMs on agentic tasks for the cost and latency arithmetic of model choice.
- Agentic AI security and prompt injection for hardening autonomous coding agents.
- Cognition: Introducing SWE-2, Pushing the Pareto Frontier, the primary announcement.
- OpenAI: Why we no longer evaluate SWE-bench Verified, on contamination and flawed tests.
By Riju — about
