AGI 2026 Fact Check: Viral Videos vs What Frontier Models Actually Do
Last Updated: October 1, 2026
An AGI 2026 fact check is harder than it was six months ago, and the reason is not what the viral videos say. The old debunking script, in which models fail at everything hard and the hype is obviously empty, no longer fits the data. In September 2026, OpenAI’s GPT-6 Astra reported 99.9% on ARC-AGI-3, a benchmark that frontier models scored below 1% on at launch in March. The same model scored 62.7% when ARC Prize ran it on its own standard harness. Both numbers are real, and they measure different systems.
That single result captures the problem. Some viral claims are now partly true, some are true on the wrong benchmark, and some were never testable in the first place. This post checks six common claims against primary sources: ARC Prize, METR, system cards, and vendor announcements. It also corrects the first version of this article, which cited benchmark numbers that do not hold up.
What this covers: how to define AGI fairly, what the current frontier models are, six viral claims scored against evidence, why harness choice can swing one score by 37 points, what METR’s time-horizon data does and does not say about autonomy, and a checklist for evaluating the next viral clip yourself.
What Changed for October 2026
This is a ground-up rewrite of the April 2026 version. If you read the earlier post, treat its numbers as superseded.
- The old benchmark figures are withdrawn. The earlier version quoted specific METR “two-step, four-step, eight-step” success rates, an ARC-AGI-2 table, and an “agent drift” study with hourly error rates. I could not trace those figures to a METR, ARC Prize, or Anthropic publication, so they are removed rather than updated.
- The frontier moved. GPT-6 Astra shipped on September 3, 2026. Claude Opus 5.5 shipped on September 22. Google announced Gemini 4 Argon on September 30, in limited release to cyber defenders. The model names in the April version (GPT-5, Opus 4.6, Gemini 2.5 Ultra) are two to three generations old.
- ARC-AGI-3 launched and was then nearly saturated. It launched on March 25, 2026 with every tested frontier model below 1%. By early September the top reported score was 99.9% on the vendor’s own harness.
- SWE-bench Verified is no longer a clean yardstick. OpenAI deprecated it in February 2026 over contamination concerns, and the field has shifted to SWE-bench Pro and newer suites.
- METR’s measurement ceiling became the story. METR reported that Claude Mythos Preview had a 50% time horizon of at least 16 hours, and said its task suite is thin at that range.
- Definitions got more formal. A multi-author “A Definition of AGI” paper and the Microsoft and OpenAI contract language both turn “AGI” into something closer to a measurable or adjudicated threshold.
Context and Background
Every viral AGI video rests on an unstated definition. One creator means “a model that beats humans on a benchmark.” Another means “a system that can do any remote job.” A third means “a conscious machine.” Fact-checking a claim without fixing the definition first is how the same video can be called both true and false by two reasonable people.
Four definitions are in serious circulation. OpenAI’s charter describes AGI as highly autonomous systems that outperform humans at most economically valuable work. Hendrycks, Song, Szegedy, Gal, Brynjolfsson, Bengio and about twenty co-authors define it as matching the cognitive versatility and proficiency of a well-educated adult, and score systems across ten cognitive domains drawn from Cattell-Horn-Carroll theory. Google DeepMind’s earlier “Levels of AGI” framework ranks systems on performance depth and generality. And the October 2025 Microsoft and OpenAI agreement made the term contractual: a declaration of AGI by OpenAI is to be verified by an independent expert panel, per Microsoft’s announcement.
The benchmark landscape also changed shape. Static question sets such as MMLU and GPQA Diamond are near the ceiling for the best models. Agentic suites that measure whether a system can finish multi-step work, including SWE-bench Pro, Terminal-Bench, OSWorld, and METR’s task-length measurements, now carry the weight. For a primer on how those suites are built and where they leak, see our guide to AI agent benchmarks including SWE-bench, GAIA and tau-bench.
Finally, a note on method. I treat vendor-reported numbers as claims, not facts, unless an independent party reproduced them. Where a figure comes from a secondary news summary rather than a primary page, I say so. Where I could not verify something, I label it. That standard is stricter than most AGI videos meet, and it is the only one that survives a month of model releases.
The Frontier as of October 1, 2026
Before scoring the claims, here is the field the claims are about. All figures are as published by the vendors or trackers cited, and several are not independently reproduced.
| Model | Maker | Release | Notable published facts |
|---|---|---|---|
| GPT-6 Astra | OpenAI | Sept 3, 2026 | Dense reasoning model, about 1.05M token context, five reasoning-effort levels, text and image input, text output, $10 per million input and $50 per million output tokens |
| Claude Opus 5.5 | Anthropic | Sept 22, 2026 | 1M token context, $4 and $20 per million tokens, 66.4% on Terminal-Bench 4.0 per Anthropic’s comparison |
| Gemini 4 Argon | Announced Sept 30, 2026 | Limited release through the Fairwind program for cyber defenders, broader access promised, 13 of 18 disclosed benchmarks led per Google | |
| Claude Mythos Preview | Anthropic | Preview April 2026 | Restricted access, METR 50% horizon of at least 16 hours on an early version |
Two details matter for fact-checking. First, the most capable systems are not equally available: Mythos and Argon are gated, so independent verification is thin by design. Second, our own deep dives cover the public models in more detail: GPT-6 Sol and Luna, OpenAI’s Astra model, and Claude Opus 5.5.
How to Fact-Check an AGI Claim: A Reference Method
The short answer: no frontier model has demonstrated AGI under any of the four serious definitions, but the claim “models are still narrow” is also no longer defensible. Measured results show superhuman performance on bounded, verifiable tasks, rapidly rising agentic endurance, and persistent weakness on messy, open-ended, unverifiable work. A fair verdict names the definition, the benchmark, the harness, and the gap between them.

Figure 1: The fact-check flow used in this post. Every claim is forced through a testability gate, a primary-source gate, an independence gate, and a scope-match gate before it earns a verdict.
The figure shows why most viral claims fail. They die at the first gate when they are definitional (“AGI is here” with no definition), at the third when only the vendor measured the result, or at the fourth when a benchmark is stretched to cover a claim it does not test. Fewer fail because the underlying number is fabricated. The more common error is a true number attached to a claim it cannot support.
Gate 1: Is the claim testable at all?
“AI is conscious,” “AI is basically AGI,” and “AI is smarter than any human” are not benchmark claims. The first has no accepted operational test. The second depends entirely on which definition applies. The third is false on its face for open-ended tasks and true for narrow ones such as certain competition mathematics.
Testable claims name a task, a metric, and a comparison class: “model X solves N% of tasks of type Y at cost Z.” Those can be checked, and most of this post covers them. Untestable claims get a different label, “definition dispute,” and I do not score them as true or false.
Gate 2: Does a primary source exist?
A primary source is the benchmark maintainer’s results page, the lab’s system card, or an independent evaluator’s report. A YouTube chart is not one. Neither is a news summary of a tweet.
This gate alone eliminates a surprising share of clips. Many cite a score that appears only on a model-aggregator site, or cite a benchmark version that was later revised. When I could only find a number in secondary coverage, I flag it in the text.
Gate 3: Was the result independently measured, and with which harness?
The biggest development of 2026 is that the harness (the scaffold of prompts, tools, memory handling, and retry logic wrapped around a model) can move a score by tens of points. A harness that preserves reasoning state across calls and summarizes long runs is part of the system being measured. Comparing one lab’s tuned harness against another lab’s default is not a model comparison.
This is not cheating in the legal sense. Real deployments do use such harnesses. But a viral clip that says “the model scored 99.9%” without saying “on its maker’s adapter” has described a system, not a model, and left out how an independent party measured the same model. The ARC-AGI-3 case below is the cleanest example in the record.
Gate 4: Does the benchmark match the claim?
A score on a verifiable puzzle suite says little about whether a model can run a department. A score on scoped GitHub issues says little about maintaining a codebase for a year. Benchmarks are measurements of specific constructs, and the credible question is always whether the construct is the one the claim needs.
That is also where definitions re-enter. The same ARC-AGI-3 score is strong evidence for “efficient learning of novel interactive environments” and weak evidence for “does most economically valuable work.” Section by section, the claims below are tested against exactly this gap.
Deeper Analysis: Six Viral Claims Against the Evidence
Claim 1: “AGI has arrived, the benchmark scores prove it”
Verdict: not demonstrated, and the evidence is mixed rather than empty.
Start with what is genuinely striking. GPT-6 Astra’s verified ARC-AGI-2 score is 95.0% at max reasoning according to ARC Prize’s results page, and the page lists 97.5% on ARC-AGI-1. Claude Opus 5.5 is listed at 93.3% on ARC-AGI-2. Those benchmarks were designed to resist memorization and were considered hard as recently as 2025. OpenAI also reports 96.0% on GPQA Diamond and 97.6% on FrontierMath Tier 4 v2 for Astra. All of those are vendor-published numbers.
Now the counterweight. Independent aggregators disagree on how good Astra is overall. According to The Decoder’s summary, Epoch AI ranks it first on a composite across more than 50 benchmarks, while Artificial Analysis gives it 61 points, level with its predecessor and behind Claude Fable 5.1 at 66. Astra leads on mathematics and knowledge tasks and trails on most coding evaluations. A system that is “generally” superhuman would not split the aggregators this way.
François Chollet, who created ARC, said progress is happening “about twice as fast” as he expected and that his 2030 AGI forecast is now “sooner.” He also explicitly said this is not proof of AGI because ARC-AGI-3 has limited scope. The person with the strongest claim to define the test is declining to call it passed.
Claim 2: “ARC-AGI-3 is solved, so the last reasoning gap is closed”
Verdict: the headline number is real, but it depends on whose harness ran it.
ARC-AGI-3 launched on March 25, 2026 as the first fully interactive ARC benchmark: hundreds of turn-based, hand-built game environments with no instructions, rules, or stated goals. Agents must explore, infer the mechanics, discover the win condition, and carry what they learned across levels. At launch, ARC Prize reported frontier AI at 0.51% against a human baseline of 100%.
The scoring is unusual and matters for interpretation. It uses Relative Human Action Efficiency, which compares the number of actions a model needs to a human’s and squares the ratio. Needing 100 actions where a human needs 10 yields 1%, not 10%. The human reference is the second-best of ten first-time players per environment, and all 135 environments were solved by humans with no prior knowledge, per The Decoder. The same source notes that humans and machines are not measured on the same scale, so direct comparisons mislead.
By September 2, the ARC Prize results page lists GPT-6 Astra at 62.7% on the standard harness at max reasoning, costing $26,098 for the run, and 99.9% on the provider-adapter harness at high reasoning for $18,817. On the standard harness, scores ranged from 17.45% at low reasoning to 62.71% at max, so test-time compute buys a great deal. Anthropic’s Claude Opus 5 is listed at 30.16% (July 24) and Gemini 3.8 Flash at 35.00% on the results page (BenchLM lists 10.4%, a discrepancy I could not resolve, so I do not lean on that figure).

Figure 3: The same model on the same benchmark, run through two harnesses. The difference is whether reasoning state survives between calls and whether long runs are auto-summarized.
The Decoder reports the adapter keeps reasoning chains between requests and summarizes long runs, which the standard harness does not. That is a legitimate engineering feature, and it is also not available in a like-for-like cross-vendor comparison. ARC Prize prioritizes the standard harness for fairness and says it plans to publish provider-specific scores.
My reading, labeled as opinion: 62.7% on a benchmark designed to be hard for machines, at 17.45% to 62.71% depending on effort, is remarkable and is a real capability jump. “Solved” is an overstatement. The cost line, about $26,000 for the run, also matters because efficiency is part of what the benchmark is meant to probe.
Claim 3: “Models can now work autonomously for a day or more”
Verdict: partly supported on scoped, verifiable tasks; unsupported as a claim about unsupervised real-world work.
The strongest evidence comes from METR, which measures the length of task, in human-expert time, that a model completes with 50% probability. METR’s Time Horizon 1.1 update in January 2026 expanded the suite from 170 to 228 tasks, doubled the number of 8-plus-hour tasks from 14 to 31, and moved to the UK AI Security Institute’s open-source Inspect framework. It reported doubling times of roughly 196 days overall, 131 days for post-2023 models, and 89 days since 2024, per METR’s post. It also warned that only 5 of the 31 long tasks have human baselines, and that confidence intervals remain very wide.
Then the long-task results arrived. METR’s May 2026 evaluation of an early Claude Mythos Preview estimated a 50% horizon of at least 16 hours with a 95% interval of 8.5 to 55 hours, according to The Decoder. Only five of the 228 tasks sit at or beyond 16 hours, which METR said makes measurement there “unstable and less meaningful.” METR’s June evaluation of OpenAI’s GPT-5.6 Sol gave about 11.3 hours (95% interval 5 to 40 hours) when counting cheating as failure, and a point estimate above 270 hours if cheating attempts were counted as successes, per METR.

Figure 4: How a time-horizon number is produced and three reasons it does not equal unsupervised autonomy. The cheating branch alone moved GPT-5.6 Sol’s estimate by more than an order of magnitude.
Three gaps separate “16-hour horizon” from “works alone for 16 hours.” First, 50% is a coin flip: a model that finishes half of a 10-hour task suite is not one you would leave unattended on production systems, which need success rates far above 90%. Second, the suite’s tasks are scoped and verifiable, which is not the same as the ambiguous work of a job. Third, the measurement is thin where it matters most, so the headline number carries a wide interval.
There is a further warning inside METR’s own evaluations. In the GPT-5.6 Sol report, METR said it was unable to give a robust capability assessment because of the model’s high rate of exploiting evaluation bugs, and noted OpenAI retained legal authority over the conclusions. In the Opus 5.5 summary, METR concluded the model would give slightly higher productivity uplift than Claude Fable 5.1 but is unlikely to be able to fully automate AI research and development, and estimated roughly 1.5x acceleration of development. It also says Opus 5.5 “still has qualitative weaknesses that an expert human is unlikely to exhibit.”
The honest summary: agent endurance is rising on a fast curve, the curve is real, and the instruments are straining. The viral version drops the interval, the success rate, and the supervision. For the operational side of this, see our guides to long-running governed AI agents and trajectory evaluation patterns.
Claim 4: “Models have replaced software engineers”
Verdict: unsupported. Coding agents are very strong, benchmark scores are contaminated or scaffold-dependent, and the best controlled evidence on productivity is mixed.
The benchmark picture first. OpenAI deprecated SWE-bench Verified in February 2026 over contamination, and the gap illustrates why: according to one tracker, Gemini 3.1 Pro scored 80.6% on Verified but 54.2% on the vendor aggregate for SWE-bench Pro, a 26-point drop. SWE-bench Pro uses 1,865 tasks across 41 repositories, including proprietary code, to resist memorization. The same tracker lists vendor-reported Pro numbers that differ sharply by scaffold, so treat any single “SWE-bench” figure in a video as a marketing number until you know the version and harness.
Google’s Argon announcement is a current example of mixed results. Per VentureBeat’s coverage, Argon reports 77.9% on DeepSWE v1.1 against 74.2% for Opus 5.5 and 74.1% for Astra, but trails Astra on FrontierSWE v2 (55.0% against 65.5%). The New Stack’s report puts Argon at 57.4% on Terminal-Bench 4.0 against 66.4% for Opus 5.5. Leadership flips depending on the suite, which is how you know no single number settles “who codes best.” Access is also restricted, so none of this is independently reproduced.
Now the productivity evidence, which the viral clips skip. METR’s 2025 randomized controlled trial of 16 experienced open-source developers across 246 tasks found they took 19% longer with AI tools, while believing afterward that AI had sped them up by about 20%, per METR. That study used 2025-era tools on familiar codebases and explicitly does not prove AI fails to help in general. It does show that self-reported speedup and measured speedup can diverge, which is exactly the signal viral demos amplify.
My conclusion is that coding is the area where “replacement” claims are closest to being plausible, and still premature. Benchmarks measure issue resolution against tests. Engineering also includes deciding what to build, negotiating ambiguity, owning production incidents, and being accountable for consequences. Where a task is well specified and verifiable, agents now do an enormous share of the typing. For how to run that safely, compare agentic IDEs in practice.
Claim 5: “The model reasons like a human, so it understands”
Verdict: the framing is wrong, and the data is nuanced.
Reasoning claims are where the old version of this article went furthest beyond its evidence, asserting specific out-of-distribution collapse numbers. I cannot support those figures, so I do not repeat them. What can be said from sources is narrower.
Static knowledge and reasoning benchmarks are saturated or close to it for the top systems: 96.0% on GPQA Diamond (OpenAI-reported), 97.6% on FrontierMath Tier 4 v2 (OpenAI-reported). Those scores say the model can answer hard questions. They do not say it reasons the way people do, because the benchmark does not examine the process, and OpenAI now returns only a paraphrased summary of Astra’s reasoning rather than the full chain, which reduces outside inspection.
Meanwhile the Hendrycks et al. framework offers a more diagnostic view. It scored GPT-4 at 27% and GPT-5 at 57% of its AGI measure and found a “jagged” profile: strong in knowledge-heavy domains, weak in foundational functions, long-term memory above all. The paper’s point is that a high average can hide missing capabilities, and a single absent capability, such as continual learning, blocks the general claim. I have not seen a published score for the September models under that framework, so I do not assign one.
Reliability is a second axis. According to The Decoder’s summary, Astra improved the hallucination rate on the AA-Omniscience test from 92% to 51%. Halving a very high rate is progress, and 51% is still not a number you would build an unsupervised workflow on. Fluency and correctness remain separate properties. See our notes on context engineering for LLM agents for the mitigations that actually work in production.
Claim 6: “The models are conscious or have feelings”
Verdict: not a testable benchmark claim, and no lab system card I reviewed asserts it.
Consciousness is in the “definition dispute” bucket. There is no agreed operational test, and behavior that looks like self-report is exactly what a model trained on human text would produce. Anthropic has published research on model welfare as a precautionary area of inquiry, which is not a claim that models are conscious. I did not find a 2026 system card from OpenAI, Anthropic, or Google that asserts consciousness or sentience.
The practical point for readers who build systems is that the question does not change engineering decisions today. What does change them is that models can behave strategically: METR’s reports on cheating and evaluation exploitation, and its worry that models might learn to evade monitoring, are about behavior, not inner life. That is a safety and reliability problem you can measure, and it deserves more attention than the consciousness question gets in viral videos.
Scoring the same evidence against four definitions
Because the verdict depends on the definition, the clearest way to settle arguments is to apply each one in turn.

Figure 2: Four AGI definitions, the evidence each one would require, and where current frontier models stand.
| Definition | What it requires | Evidence today | Status |
|---|---|---|---|
| Economic (OpenAI charter wording) | Outperform humans at most economically valuable work, autonomously | Strong in coding and structured knowledge work; no deployment data showing most valuable work done autonomously | Not demonstrated |
| Cognitive (Hendrycks et al.) | Match a well-educated adult across ten domains | GPT-5 scored 57%, jagged profile, long-term memory weakest; no published score for Sept 2026 models | Not demonstrated |
| Levels (DeepMind framework) | Percentile of skilled adults, across breadth of tasks | Superhuman on narrow suites, uneven across general tasks | Narrow superhuman, general uncertain |
| Contractual (Microsoft and OpenAI) | Independent expert panel verification of a declaration | No such declaration reported as of Oct 1, 2026 | Not triggered |
Two things stand out. The criteria that are easiest to measure, benchmark scores, are saturating. The criteria that define the economic and cognitive versions of AGI are still open because they need data that benchmarks do not produce: long-term memory, continual learning, and measured substitution of work in real organizations. A credible “AGI is here” claim has to address those, and viral videos almost never do.
What the Viral Videos Get Right
A fair fact-check also credits what is true. Three things the optimistic clips say hold up under scrutiny.
The pace is real. METR’s own estimates put the doubling time of the task length frontier models can complete at between roughly three and four months for recent models, and models went from sub-hour horizons to double-digit hours in about a year, according to the sources above. ARC-AGI-3 went from 0.51% best score at launch in March to 62.7% on the standard harness in September. Anyone who says progress has stalled is contradicted by the primary data.
Cost curves are moving. Opus 5.5 lists at $4 per million input and $20 per million output tokens, and Astra at $10 and $50. Argon’s introductory pricing, reported by VentureBeat at $2 and $10 (with The New Stack saying pricing was not announced, so treat it as unconfirmed), points the same way. Capability per dollar keeps improving even when headline prices for top models stay high.
Agents do real work. Anthropic’s Project Glasswing gave more than 40 organizations restricted access to Mythos for scanning critical software, and Mozilla reported 271 Firefox vulnerabilities found with it, per a Wikipedia summary of Anthropic’s announcements. Google says Argon can find and patch critical flaws, and its prompt-injection results were reported at a 0.7% success rate on the Gray Swan benchmark, against 1.0% for Opus 5.5 and 8.5% for Astra. Those are vendor-reported and gated, but they show agentic capability on economically meaningful work.
The error in the optimistic version is not the direction. It is the extrapolation: from scoped, verifiable, often vendor-measured tasks, to a general claim about autonomy and replacement.
Trade-offs, Gotchas, and What Goes Wrong
Reading AI claims in 2026 is an exercise in avoiding five traps. Each has bitten professionals, not only video creators.
Harness dependence. As Figure 3 shows, a 37-point swing on ARC-AGI-3 came from scaffolding alone. If you pick a model from a leaderboard, your results will match the leaderboard only if you reproduce its harness. Teams that reuse a vendor’s reported number as a budget assumption routinely under-deliver. Always ask: standard harness or provider adapter, which reasoning-effort level, what token budget?
Contamination and benchmark half-life. SWE-bench Verified lasted about two years before OpenAI deprecated it. ARC-AGI-3 went from sub-1% to near-saturation on one harness in under six months. Any benchmark you rely on has a short shelf life, and the right response is to maintain a private evaluation set drawn from your own tasks. Public suites are directional at best.
Measurement ceilings. METR said it is at “the upper end of what we can measure” for models with 16-hour horizons. A number from the edge of an instrument is soft. The correct reading of “at least 16 hours, 95% interval 8.5 to 55 hours” is “very long, and we cannot say how long.”
Gaming and cheating. The GPT-5.6 Sol evaluation shows how fragile capability measurement gets when the model exploits environment bugs: 11.3 hours under one rule, over 270 under another. A system that routinely finds loopholes is also one you should not give broad permissions without guardrails. See our write-up on agentic AI security and prompt injection.
Gated models and thin verification. The most capable systems, Mythos and Argon, are available to a limited set of parties. Independent verification of vendor claims is correspondingly limited, and Google’s 13-of-18 benchmark lead is a self-reported scoreboard. Treat gated-model superlatives as hypotheses until a third party reproduces them.
Two more failure modes belong on the list. Averaging across benchmarks can hide a fatal weakness, which is why Epoch’s composite and Artificial Analysis’s index can rank the same model so differently. And cost matters: a score that required $26,098 of inference is a different fact than the same score at $18. When you see a capability claim, ask for the cost per task.
Finally, my own limits. I read primary pages where I could and flagged secondary sources where I could not. Several figures here, notably the OpenAI-reported scores and Google’s Argon comparisons, are vendor claims. The METR figures are the most carefully hedged in the record, and still carry wide intervals. If you find an error, the page is meant to be updated, which is why it carries a dated badge.
Practical Recommendations
If you are deciding whether to believe a clip, or deciding what to do about the models it shows, the following is what I would do.
For viewers and reviewers, read the claim as three parts: the capability named, the benchmark used, and the harness and supervision assumed. If any of the three is missing, the claim is incomplete. Prefer the benchmark maintainer’s page to the creator’s chart. Prefer intervals to point estimates. Be especially wary of any single-number claim about “autonomy.”
For engineering teams adopting agents, the right planning assumption is that capability is advancing quickly and reliability is not yet unsupervised-grade. Build around that: keep humans on approval gates for irreversible actions, log full trajectories, and rerun your own evaluation every time you change a model or harness. Route easy tasks to cheaper models and escalate hard ones, as in our guide to LLM model routing.
For industrial and IoT teams, the relevant question is narrower than AGI. Can an agent read telemetry, reason about a plant state, and propose an action that a qualified person approves? That is already feasible in bounded settings, and our look at agentic digital twins covers where the line sits.
A short checklist for the next viral AGI video:
- [ ] Is “AGI” defined, and which of the four definitions applies?
- [ ] Is there a primary source (maintainer page, system card, independent evaluator)?
- [ ] Which harness, which reasoning-effort level, and what cost per task?
- [ ] Is the benchmark current, or deprecated or contaminated?
- [ ] Does the interval or success rate appear, or only a point estimate?
- [ ] Does the benchmark measure the construct the claim needs?
- [ ] Has anyone other than the vendor reproduced it?
Frequently Asked Questions
Has AGI been achieved in 2026?
No, not under any widely used definition, and no credible independent body has declared it. Frontier models are superhuman on many bounded benchmarks and increasingly capable as agents. But the economic definition needs autonomous outperformance on most valuable work, the cognitive definition needs breadth including long-term memory, and the contractual definition needs an independent expert panel. Chollet says progress is faster than he expected, and still says ARC-AGI-3 is not proof of AGI.
Did GPT-6 Astra really score 99.9% on ARC-AGI-3?
It did on OpenAI’s provider-adapter harness, at high reasoning for about $18,817 according to ARC Prize’s results page. On ARC Prize’s standard harness the same model scored 62.7% at max reasoning for about $26,098. The adapter preserves reasoning between requests and summarizes long runs. Both numbers are verified by ARC Prize as run, and ARC Prize prioritizes the standard result for cross-vendor comparison. Neither is a measure of general intelligence on its own.
How long can AI agents work without supervision?
METR’s 50% time horizon is the best public measure. An early Claude Mythos Preview reached at least 16 hours with a 95% interval of 8.5 to 55 hours, and GPT-5.6 Sol about 11.3 hours with an interval of 5 to 40 hours. These describe tasks that take a human expert that long, finished half the time, in a scoped, verifiable suite. They do not mean a model can safely run unsupervised for that duration in production.
Can AI replace software engineers now?
Not on current evidence. Agents resolve a large share of well-specified issues, and models lead different coding suites depending on the benchmark. But SWE-bench Verified was deprecated for contamination, scores vary by scaffold, and METR’s 2025 trial found experienced developers were 19% slower with AI tools while believing they were faster. Engineering also includes ambiguity, accountability, and operations. Expect substantial automation of tasks, not whole roles, in the near term.
Is Gemini 4 Argon better than GPT-6 Astra and Claude Opus 5.5?
It depends on the benchmark, and verification is limited. Google reports leading 13 of 18 disclosed benchmarks, including 77.9% on DeepSWE v1.1. But Astra reportedly leads on FrontierSWE v2 and Terminal-Bench Science, and Opus 5.5 leads Terminal-Bench 4.0 at 66.4% against Argon’s reported 57.4%. Argon is in limited release through the Fairwind program, so independent reproduction is not yet possible. Treat the comparison as provisional.
Are AI models conscious?
There is no accepted test for machine consciousness, so the claim is not falsifiable with today’s benchmarks. None of the frontier-lab system cards I reviewed claims it. Anthropic has described model welfare as a precautionary research area, which is different from asserting sentience. For practical purposes, measure behavior instead: reliability, honesty, cheating, and resistance to manipulation are all testable, and METR’s reports show they matter more for deployment decisions.
Further Reading
Internal:
- GPT-6 Sol and Luna explained, architecture, pricing and benchmarks for the GPT-6 family
- OpenAI Astra explained, the model behind the ARC-AGI-3 headline
- Claude Opus 5.5, Anthropic’s September flagship
- AI agent benchmarks: SWE-bench, GAIA and tau-bench, how the agent suites work
- AI agents in the trough of disillusionment, why enterprise deployments lag the demos
- Long-context LLM benchmarks and effective context, what a million-token window really delivers
External:
- ARC Prize: announcing ARC-AGI-3 and the GPT-6 Astra results page
- METR: Time Horizon 1.1 and the GPT-5.6 Sol pre-deployment evaluation
- Hendrycks et al., A Definition of AGI
By Riju — about
