OCXLY Tech · Explainer

Artificial general intelligence: what we're actually arguing about

Everyone from lab CEOs to newspaper columnists now talks about AGI as though it were a fixed destination. It isn't. The single most useful thing to understand about artificial general intelligence is that nobody agrees on what it means — and almost every argument about timelines, risk, and policy flows from that one unresolved question.

OCXLY Tech Published 10 September 2026 ~13 min read Sources linked throughout
A luminous half-formed sphere of intelligence — part glowing brain, part dissolving into particles — suspended in a dark hall as many hands reach toward it, evoking a concept everyone defines differently.

In early 2025, OpenAI's leadership wrote that the company was "now confident we know how to build AGI." Within days, researchers pointed out that the claim was almost impossible to evaluate, because there is no shared definition of the thing being built.1 That gap — between a term used to raise tens of billions of dollars and a term nobody can pin down — is where this whole debate lives.

This is a plain-English map of the AGI conversation: what the phrase has meant, how the field is trying to measure it, where the technology actually stands, why serious researchers place its arrival anywhere from "a few years" to "not with this approach," and what is genuinely at stake either way. Every load-bearing claim links to a primary or reputable source.

01A term with no owner

The dream is old. In 1950, Alan Turing opened Computing Machinery and Intelligence with "Can machines think?" — then, judging the question too vague, swapped it for a behavioural test: could a machine hold a conversation well enough to be mistaken for a person?2 That move — define intelligence by what a system does, not by what it is — still structures the argument seventy-five years later. The word "artificial intelligence" itself was coined by John McCarthy for the 1956 Dartmouth workshop; the general qualifier came later, to separate a hypothetical all-purpose mind from the "narrow" systems that actually existed.3

An empty museum plinth with a blank brass nameplate under a single spotlight, a crowd of shadowy figures debating around it — a famous term that no one owns or has defined.
AGI is a headline term with no agreed definition — a pedestal everyone argues over but no one can fill.

Here is the problem in one sentence: there is still no agreed definition of AGI, and the ones in circulation measure different things. OpenAI's charter calls it "highly autonomous systems that outperform humans at most economically valuable work" — a labour-market test that says nothing about understanding, consciousness, or reasoning about the physical world.4 The research nonprofit METR has catalogued how these rival definitions pull in different directions — economic value, cognitive breadth, autonomy, human-likeness — and argues that "we've reached AGI" is close to unfalsifiable until you say which one you mean.5 At the far end sits the philosopher Nick Bostrom's superintelligence: "any intellect that greatly exceeds the cognitive performance of humans in virtually all domains of interest."6

02From "can it think" to "how general, how good"

The most useful response to that chaos was to stop treating AGI as a yes/no threshold. In a 2024 paper, a Google DeepMind team led by Meredith Ringel Morris proposed Levels of AGI, grading systems on two axes at once: performance (how well it does a task — Emerging, Competent, Expert, Virtuoso, Superhuman) and generality (how wide a range of tasks it covers).7 On that grid a chess engine is "Superhuman Narrow AI," while the best chatbots of 2023 landed only at "Emerging AGI" — broad in scope, but roughly at the level of an unskilled human across most tasks. The framework adds a third dimension, autonomy, and its real contribution is to reframe the question from "is it AGI?" to "how general, at what level, with how much independence?" — which is at least answerable.

A single glowing lightbulb on the left dissolving rightward into a wide spectrum of many small luminous nodes of varying brightness, a person silhouetted before it — a shift from a yes/no question to a graded scale.
From "can it think?" to "how general, at what level?" — grading intelligence on a spectrum rather than a switch.

03How you measure a moving target

If AGI is graded rather than declared, benchmarks matter — and recent numbers are startling. Stanford's 2025 AI Index found that on three hard tests introduced in 2023 — MMMU, GPQA (graduate-level science), and SWE-bench (real software tasks) — scores jumped 18.8, 48.9, and 67.3 points in a single year. On SWE-bench, systems went from solving 4.4% of problems to 71.7%.8 The catch is that benchmarks keep dying of success: tests meant to last years now saturate in months.8

A translucent measuring caliper trying to gauge a shifting cyan cloud of particles that keeps changing shape and slipping past the marks — a benchmark chasing a moving target.
Benchmarks keep dying of success: tests built to last years now saturate in months.

That saturation is exactly why François Chollet built ARC-AGI — a test of abstract, on-the-fly reasoning where each puzzle is novel, so a model can't win by memorising.10 For years, models scored in the low double digits. Then, in December 2024, OpenAI's o3 reached 87.5% in high-compute mode (and 75.7% at roughly $20 of compute per task), which Chollet called a genuine breakthrough in adapting to novel tasks.9 And yet he added the crucial caveat: "passing ARC-AGI does not equate to achieving AGI… I don't think o3 is AGI yet."9 His team then released ARC-AGI-2, a harder set on which the same reasoning models fall back under 30% while ordinary humans still score above 95%.11 A separate analysis argued the point in its title — OpenAI's o3 Is Not AGI — precisely because a system can ace a benchmark and still lack the robust, sample-efficient generality that benchmark was meant to proxy.12 The pattern repeats: every test that was supposed to be the finish line turns out to be a lap marker.

04The engine, and the argument about the engine

The capability jumps of the last few years rode on one unreasonably effective idea: make the models bigger and feed them more. From 2020, work on scaling laws showed model error falling as a smooth power law in parameters, data, and compute, holding across many orders of magnitude — which is why the industry spent the decade building ever-larger data centres.13 Predictable size gains produced an unpredictable side effect: emergent abilities, skills like multi-step arithmetic that seemed to appear abruptly past a certain scale. That observation fuelled a lot of "AGI is imminent" reasoning — but a 2023 rebuttal, subtitled Are Emergent Abilities a Mirage?, showed many "emergences" are artifacts of harsh, all-or-nothing scoring; measure smoothly and they improve gradually.13 Whether the curve is a cliff or a ramp is the whole ballgame: it's the difference between "keep scaling and generality will emerge" and "scaling buys competence, not a new kind of mind."

A tall stack of glowing translucent layers forming an abstract neural-network engine core, with two silhouetted figures on either side gesturing in disagreement about it.
Scaling is the engine — and whether the curve is a ramp to generality or a plateau is the whole argument.
The honest one-line summary of the AGI timeline debate: capabilities are improving fast on tasks we can measure, and we do not know whether the last, hardest part of generality is a short ramp away or a different problem entirely.

05So when? Ask the builders — carefully

The most-cited data point on timing comes from Katja Grace and colleagues, who since 2016 have surveyed researchers who publish at the top AI conferences. Their 2024 round polled 2,778 experts and found an aggregate forecast of a 10% chance of high-level machine intelligence by 2027 and 50% by 2047, where HLMI means machines that can do every task better and more cheaply than human workers.14 The striking part is the trend: the survey's median jumped forward by roughly 13 years between its 2022 and 2023 editions — a decade of expected timeline erased in a single year of progress.15

A receding row of glowing calendar panels labelled 2024 through 2030-plus dissolving into fog and question marks, a lone builder pointing cautiously toward the haze.
Expert forecasts span from "a few years" to "not this way" — a distribution over dates, not a date.

Two cautions belong beside those numbers. First, aggregating expert guesses is not the same as knowing; the same researchers gave very different answers depending on how the question was phrased, and the field has a poor forecasting record on itself. Second, "50% by 2047" is a distribution, not a prediction — it bundles people who expect AGI this decade with people who expect it never. Lab leaders skew far more aggressive than the survey median, but that is a claim about their own roadmaps, not a neutral estimate.

06The dissent: fluency is not comprehension

For every scaling optimist there is a serious researcher who thinks the current path tops out short of AGI. The most prominent is Yann LeCun, a Turing Award winner and, until late 2025, Meta's chief AI scientist. LeCun has called large language models "an off-ramp, a distraction, a dead end," argued that "just train up the LLMs" as a route to superintelligence "is never going to work," and left Meta to build "world models" that learn how reality behaves rather than predicting the next word.16 His objection is structural: a system trained purely to predict text has no grounded model of the physical world, and no quantity of text supplies one.

A sculpted human head in profile with articulate words streaming from its mouth in ribbons of light, but a hollow empty space where the brain should be — fluent output without comprehension.
The skeptics' core objection: eloquence is not understanding, and no amount of text supplies a grounded model of the world.

The cognitive scientist Gary Marcus has made a version of this case for over a decade — that pattern-matching at scale yields fluent output without reliable reasoning, and the field keeps mistaking the former for the latter.17 The best-known academic statement of the worry predates the boom: in 2021, Emily Bender, Timnit Gebru and colleagues described large language models as "stochastic parrots" — systems that restitch linguistic form they've seen without grounding in meaning.18 Whether today's models have transcended that description or merely become more convincing parrots is genuinely unresolved — and the ARC-AGI-2 result, where reasoning models collapse on novel puzzles a child can solve, is the skeptics' strongest evidence that something real is still missing.11

07Why AGI would be different: the intelligence explosion

Why does any of this rate as more than a technology-adoption story? Because of an argument first made in 1965 by the statistician I. J. Good. An "ultraintelligent machine," he wrote, could design even better machines — "there would then unquestionably be an intelligence explosion, and the intelligence of man would be left far behind."19 The mechanism is recursive self-improvement: each round of improvement makes the next round faster, so capability could climb from human-level to far beyond it over a short stretch. Bostrom later formalised this, mapping the pathways and the difference between a slow, governable takeoff and a fast one that leaves no time to react.6 This is the load-bearing assumption behind treating AGI as a civilisational event rather than a very good product: if generality is a threshold that feeds back on itself, crossing it once could be decisive.

A single spark at ground level igniting an exponential cascade of ever-brighter linked spheres of light racing upward and outward, a tiny human figure below for scale — a runaway chain reaction of self-improving intelligence.
The intelligence-explosion argument: a system that improves itself could climb past human level fast enough to be decisive.

08The stakes, stated plainly

In May 2023, the Center for AI Safety published a one-sentence statement: "Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war." It was signed by hundreds of researchers and executives, including Turing laureates Geoffrey Hinton and Yoshua Bengio and the CEOs of OpenAI, Anthropic, and Google DeepMind.20 Reasonable people read that two ways — a sincere warning, or incumbents talking up their product's power while inviting favourable regulation — and both can be partly true.

A balanced metal scale on a stone plinth above a foggy chasm, one pan holding a glowing seedling of promise, the other a dark heavy weight of risk, held in tense equilibrium.
The upside and the risk are two tails of the same distribution — both grow as systems become more general.

The near-term stakes are economic before they are existential. Goldman Sachs estimated in 2023 that generative AI could expose 300 million full-time jobs worldwide to automation, perform tasks equal to a quarter of current US work hours, and also lift global GDP by around 7% over a decade.22 "Exposed" is not "eliminated," and most jobs are complemented rather than replaced — but even the conservative reading describes a large, uneven reallocation of work.

09Governing something you can't define

Policy has moved unusually fast, which is itself a signal. The 2023 Bletchley Park summit produced the first multinational declaration on frontier-AI risk and led to a network of national AI Safety (now Security) Institutes and to the International AI Safety Report — a January 2025 assessment by 96 experts chaired by Bengio, commissioned by the 30 nations at Bletchley to inform the 2025 Paris summit.21 Its through-line is sober: the risks of general-purpose AI are real and unevenly understood, and the science of measuring and mitigating them lags the pace of capability. The European Union's AI Act, the first comprehensive statute of its kind, adds specific obligations for "general-purpose AI" — a legal category invented precisely because these systems refuse to stay in one application. The pattern everywhere is regulators reaching for handles on capabilities and compute rather than on "AGI" as such.

A neoclassical hall of justice with columns and a scales emblem trying to contain a glowing cyan mist that seeps between the pillars, officials debating on either side amid labels like standards, accountability and enforcement.
Regulators are reaching for handles on capabilities and compute — because "general-purpose AI" refuses to stay in one application.

10The optimists aren't naïve — they're betting on the other tail

It would be a mistake to treat safety as the only serious conversation. Anthropic's CEO Dario Amodei — himself a signatory of the extinction statement — published a long essay, Machines of Loving Grace, arguing that "powerful AI" within a decade could compress a century of progress in biology and medicine, transform mental health and economic development, and strengthen rather than erode democratic institutions, if the transition is managed well.23 His framing is deliberately two-sided: the same capability that makes the risk severe is what makes the upside enormous. The optimist case and the safety case are not opposites; they are two tails of the same distribution, and both grow more consequential the more general the system becomes.

An infographic of a bell curve titled 'The optimists aren't naïve — they're betting on the other tail', with most probability mass in the middle and a bright highlighted far-right tail where a hiker on a cliff reaches toward a small glowing point marking a much brighter outcome.
Calculated optimism, not blind faith: the optimist case bets on the bright far tail of the same uncertainty everyone shares.

11What to actually watch

Strip away the noise and a few honest through-lines remain. AGI has no agreed definition, so treat any confident "we've built it" or "it's impossible" as a claim about a private definition — ask which one.5 Progress on measurable tasks has been real and fast, but the benchmarks keep saturating, so trust the trend of novel-task performance over any single headline score. The central technical bet — that scaling current architectures yields general intelligence — is exactly what the field's most credible skeptics dispute, and the honest answer to "who's right" is that we don't yet know. Timelines have compressed, but forecasting AI is something the field has repeatedly done badly. And the stakes are large enough that the right posture is neither the certainty of the boosters nor the dismissal of the cynics, but sustained, measured attention.

A dark analytical dashboard titled 'What to actually watch', with labelled gauges for capabilities, agentic behaviour, safety and alignment, economic impact, compute, adoption, societal effects and risk indicators, a trends-over-time chart, and a watchful eye motif, viewed by a person at a desk.
Track the signals that matter — novel-task performance, autonomy, safety, and real-world impact — and update your view as they move.

That is an unsatisfying conclusion, and it is the correct one. The most useful thing you can do with the AGI debate is refuse to resolve it prematurely — and keep watching the one number that still separates the machines from us: how well they handle a problem they have never seen before.

About this piece. This is an editorial explainer from OCXLY, written for general readers trying to think clearly about a term that is used loosely and often. Every load-bearing claim links to a primary or reputable source in the references below — definitions to the organisations that wrote them, benchmark figures to Stanford's AI Index and the ARC Prize, forecasts to the researcher surveys, and risk claims to the scientists and institutions making them.

References

  1. ACS Information Age — "OpenAI says 'the AGI era' is here. Experts disagree" (on the definitional gap)
  2. Alan Turing (1950), "Computing Machinery and Intelligence," Mind — the imitation game / Turing Test
  3. Encyclopaedia Britannica — History of artificial intelligence (McCarthy coins "AI"; the 1956 Dartmouth workshop)
  4. OpenAI — Charter (2018): AGI as "highly autonomous systems that outperform humans at most economically valuable work")
  5. METR — "AGI: Definitions and Potential Impacts" (competing definitions and why claims are hard to falsify)
  6. Nick Bostrom, Superintelligence (2014) — definition; recursive self-improvement; slow vs fast takeoff
  7. Morris et al. (Google DeepMind) — "Levels of AGI: Operationalizing Progress on the Path to AGI," ICML 2024
  8. Stanford HAI — 2025 AI Index Report (benchmark gains on MMMU/GPQA/SWE-bench; benchmark saturation)
  9. VentureBeat — OpenAI o3 scores 87.5% on ARC-AGI (Chollet: "I don't think o3 is AGI yet")
  10. ARC Prize / François Chollet — ARC-AGI-1 (a test of novel, on-the-fly reasoning)
  11. Chollet et al. — "ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems" (2025)
  12. Understanding and Benchmarking Artificial Intelligence: "OpenAI's o3 Is Not AGI" (arXiv, 2025)
  13. Overview — Kaplan (2020) & Chinchilla (2022) scaling laws; Wei (2022) vs Schaeffer (2023) on emergent abilities
  14. Grace et al. — "Thousands of AI Authors on the Future of AI" (2024): 10% HLMI by 2027, 50% by 2047
  15. 80,000 Hours — "When do experts expect AGI to arrive?" (shrinking-timeline review)
  16. The Decoder — The case against predicting tokens to build AGI (Yann LeCun on LLMs and world models)
  17. Gary Marcus — on scaling, LLMs, and the limits of pattern-matching
  18. Bender, Gebru, McMillan-Major & Mitchell — "On the Dangers of Stochastic Parrots" (FAccT 2021)
  19. I. J. Good (1965), the "intelligence explosion," and recursive self-improvement
  20. Center for AI Safety — "Statement on AI Risk" (2023; signed by Hinton, Bengio, Altman, Amodei, Hassabis)
  21. International AI Safety Report (2025) — 96 experts, chaired by Yoshua Bengio; commissioned via the Bletchley process
  22. Goldman Sachs (2023) — 300M jobs exposed; ~25% of US work hours automatable; ~7% lift to global GDP
  23. Dario Amodei (Anthropic) — "Machines of Loving Grace" (2024): the optimistic case for powerful AI