AI did not simply move up a ladder. It spread across forms—and began to act through tools, software, and delegated workflows.
4 MINUTES 45 SECONDS
From producing outputs to pursuing outcomes.
The video never autoplays. Captions and the complete transcript are available on this page.
THE UPDATED MAP
Three layers, not one intelligence ladder.
The layers overlap. A single product may combine a generative model, specialized classifiers, and an agent harness.
Recognize · rank · forecast
Specialized + predictive
Task-specific systems classify, recommend, detect patterns, and predict. Limited memory may carry recent data or state, but competence stays bounded.
recommendations
fraud flags
forecasting
game play
WHAT CHANGED SINCE LAST YEAR?
The future category became a visible capability profile.
2024 MAP
A ladder toward “capable AI”
Tool use, computer control, long-running work, and delegation sat mostly in a future-facing category above generative AI.
→
2026 MAP
Capabilities distributed through products
Those behaviors are observable in current systems—but unevenly, through permissions and harnesses, and without proving AGI. The field guide below tracks every box.
THE FIELD GUIDE · REALIZED VS. THEORETICAL
Two tiers still hold: what exists, and what is still a forecast.
Last year’s version of this map drew the same two tiers. The tiers survived the year; almost every box inside them changed. Every claim below carries its source inline.
Tier 1 · Realized AI
Systems that exist and ship.
Realized AI is anything you can actually use: the narrow, task-specific systems that quietly run daily life, and the generative and agentic systems that made the last three years feel so fast. “Realized” does not mean flawless. It means observable.
1.1Reactive machinesNarrow AI
The oldest pattern in the tier. A reactive machine responds to the situation in front of it using fixed rules or trained patterns—no memory of past encounters, no model of the future. IBM’s Deep Blue, which defeated world chess champion Garry Kasparov in 1997, is the classic example; simple spam filters and static recommendation rules still work this way.
Classroom echo: basic drill-and-practice software has behaved like this for decades—respond to the answer in front of it, remember nothing.
1.2Limited memoryNarrow AI
Nearly every deployed AI system today adds short-term memory: recent data, stored patterns, a running state. Self-driving systems such as Waymo’s track surrounding vehicles and predict behavior seconds ahead; fraud detection weighs your recent transactions; forecasting models digest last season. This is the quiet infrastructure under search, logistics, medicine, and finance—astonishing inside its task, helpless one step outside it.
1.3Generative + multimodalThe visible frontier
The category that redrew the public map. In mid-2024 this box was mostly chatbots and image tools. As of July 2026 it spans text, image, video, music, code, and science—and the “reasoning model” idea that arrived with OpenAI’s o1 in September 2024 and went open-weight with DeepSeek-R1 has been folded into nearly every flagship: models now decide when to answer fast and when to think longer. The table below is the July 2026 frontier. It will age; the links will tell you how.
Frontier generative models as of July 2026: model, category, developer, primary use, and source links
Model
Category
Developer
Primary use
Source
GPT‑5.6 (Sol · Terra · Luna)
Language + reasoning
OpenAI
Flagship family released July 9, 2026. One router switches between fast answers and deeper reasoning; Sol adds a max reasoning effort and an ultra mode that delegates to subagents.
Frontier model for long-horizon reasoning and multi-day agentic projects, with always-on adaptive thinking, a 1M-token context window, and subagent delegation. Announced June 9, 2026.
The everyday tiers: Sonnet 5 (June 30, 2026) is the default assistant model; Haiku 4.5 is the fast, low-cost option. Opus 4.8 sits between Sonnet and Fable.
Gemini 3.5 Flash (GA May 2026) targets sustained agentic and coding work at Flash cost; 3.1 Pro is the reasoning-first tier with a 1M-token context window.
Natively multimodal mixture-of-experts family with open weights; Scout offers a 10M-token context window. The largest variant, Behemoth, remains unreleased.
Open-weight MoE family with a unified thinking toggle (V4) and a specialized reasoning line (R2); successors to the R1 release that made reasoning models open in January 2025.
European lab shipping the largest open-weight MoE flagship (Large 3) plus Magistral, a reasoning line with transparent, multilingual chains of thought.
Released April 21, 2026; plans, optionally searches the web, and self-checks before drawing. Replaced DALL·E as OpenAI's default image model in May 2026.
Physically accurate video with synchronized dialogue and sound. Included deliberately as a caution: the consumer app was retired April 26, 2026, and the API ends September 24, 2026. Frontier products can also disappear.
Full songs from text prompts. A November 2025 settlement with Warner Music commits Suno to licensed models and retires those trained on unlicensed catalogs; other suits continue.
Predicts the structure and interactions of proteins, DNA, RNA, and small molecules. Demis Hassabis and John Jumper shared the 2024 Nobel Prize in Chemistry for this line of work.
Capabilities and dates are the makers’ claims unless an independent source is linked. Model names churn quickly; treat the categories as the durable part of the table.
For the classroom: any artifact a student can hand in—essay, image, video, song, working code—can now be generated to a competent standard. That fact sits underneath every other page on this site.
1.4The agentic waveNew since 2024
The biggest shift since mid-2024 is not a new kind of output; it is action. Given permission, current systems plan multi-step work, use tools, browse, operate a computer, check their own results, and keep going. OpenAI describes GPT‑5.6 Sol coordinating subagents; Anthropic says Claude Fable 5 can run projects for days and delegate to subagents; Google’s Gemini computer use spans browser, mobile, and desktop with confirmations for sensitive actions.
For the classroom: “show your work” changes meaning when software can do the work, step by step, on a student’s behalf. Supervision, permissions, and honest disclosure become part of the assignment design.
1.5Emotion AIApplication field
Systems that infer emotional signals from faces, voices, text, or physiology. Inference is the honest word: the system observes proxies, not feelings. Validity varies by context and population, and bias, privacy, and consent remain unresolved—reasons for schools to treat emotion-sensing products with particular care rather than as settled science.
1.6Neural decoding + brain imagingApplication field
Research systems can reconstruct approximations of language a person is hearing or imagining from brain activity—the University of Texas semantic decoder is the landmark public example. It required hours of consenting cooperation and fails without it. This is not mind-reading; it is a field where consent and neural privacy deserve rules before products arrive.
Tier 2 · Theoretical AI
Categories that remain forecasts—one of them now blurry.
Theoretical AI is the tier of proposed capabilities: useful for orientation, dangerous when treated as achieved. Mustafa Suleyman’s The Coming Wave gave the 2024 version of this map its ladder—ACI, then AGI, then ASI. The ladder is still the clearest way to talk about what has not happened. One rung, though, no longer sits entirely in this tier.
2.1ACI · Artificial Capable IntelligenceThe card that moved
Suleyman—DeepMind co-founder, and since March 2024 the CEO of Microsoft AI—proposed ACI as the near-term category worth watching: AI that can accomplish complex, multi-step goals in the real world. His “Modern Turing Test” made it concrete: give an AI $100,000 and ask it to turn it into $1 million, legally and mostly on its own.
As of July 2026, no lab has published a verified, autonomous run of that test. But the capability profile behind it—planning, tool use, delegation, transactions under permission—is now observable in shipping products. That is the argument of this keynote’s opening film, The Card That Moved: the ACI card did not get discarded; it slid across the board, from theoretical toward realized. The honest phrasing is “begins to.” Capability is jagged, supervision is required, and a moved card is not a certified milestone.
2.2AGI · Artificial General IntelligenceDisputed
OpenAI’s charter defines AGI as highly autonomous systems that outperform humans at most economically valuable work. There is no settled test, and no verified public milestone lets this page mark the box achieved. The evidence stays jagged: near-human scores on some computer-use benchmarks, one-in-three failure on the same benchmarks, and expert disagreement about what would count. Confident “AGI is here” claims are marketing until they arrive with a test the claimant did not design.
2.3ASI · Artificial SuperintelligenceTheoretical
Systems beyond human capability across nearly every domain remain a proposed future category. What changed is the language of the labs: Microsoft AI now runs an in-house team pursuing what Suleyman calls “humanist superintelligence”. A team name is an ambition, not evidence. This page tracks capabilities, and there are none to report in this box.
2.4Who is checking the work?Safety, as of 2026
The 2024 version of this map could point to OpenAI’s Superalignment team, which pledged a fifth of the company’s compute to aligning superhuman AI. It cannot anymore: that team dissolved in May 2024, and the successor Mission Alignment team was disbanded in February 2026, with safety work distributed into product and research teams. OpenAI’s public commitments now run through its Preparedness Framework and safety pages.
Elsewhere: Anthropic publishes a Responsible Scaling Policy (version 3.0, February 2026, since updated to 3.1) with public Frontier Safety Roadmaps and periodic Risk Reports; Google DeepMind maintains a Frontier Safety Framework; and the EU AI Act’s obligations for general-purpose models phase in through 2026 and 2027. The honest summary for educators: real safety work exists, the structures around it keep churning, and a lab’s safety claim deserves the same sourcing scrutiny as its capability claim.
PRODUCT-CAPABILITY TIMELINES
Logos stayed. Capability stacks changed.
The timeline tracks products and model releases; it does not imply that a brand mark itself “evolved.”
ChatGPT + Codex
Conversation becomes the public doorway
Multimodal work and connected tools expand
Codex operates in repositories and handles longer tasks
GPT-5.6 Sol adds stronger computer use and multi-agent workflows
Do not confuse an application field with a higher mind.
Each label states what can honestly be claimed now—and what remains unresolved.
Application field
Emotion AI
Infers signals from faces, voices, or bodies. It does not directly observe emotion. Context, validity, bias, privacy, and consent remain central.
Application field
Neural decoding
Infers task-specific information from brain activity. It is not literal mind-reading; privacy, consent, validity, and misuse require unusually strong safeguards.
Disputed milestone
AGI
There is no settled test or verified public milestone that lets this chapter mark artificial general intelligence as achieved.
Disputed / theoretical
Theory of Mind AI
Models can generate plausible accounts of beliefs and intentions. That is not proof that a system possesses a human-like theory of mind.
Theoretical
ASI
Artificial superintelligence remains a proposed future category, not a verified present capability.
SOURCE TRAIL
Who is making each claim?
Provider claims describe shipped products. Independent evidence tests reliability. Speaker framing makes the map usable without pretending it is settled science.
Provider claim
OpenAI: GPT-5.6
Computer use, agent coordination, and tool support as described by the maker.
The transcript matches the approved narration script. Captions are a timed presentation of the same words.
The map redraws itself
Last year, I showed artificial intelligence as a ladder: narrow systems at the bottom, generative AI above them, and a future category called Artificial Capable Intelligence somewhere ahead. Twelve months later, that map no longer works. The categories did not simply climb. They spread, overlapped, and—most importantly—began to act. So let’s redraw the landscape in three layers.
Layer one: specialized and predictive AI
Layer one is specialized and predictive AI. These systems recognize a face, recommend a song, flag a suspicious transaction, forecast demand, or choose the next move in a game. Some react only to what is in front of them. Others use limited memory: recent data, stored patterns, or a running state. They can be astonishing inside a defined task and helpless one step outside it. This layer is the quiet infrastructure beneath search, logistics, medicine, finance, and daily life.
Applications, not higher minds
Emotion AI and neural decoding belong beside this layer as application fields—not as higher forms of intelligence. Systems can infer signals from faces, voices, bodies, or brain activity, but an inference is not mind-reading. Validity varies. Context matters. Privacy, consent, bias, and misuse are central design questions.
Layer two: generative and multimodal AI
Layer two is generative and multimodal AI. A year ago, many people still pictured a chatbot producing paragraphs. Now models work across text, images, audio, video, code, diagrams, and documents—and increasingly reason across several of those forms at once. A prompt can become a lesson, podcast, software prototype, translated video, or visual explanation.
But keep four things separate. A model is the learned engine. A product—such as ChatGPT, Claude, or Gemini—is the experience wrapped around it. An agent harness gives a model tools, memory, permissions, and a loop for taking steps. An application puts all of that into a particular workflow. Confusing those layers makes every new demo sound like a new kind of mind.
Layer three: agentic or capable AI
Layer three is agentic—or capable—AI. The shift is from producing an artifact to pursuing an outcome. With permission, current systems can plan, browse, use a terminal, edit files, operate software, check their work, and continue through a long task.
Watch the product timelines. ChatGPT moved from conversation to multimodal work and tool use. Codex moved from suggesting code to operating in repositories and coordinating parallel agents. In July 2026, OpenAI described GPT-5.6 Sol as supporting stronger computer use and multi-agent workflows. Anthropic says Claude Fable 5 can run projects for days and delegate to subagents. Google’s Gemini computer-use system spans browser, mobile, and desktop environments, with confirmations for sensitive actions.
This is close to the capability profile called Artificial Capable Intelligence last year: tools, computer control, longer-running work, and delegation. The profile is now observable. The label is not a scientific milestone, and capability is not universal reliability. These systems act through products, permissions, connected tools, and security boundaries. They do not possess unlimited authority. An agent may draft a support message or navigate a service only when the user and connected system authorize it.
The jagged frontier
The frontier is jagged. An agent can complete a complex workflow, then miss an obvious button or follow a bad instruction. Stanford’s 2026 AI Index reports top systems succeeding on about two-thirds of structured computer-use tasks—remarkable progress, and failure roughly one time in three. Supervision, confirmations, logs, and secure defaults still matter. AGI is not a settled, verified milestone. Theory-of-Mind AI and superintelligence remain disputed or theoretical—not boxes we can honestly color in as achieved.
The educational turn
That brings us to education. AI can now produce and act. A polished essay, solved problem, finished slide deck, or working app tells us less than ever about what happened inside the learner. The question is no longer only, “Did AI make this?” It is, “What did the learner have to notice, remember, judge, explain, struggle with, and choose?” When the artifact becomes easy, we have to become much more precise about the human capacities the work was meant to build.
Final asset disclosure: Narration generated with ElevenLabs using Matthew Zinn’s consented voice clone.
THE EDUCATIONAL TURN
When AI can produce and act, the artifact proves less.
Protect memory, judgment, dialogue, productive struggle, and agency by making the learning—not merely the output—visible.