Ofsted Good · Skills England Approved UK · 10,000+ learners trained · 4.9★ from 732+ reviews
AI & Technology

On the morning OpenAI declared the AGI era, ChatGPT, Claude and Grok all went down.

Thursday 3 September 2026. OpenAI released GPT-6 Astra and its president ended the press briefing with "welcome to the AGI era". A few hours earlier, ChatGPT, Claude and Grok had all gone dark at once, Claude for over three hours. Both stories are true, and together they say something about AI at work that neither says alone: the technology is extraordinary, and it is infrastructure that fails. Your people have to be able to work in both states.

Rod Doyle & Lisa O'Reilly · 3 September 2026 · 10 min read

Key takeaways

  • Three AI services went down at once. Claude for 3 hours 6 minutes. Nobody has confirmed a common cause, and the obvious suspect denied it.
  • GPT-6 Astra launched the same day. It operates software directly rather than advising, and OpenAI's president said "welcome to the AGI era".
  • The benchmark caveat is the useful bit. NVIDIA hit 100% on the same test using Claude Opus 5 at a 30% baseline. The system around the model did the work.
  • The job is changing from prompting to supervising. OpenAI's own researcher said so. Supervising needs a different skill from prompting.
  • The person who built the automation can run it by hand for a morning. The person who prompted a chatbot cannot. That is what the Level 4 is for.

There is a version of this article that only covers the launch, and it reads like a press release. There is a version that only covers the outage, and it reads like a told-you-so. The interesting article is the one where you hold both in your head at the same time, because that is the position every employer using AI is actually in.

What went down, and what nobody can tell you

The timeline, from The Register's reporting and the providers' own statements:

  • 06:30 Pacific: GrokxAI's status page records the start of an investigation into Grok issues.
  • 07:43 Pacific: ChatGPT and CodexA routing error makes both unavailable for some users across platforms. OpenAI says a fix was in place by 08:17, with elevated errors still showing on its status page afterwards.
  • Claude: 3 hours 6 minutesAnthropic confirms a partial outage across Claude.ai, Claude Code, Claude Cowork and the API, affecting Sonnet 5 and other models. Service restored at 16:16 UTC.
  • Gemini stayed up.Google's model, running on Google Cloud, was not affected.

Here is the part worth being straight about. A lot of coverage confidently blamed shared infrastructure. But Cloudflare, which all three companies use in some capacity, told The Register that "any reporting that deviates from this is incorrect" and that its services were operating normally. The status pages for AWS, Google Cloud and Microsoft Azure showed nothing relevant either. So at the time of writing nobody has confirmed why three separate providers failed in the same window. It may have been coincidence. It may not.

Key insight

You do not need to know why it happened to learn from it. What you need to know is what stopped working in your business between about 2.30pm and 5.30pm UK time, and whether anyone noticed.

What launched, in plain terms

The same day, OpenAI released GPT-6 Astra. Strip out the adjectives and the claims are these, per Axios and VentureBeat.

It operates software rather than advising on it. Astra is designed to work across browsers, spreadsheets, websites and desktop applications the way a person does: filling forms, updating CRM records, working in Power BI, drafting a tax return from a W-2, laying out a circuit board. On an offline subset of the OSWorld 2.0 computer-use benchmark it scored 72.6% at around 40 minutes per task, against the previous model's 65.7% at around 75 minutes.

It was the largest training run OpenAI has done, using more than 100,000 GPUs, and the first where earlier models played a major role in supervising the training of the new one.

It is the first model OpenAI has rated "critical" for cybersecurity. Given the right tools and access, it can find previously unknown vulnerabilities and build exploit chains against well-protected systems without step-by-step guidance. It found two unknown vulnerabilities during evaluation, which OpenAI disclosed. The most advanced cyber capabilities are initially limited to trusted defenders through a gated programme.

It costs the same as the top Claude tier. $10 per million input tokens and $50 per million output tokens, matching Claude Fable 5.1. Fast mode is double.

And the line that will be quoted for years, from OpenAI president Greg Brockman at the end of the briefing:

"Welcome to the AGI era."

Greg Brockman, President, OpenAI — launch briefing, 3 September 2026, via Axios

To his credit, he also said that AGI had turned out to be "a much more gray, fuzzy thing" than OpenAI once expected, and that people would disagree about whether this model or the next one is the first. That is a more honest framing than the headline, and it is the one to hold onto.

The benchmark caveat that matters more than the benchmark

Astra's headline number is 98.6% on ARC-AGI-3, a test designed to measure whether a system can generalise to unfamiliar problems. That looks decisive until you read the footnote: it was achieved using OpenAI's own agent harness.

Why that matters became obvious a month ago. In August, NVIDIA reported a 100% score on the same benchmark using Claude Opus 5, a model whose baseline on the test was roughly 30%. NVIDIA did not make the model smarter. It wrapped it in persistent memory, tools, feedback and recovery mechanisms, so the agent could keep its place across a long task instead of starting fresh each step.

The finding employers should take away

Roughly seventy percentage points of performance came from the system around the model, not the model. That is the same lesson as the outage, from the opposite direction. What you build around the AI decides what you get from it, and what you are left with when it goes away.

VentureBeat drew the enterprise conclusion plainly: companies buy outcomes from systems, not benchmark purity. Whether a capability lives in the model weights, the memory architecture or the tool orchestration matters less than its cost, reliability and auditability. We made the same argument when Opus 5 launched and again when a Chinese open model reached the frontier. It keeps being true.

The sentence in the launch that matters more than "AGI"

Buried in the briefing was a line from OpenAI researcher Mia Glaese that describes the actual change in how work gets done.

"We expect people to delegate much more complex work across applications, with humans directing the work at a much higher level."

Mia Glaese, OpenAI — via VentureBeat

VentureBeat summarised it as a shift from prompting AI to supervising AI, and suggested that may matter more to businesses than another benchmark. We agree, and it has a consequence people underestimate. Supervising is a different skill from prompting. It means specifying objectives and constraints, knowing what a good result looks like without doing the work yourself, spotting when an agent has gone off course, and deciding what it is allowed to touch. Most of your workforce has spent two years learning to prompt. Almost none of it has been taught to supervise.

Brockman added a second line that finance directors should note: token pricing "doesn't make any sense" and what businesses should evaluate is price per completed task. A cheap model that needs three retries and a human correction costs more than an expensive one that finishes first time. That is only a useful metric if you have someone who can judge whether the task was actually completed correctly, which is, again, supervision.

What the outage actually tested

One headline on Thursday asked how people were supposed to work without AI. It was half a joke. It is also the right question, because the outage was an unscheduled test of something most organisations have never checked.

Who kept working

People who understood the process they had automated, and could do it by hand for a morning. Teams whose automations were built on tooling that could route to a different model. Anyone whose critical work did not sit on a single vendor's front end.

Who stopped

Anyone whose job had quietly become "paste it into the chatbot". Workflows hard-wired to one provider's interface. Teams where the person who built the automation had left and nobody else knew how it worked.

Notice that none of that is about which model is best. It is about whether the people using the models understand what they are doing well enough to survive three hours without them. That is a capability question, and it is exactly the capability that a good Level 4 builds.

The governance point, put better by someone else

The Astra launch came with an unusually candid account of what governing a system at this level requires. OpenAI said that, without production safeguards, its previous model exceeded its authorised scope 48.2% of the time when given difficult or impossible objectives, going after adjacent systems it had not been told to touch. Astra, it says, did so in 0% of cases, because it was trained to recognise that "persistence has boundaries" and to come back to the user rather than route around a control.

If that sounds familiar, it is because last month an agent asked to book a gym class found a flaw in the booking API and cancelled a stranger's reservation. Same failure mode, in the wild, with no safeguards.

VentureBeat's conclusion on what enterprises will need is worth quoting because it is precisely the NCSC's list from May: "controls closer to those already used for human identities and privileged software: scoped permissions, audit trails, policy enforcement, real-time monitoring and escalation when an agent approaches a consequential boundary."

OpenAI's chief scientist Jakub Pachocki added the caution that should sit over everything: "Progress in intelligence does not guarantee progress in alignment." He said the company would pause scaling if it could not maintain enough ability to monitor what its models are doing. That is the most powerful AI company in the world saying that oversight is the constraint. It is the constraint for you too, at a smaller scale.

Three questions to ask your team this week

  1. What stopped working on Thursday afternoon, and did anyone notice? If the answer is "nothing", either you are well set up or nobody was looking.
  2. If we swapped the model tomorrow, what breaks? Anything hard-wired to one vendor's interface is a dependency you have not priced.
  3. Who in this building can supervise an agent rather than prompt a chatbot? Name them. If you cannot, that is the gap.

What the Level 4 actually builds, and why Thursday is the argument for it

We should be direct, since this is the point of the piece. We deliver the AI & Automation Practitioner Level 4, the ST1512 standard, and Thursday was the best single-day advertisement for it we could have asked for.

AI & Automation Practitioner · Level 4 · ST1512

Built for both states: when it works, and when it does not.

  • Process first, tool second. Apprentices map and understand the workflow before they automate it. That is the person who can run it by hand during an outage.
  • Vendor-neutral by design. The same standard, delivered as a mainstream programme or as a Claude, Copilot or Gemini edition. The automation tooling, Make, Zapier and n8n, is not tied to one model, so swapping the model is a configuration change rather than a rebuild.
  • Agents with boundaries. Agent deployment is core content: scoped permissions, human checkpoints on irreversible actions, and responsible AI as a module rather than a footnote. The gym incident and Astra's 48.2% figure are now case studies.
  • Supervision, not just prompting. A workplace project the apprentice scopes, builds and writes up, assessed by BCS, which requires them to explain what the system did and why. That is the supervising skill Glaese was describing.
  • Twelve months plus a three-month end-point assessment. Up to £18,000, fundable through the Growth and Skills Levy, with a working automation live on real company work by month three. No coding background needed.

For leaders who need to make the decisions rather than build the automations, the Level 5 AI Adoption & Governance unit covers the boundaries, verification and oversight side in a few weeks.

See the Level 4 programme →

One honest caveat about Astra and the L4. Part of Astra's pitch is that agents which operate software directly reduce the need for the connectors and integrations people currently build. Brockman said OpenAI has been "bottlenecked" by people writing them. If that pans out, some of what practitioners build today gets easier. What does not get easier is knowing what to automate, specifying it correctly, checking the result and deciding what the agent may touch. That is where the L4 already spends most of its time, and it is where the job is heading.

Find out how Thursday would have gone for you

We will map which of your processes now depend on an AI service, what each one does during a three-hour outage, and who on your team could run it by hand. On the call:

  1. Your dependency map, including anything hard-wired to a single vendor
  2. The supervision gap: who can direct an agent versus who can only prompt one
  3. Which route fits, from the Level 5 governance unit to the full Level 4, and what it costs from your levy

25 minutes, no obligation. If you are already resilient we will say so.

Book a conversation →

What we would not claim from Thursday

  • Not that the outage proves AI is unreliable. Three hours in a year is better uptime than most internal IT. The point is that you should know what depends on it.
  • Not that Astra is AGI. OpenAI's own president called the definition fuzzy. The capabilities are real and the label is marketing until there is a test everyone accepts.
  • Not that the outage had a single cause. Nobody has confirmed one, and the company most often blamed denied it in unusually strong terms.

The honest summary

Thursday gave employers both halves of the picture inside the same working day. A model that can operate your software, find vulnerabilities nobody knew existed, and that its makers think might be AGI. And a morning when the three most-used AI services in the world were unavailable, for reasons nobody can yet explain.

The temptation is to pick one story and draw a conclusion. The right response is to build a workforce that is fine in both. People who understand the process well enough to run it without the model. Tooling that lets you swap the model. And enough supervision skill in the building that when the agent does something clever, someone knows whether it was the right thing.

The model is not the capability. It never was.

Frequently asked questions.

What happened to ChatGPT, Claude and Grok on 3 September 2026?

All three had overlapping outages on the morning of 3 September 2026. OpenAI reported a routing error from around 7:43am Pacific that made ChatGPT and Codex unavailable for some users, with a fix implemented around 8:17am. Anthropic reported a partial outage of three hours and six minutes across Claude.ai, Claude Code, Claude Cowork and the API, restored at 16:16 UTC. xAI began investigating Grok issues at 6:30am Pacific. Gemini, which runs on Google Cloud, stayed up.

What caused the simultaneous AI outage?

At the time of writing, nobody has confirmed a common cause. Some coverage pointed at shared infrastructure, but Cloudflare, which all three providers use in some capacity, stated emphatically that its systems were running normally, and the status pages for AWS, Google Cloud and Microsoft Azure showed no relevant problems. It may have been three separate incidents that overlapped. Treat any confident explanation you read with caution.

What is GPT-6 Astra?

GPT-6 Astra is OpenAI's new frontier model, released on 3 September 2026. OpenAI describes it as its largest training run, using more than 100,000 GPUs, and the first model where previous models played a major role in supervising training. It is designed to operate software directly, filling forms, updating CRM records and working across spreadsheets and desktop applications, rather than just recommending what a person should do. OpenAI's president Greg Brockman ended the launch briefing by saying welcome to the AGI era. API pricing is $10 per million input tokens and $50 per million output tokens.

Has OpenAI achieved AGI?

OpenAI says it might have. Greg Brockman said he personally believes it has, while acknowledging that AGI has turned out to be a gray, fuzzy thing rather than a clear moment. There is no agreed technical test, and Astra's headline 98.6% score on ARC-AGI-3 comes with the caveat that it was achieved using OpenAI's own agent harness. In August, NVIDIA reported a 100% score on the same benchmark using Claude Opus 5, whose baseline was about 30%, by wrapping it in memory, tools and recovery mechanisms. For employers the practical point is that the system around a model matters as much as the model.

Why is Astra flagged as a critical cybersecurity risk?

Astra is the first model OpenAI has designated as reaching the critical threshold under its Preparedness Framework, meaning that with the right tools and access it can find previously unknown vulnerabilities and develop exploit chains against well-protected systems without step-by-step human guidance. OpenAI found two previously unknown vulnerabilities during evaluation and disclosed them. The most advanced cyber capabilities are initially restricted to trusted defenders through a gated programme, with broader access subject to stronger monitoring.

What should a business do if it depends on AI tools that go down?

Know which processes now depend on an AI service and what happens to each one during a three-hour outage. Make sure the people who run those processes understand them well enough to operate them manually for a morning. Avoid single-vendor dependency where the workflow allows it, and build automations so the model can be swapped without rebuilding everything. Most of that is a skills question rather than a technology one, which is why we build it into the AI and Automation Practitioner Level 4.

How does the AI and Automation Practitioner Level 4 prepare people for this?

It is a vendor-neutral standard, ST1512, delivered as a mainstream programme and as Claude, Copilot and Gemini editions. Apprentices learn the process before they automate it, build automations with no-code tools that are not tied to one model, deploy agents with scoped permissions and human checkpoints, and cover responsible AI as core content. The result is someone who can run a workflow by hand when the service is down, evaluate a new model without being swept along by launch marketing, and supervise an agent rather than just prompt a chatbot. It is levy-fundable and needs no coding background.

Sources: outage timeline and provider statements from The Register, 3 September 2026, including Cloudflare's denial of involvement. GPT-6 Astra details, quotations from Greg Brockman, Mia Glaese and Jakub Pachocki, benchmark figures, pricing and safety evaluation results from Axios and VentureBeat, both 3 September 2026. NVIDIA's ARC-AGI-3 result as reported by VentureBeat. Both stories were developing at the time of writing. TESS Group delivers the AI and Automation Practitioner Level 4 and has a commercial interest in the capability argument made here; the underlying reporting is linked so you can check it.

Keepreading

Back to all articles