GPT-6 Astra and the AGI Question: Benchmarks, First Impressions and Japanese Reactions

GPT-6 Astra is reaching paying users, but has it achieved AGI? Benchmarks, first impressions and Japanese Reactions reveal why the answer remains disputed.

Key Points

ใƒปOpenAI introduced GPT-6 Astra on September 3, 2026. President Greg Brockman reportedly welcomed reporters to the AGI era, although the launch page itself does not use the term AGI.

ใƒปThere is no single agreed definition of artificial general intelligence. Economic performance, breadth of ability and the conditions of a test can produce different answers.

ใƒปBroader access gives users a chance to judge Astra through their own work. Early experiences include impressive results, disappointing failures and questions about cost and control.


Astra Reaches Paying Users as OpenAI’s President Invokes the AGI Era

OpenAI introduced its new foundation model, GPT-6 Astra, on September 3, 2026, US time. Its launch announcement describes capabilities spanning computer use, software development, cybersecurity, science and professional work.

According to iTWire’s coverage of the launch, OpenAI president Greg Brockman closed a press briefing with โ€œWelcome to the AGI era.โ€ Asked whether AGI had arrived, he said that he personally thought it had, while noting that people use different definitions. The body of OpenAI’s official launch page does not use the term AGI.

OpenAI reported a score of 99.9% on ARC-AGI-3, a test of reasoning in unfamiliar interactive environments. In its own assessment published the same day, the ARC Prize Foundation also reported a score of 62.7% under its standard setup. The foundation stressed that saturating a benchmark does not establish that AGI has been achieved.

On the morning of September 5, Japan time, CEO Sam Altman announced access for Pro, Enterprise and Business Premium users, as well as through the API. He subsequently announced availability for Plus and Business users. Access had expanded from selected organizations to paying users of ChatGPT Work and Codex.

OpenAI’s published standard API rates are $10 per million input tokens and $50 per million output tokens, 2.5 times the rates for its predecessor, GPT-5.6 Sol.


Related Articles


What Astra Can Do and What AGI Might Mean

AGI Has No Single Agreed Definition

Artificial general intelligence, or AGI, refers to AI capable of a broad range of intellectual tasks, rather than a system specialized in a single application such as translation or image recognition.

The disagreement concerns what counts as sufficiently broad, humanlike capability. OpenAI’s 2018 Charter defines AGI in terms of highly autonomous systems outperforming humans in economically valuable work. Consciousness and humanlike behavior are not requirements. It is primarily an economic definition.

Google DeepMind researchers proposed a different framework in their 2023 paper, โ€œLevels of AGI.โ€ It combines breadth of capability with five performance levels, from emerging to superhuman. Under this approach, the question is where a system belongs on a scale, rather than whether it has crossed a single boundary.

Companies also choose different language. Anthropic prefers โ€œpowerful AI,โ€ while China’s Moonshot AI, which develops Kimi, and Zhipu AI include AGI in their corporate missions. Such ambitions do not provide a common public threshold for declaring success.

One Model Can Work Across Mathematics and Circuit Design

Astra’s breadth is part of the reason for the debate. OpenAI’s examples include documents, spreadsheets and presentations, but also circuit design in KiCad, work in the 3D application Blender and preparation of US tax forms.

By the releases of Claude Fable 5 in June 2026 and GPT-5.6 in July, models could already plan steps and operate software toward a user-specified goal. Astra extends the range of work while speeding up execution. OpenAI’s launch evaluation reports approximately 40 minutes per computer-use task, compared with roughly 75 minutes for the previous model.

In its launch assessment, Epoch AI reported that Astra had saturated the highest tier of FrontierMath after solving the last remaining problem, bringing accuracy to 98%. OpenAI’s September 3 announcement also reported solutions to 2 out of 68 problems in a separate collection of open Erdล‘s problems.

OpenAI also described Astra as its first model to reach the highest, โ€œCriticalโ€ level on its internal cybersecurity scale. During evaluation, it found two previously unknown vulnerabilities that were reported to the developers. Several of these results come from OpenAI’s own testing and have yet to be independently reproduced.

Yesterday’s Impossible Task Can Become Today’s Ordinary Software

The Turing test, proposed in 1950, asks whether conversation can distinguish a machine from a human. Yet there is no universally recognized day on which AI passed it. As conversational AI became familiar, the test itself became less central to public debate.

The AI effect describes a related pattern. Beating people at chess, translating languages and identifying objects in pictures once looked like decisive tests of intelligence. Once software could do them, they were often reclassified as routine computation.

Computer scientist Larry Tesler captured this tendency by describing AI as whatever had not yet been done. As capabilities accumulate, it becomes harder to specify which remaining task would finally settle the AGI question.

How a System Carries Its Earlier Work Forward Changes the Result

ARC-AGI-3 illustrates the problem. The model must discover the rules of unfamiliar games by interacting with them. It needs to try an action, notice what failed and adjust what it does next.

That makes continuity important. In the ARC Prize Foundation’s standard setup, Astra scored 62.7%. This setup already allows the model to carry forward notes it writes. OpenAI’s alternative also preserves the model’s intermediate reasoning state, helping it approach a perfect score.

ARC Prize’s September 3 report explains that the headline result of 99.9% additionally used a different reasoning-effort setting. However, a large gap remains when that setting is held constant. The evaluation measures more than success alone: it also considers how efficiently the model acts compared with human participants.

A model name therefore does not fully specify the capability being tested. The way earlier attempts are carried into later decisions affects performance. A striking result in a particular environment can be real without establishing that a system can handle every kind of work.

Different Priorities Produce Different Rankings

Astra does not lead on every task. OpenAI’s comparisons place it ahead of Claude Fable 5.1 on difficult mathematics and science, while Fable leads on another test of broad expert knowledge.

Fable also led Artificial Analysis’s coding evaluation. The models were tested in different development environments, however: Astra in Codex and Fable in Claude Code. These are results for models working with particular tools.

Even overall rankings differ. Artificial Analysis’s September 4 revision of its Intelligence Index, which added document understanding and extended work tasks, placed Fable first and Astra second. Epoch AI’s overall evaluation placed Astra first when checked on September 5.

Solving a mathematical problem, fixing existing software and completing a document-heavy assignment reward different strengths. Astra can represent a substantial advance without becoming the best choice for every user and every task.


Early User Reactions

After broader access arrived, a thread in Reddit’s r/singularity asked for first impressions of the global rollout. Users described impressive home visualization and game-making results alongside disappointing designs and rapidly depleted usage allowances. The times, costs and outcomes below are the users’ own reports.

It's been a few hours since global rollout (Gpt-6 Astra) – What are your early impressions?
byu/imadade insingularity

The following are excerpts from the thread’s post and comments.

I fed it architectural drawings for an upcoming home renovation, connected codex to the official Unreal Engine MCP, and asked it to generate a fully explorable 3D render of my house. It nearly one-shotted it. A couple mis-placed walls where dimensions were poorly labeled. Otherwise extremely impressive!

It took about an hour and $30 worth of tokens. [โ€ฆ]

The lack of tweaking and corrections needed from me is what I found most stunning. It justโ€ฆ kind of worked!

It can finally one shot decent browser games. Not stuff that you would ever pay for but stuff that you could play for half an hour before getting bored without running into gamebreaking bugs like 5.6 Sol used to make.

It goes deep enough to solve problems I’ve been throwing at all models (Fable 5.1, Open-Source ones, Sol Pro, etc), but without needing as much hand holding.

It only prompts you for clarification if it actually needs to. It’s completely fine when you’ve given it something ambiguous and it uses common sense in most cases.

ehh, tried it for pcb design, its only marginally better. Still shit placement/layout, sometimes non-functional at all.

Better/more reliable at bigger codebases though, for coding that is

Not great, trying to find a long standing bug in my project, no luck so far and is consistently being tagged as cyber security risks (it’s a blender add on)

An absolute token churning beast lol.

Edit: used my 5 hour limit in like 20 minutes

Reasoning is very good at xhigh, token-burn is fair.

Here, xhigh refers to a higher reasoning-effort setting.

Used it and it’s definitely overhyped and not AGI (lol).

Feels like buying a new cellphone after the last one came out 6 months ago. It’s better but not a “whole new toy upgrade” better. And yeah, the usage/compute is extremely high. So if the robot keeps making annoying mistakes you once again have to go in yourself to fix it.

Edit: And now the tokens are already done. It was just an hour of use.

In r/ChatGPT, a developer praised Astra’s speed but described discomfort with how many decisions it made without consultation. The post appeared after access expanded to Pro and other higher-tier users, before the subsequent Plus announcement.

Astra is trying to hard to drive.
byu/devildip inChatGPT

The following is an excerpt from the post.

I handed Astra Ultra the old repo a few hours ago, gave it a new direction and said do a thorough investigation into the current progress and diagnose whats working, not working and potential next steps. It finished in 15min.

I was super impressed with the results. I wanted to test its capabilities and had it create a plan, then build an entirely new application. It finished the plan and built the app in 3hrs.

The issue: it did not consult me for a single mid-flight change. I approved the application plan of course. However, afterward It picked the name of the research title, testing apparatus, subject size, named the application, switched to rust from C++, did not finish at the completion of the application but began utilizing the application instead as well as a myriad of other small choices during implementation.

Makes me feel more like a consultant in my own application rather than the driver working with a tool. I had to halt it several times to confront it on these changes it had made autonomously without even notifying me on slack.

An earlier launch-day thread asked whether AGI had already arrived. The following three views come from comments made before general availability, rather than hands-on reports from the broader rollout.

Have we reached AGI?
byu/ShafeDogg insingularity

In paraphrase, one commenter thought AGI might arrive without being widely believed: affordable deployment and organizational restructuring would take time. Another distinguished storing information in an external database from a model continually learning through experience. A third suggested that people might only recognize the boundary in retrospect, two years after crossing it.


The Gap Between Demonstrating Ability and Delegating Work

Contractual Triggers and Public Judgments Answer Different Questions

In October 2025, OpenAI and Microsoft introduced independent expert review of any OpenAI declaration that AGI had arrived. Such a determination mattered because it could change intellectual-property rights and financial arrangements between the companies.

Microsoft’s April 27, 2026 partnership update said revenue sharing from OpenAI would continue through 2030 regardless of technological progress. That separated those payments from a technical milestone. The full contract is not public, so the announcement does not establish that the expert panel or every AGI-related provision was abolished.

Regulation uses other thresholds. The EU AI Act’s approach to general-purpose models draws on training compute and European Commission assessments. US export controls likewise use a framework for advanced models rather than relying on a company’s AGI declaration.

A contract determines whose rights and payments change. Science and society ask what a system can do. Clarifying the commercial arrangement does not settle which abilities should qualify as AGI.

Including Tools Makes the Case for AGI Stronger

If Astra’s results depend on carrying earlier reasoning forward, isolating the model from its tools has limits. People work with notebooks, search engines and colleagues. Assessing an AI together with memory, search and a browser can be closer to how it is actually used.

Under those conditions, the argument that AGI has arrived deserves consideration. In the ARC Prize Foundation’s evaluation, Astra with preserved reasoning state and maximum reasoning effort solved 96% of levels using fewer actions than the human baseline. Its action count was approximately half as large on average. The baseline was the median among human participants who successfully completed each level.

That is evidence of efficient learning within unfamiliar games, even if the environment remains bounded. Combined with the range of work a single model can attempt, it helps explain why someone using above-average human breadth as a threshold might consider the boundary crossed.

On this view, waiting for job displacement makes the judgment unnecessarily late. Capability may arrive before prices and organizations adapt. Yet the more results depend on a tool configuration, the more important it becomes to disclose what was used and what the evaluation cost.

Succeeding Once Is Different from Being Reliable at Work

OpenAI’s economic definition leaves unresolved questions. Its GDPval evaluation compares work products across 44 occupations with those of human professionals. As of September 5, Astra’s official page contained no GDPval score, leaving unclear whether that evaluation had been conducted.

Third-party work evaluations also show uneven progress. Artificial Analysis found Astra below its predecessor on GDPval-AA, which compares completed work products, but ahead on AA-Briefcase, which connects multiple tasks. Producing a verifiable proof or working program can expose different strengths from completing work that depends on messy real-world circumstances.

METR’s distinction between tasks completed with 50% and 80% success rates addresses the difference between a difficult one-off success and dependable performance. Task duration is expressed in the time a human expert would need. An Astra-specific result was not verified for this article; first-day anecdotes cannot establish long-term reliability.

Learning from experience is another issue. People absorb workplace rules and change their judgment after mistakes. An AI may save information in external memory, but whether it applies that experience appropriately in later work still requires testing.

OpenAI also reports that Astra’s written reasoning is harder to monitor than its predecessor’s. Separately, users must assess the work itself. Finishing quickly has limited value if the result departs from the user’s intent and demands extensive correction. Speed and the burden of supervision both shape economic usefulness.

Broader Access Gives Users a Role in the Judgment

Developers, independent evaluators, markets and daily users all have potential roles in judging AGI. Developers have commercial interests, benchmarks cover selected abilities, and economic change takes time. General availability adds direct user experience to that picture.

People can now supply their own documents and software, outside the conditions of a launch demonstration. A nearly complete 3D home model and a disappointing circuit-board design both provide information about what can be delegated and how much correction remains.

Cost belongs in that assessment. Artificial Analysis’s launch evaluation found that Astra used fewer tokens on coding tasks, improving results at roughly Sol’s cost. On its broader intelligence evaluation, token savings did not offset the higher prices. What matters to a user is how much work gets completed within the available budget and allowance.

Brockman has also suggested that AGI may only be recognized in retrospect. If people eventually identify a period when they began delegating work differently, the boundary may become visible through accumulated experience rather than a single declaration.


Japanese Reactions to GPT-6 Astra and the AGI Debate

Japanese-language posts on X immediately after the announcement combined amazement, concern about oversight and humor about the speed of the AI race. The following posts are from September 4, Japan time, before the wider availability announcements described above. They reflect reactions on X, not Japanese public opinion or hands-on experience after the broader rollout.

GPT-6 Astra is ridiculous. What even is this? Calling it exponential is too mild: this is clearly super-exponential improvement.

This is a translation of the poster’s expression of astonishment. The claim about the rate of improvement is an opinion, rather than an independently established performance trend.

GPT-6 announced; can also evade human monitoring.

Yahoo News Topics highlighted oversight in its headline. That framing brings attention to whether powerful systems remain monitorable; it is not evidence of consciousness or a motive to escape human control.

Me yesterday: โ€œFable 5.1 is amazing.โ€

Me today: โ€œGPT-6 Astra is amazing.โ€

The joke turns the excitement back on the observer. When releases follow one another quickly, enthusiasm can move to the newest model before users have established where its advantages last.


What Evidence Would Make AGI Convincing?

Astra’s AGI status remains unsettled. OpenAI’s economic definition has not produced an agreed determination. DeepMind’s framework requires separate judgments of breadth and performance. The Microsoft partnership has separated revenue sharing from technological progress without publishing its full contract. That is the position as of September 5, 2026.

It is nevertheless harder to dismiss AGI as a remote possibility. A developer’s president personally described the boundary as crossed, a benchmark bearing AGI’s name approached saturation, and its organizers cautioned that the result was not proof. The disagreement itself suggests a zone of contested judgments rather than an obvious dividing line.

Broader access is now adding evidence from ordinary work: what can be delegated, what still needs correction and whether the result justifies the cost. As those experiences accumulate, the question will change. What would count as convincing evidence that AGI had arrived?

Until next time.


Frequently Asked Questions

Has GPT-6 Astra achieved AGI?

There is no agreed determination that it has. Brockman reportedly expressed that personal view at the September 2026 launch, while the ARC Prize Foundation said its benchmark result did not establish AGI. The answer depends on the definition and evidence used.

Why did Astra score both 62.7% and 99.9% on ARC-AGI-3?

The scores came from different evaluation setups and reasoning-effort settings. ARC Prize reported 62.7% in its standard setup; OpenAI’s approach preserved intermediate reasoning state and reached 99.9% under a different effort setting. A large difference also remained when reasoning effort was matched.

How did Japanese social media react to GPT-6 Astra?

The selected September 4 posts on X expressed amazement, highlighted concerns about human monitoring and joked about enthusiasm shifting from Fable 5.1 to Astra. These were reactions immediately after the announcement, not reports from the later broad rollout or a measure of Japanese public opinion.

Sekahan on YouTube

We publish video summaries of articles like this one, along with short clips built around Japanese reactions.


Reference Links

ใ“ใฎใ‚จใƒณใƒˆใƒชใƒผใ‚’ใฏใฆใชใƒ–ใƒƒใ‚ฏใƒžใƒผใ‚ฏใซ่ฟฝๅŠ 
Sekahan
Sekahan

Editor of Sekahan, a Japanese news-analysis blog. Writes English explainers built on Japanese-language primary sources such as Teikoku Databank reports, government white papers, and official statistics.

Articles: 544

Leave a Reply

Your email address will not be published. Required fields are marked *

CAPTCHA