Key Points
ใปOn September 22, 2026, OpenAI released GPT-6 Sol and Luna at half the current price of the previous generation, and on the same day Anthropic released Claude Opus 5.5, which it describes as performing at the level of Claude Fable 5.1 on most work while costing 40% less to run than Opus 5 on typical workloads.
ใปThree weeks after GPT-6 Astra and Fable 5.1 took the top spots, models that their makers describe as carrying frontier-level capability appeared at much lower price points, and the number that matters is shifting from the price of a token to the cost of finishing a task.
ใปThe cheaper each unit of AI becomes, the more of it gets used, and electricity supply and grid connections may not grow fast enough to match. The yardstick of competition may shift toward how much useful work can be extracted from a limited supply of power.
OpenAI and Anthropic Launch New Models on the Same Day, With Top-Tier Performance at Around Half the Price
OpenAI released two new models, GPT-6 Sol and GPT-6 Luna, on September 22, 2026 (US time). According to OpenAI’s announcement, API prices per million tokens (the units into which AI models break text) are $2 for input and $10 for output for Sol, and $0.10 for input and $0.50 for output for Luna, which the company describes as a 50% cut from the current pricing of the previous generation, GPT-5.6.
OpenAI said both models were trained with similar methods to GPT-6 Astra, its top model released on September 3, and that they bring the advances behind Astra to faster, more affordable models.
On the same day, Anthropic released Claude Opus 5.5 at $4 per million input tokens and $20 per million output tokens. Anthropic said the model performs at the level of its top model, Claude Fable 5.1, on most work and, at default settings, costs 40% less to run than Opus 5 on typical workloads.
Anthropic also raised the five-hour usage limits on its paid plans. Its announcement carries a note that “at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences.”
Fable 5.1 was released on September 1 and Astra on September 3, so both companies’ new models arrived three weeks later. As of September 24, 2026, no independent evaluation comparing the two companies’ models on the same tasks under the same settings had been published.
Related Articles
How the Price of Intelligence Is Set, and Why It Falls Within Weeks
A token is a small chunk of text, often a word or part of a word, and API prices for large language models are quoted per million tokens. What became cheaper on September 22 is the pay-as-you-go API price that companies and developers pay, not the flat monthly subscriptions sold to individuals.
What “cheaper” actually means
The $20 a month ChatGPT and Claude subscriptions did not change in price. The benefit reaches those users as larger usage allowances within each five-hour window rather than as a lower bill.
API bills are built from three prices: input (what the model reads), output (what it writes), and cached input (a discounted rate for reading the same preamble again, known as prompt caching). If the price per token halves but a task uses twice as many tokens, the total cost of that task does not change.
There is more than one way to move top-tier capability into cheaper models
The best known method is distillation, a training technique in which a smaller model learns to answer the way a larger model does. The “illicit distillation” covered in our September 13 article refers to using another company’s model without permission. Distillation inside a company, using its own models, is a standard and legitimate technique used across the industry.
There are other routes as well: more efficient inference, lighter model formats such as quantization, and improvements in architecture. OpenAI said it trained Sol and Luna “with similar methods as GPT-6 Astra” and that “improvements in caching and inference let us serve these models at lower cost.” Anthropic said Opus 5.5 needs less compute to run than Opus 5. Neither company has said it used distillation to build these models.
Related article
The price ladder was rebuilt in three months
According to the companies’ published price lists, the top of the market in July was GPT-5.6 Sol at $5 input and $30 output, and Claude Fable 5 at $10 and $50. In early September, Astra and Fable 5.1 took the top spots. Three weeks later, performance that both companies describe as top-tier was placed at the $2 and $4 input price points.
| Company | Model | Input | Cached input | Output | Position |
|---|---|---|---|---|---|
| OpenAI | GPT-6 Astra | 10.00 | 1.00 | 50.00 | Top model (September 3) |
| OpenAI | GPT-6 Sol | 2.00 | 0.20 | 10.00 | Everyday workhorse (September 22) |
| OpenAI | GPT-6 Luna | 0.10 | 0.01 | 0.50 | High-volume processing (September 22) |
| Anthropic | Claude Fable 5.1 | 10.00 | 0.25 | 50.00 | Top model (September 1) |
| Anthropic | Claude Opus 5.5 | 4.00 | 0.20 | 20.00 | Described as Fable 5.1 level (September 22) |
| Gemini 3.1 Pro (preview) | 2.00 | 0.20 | 12.00 | Top model still in preview | |
| DeepSeek | V4.1 Flash | 0.30 | 0.006 | 1.20 | Budget model (price cut on September 10) |
| Moonshot | Kimi K3 | 3.00 | 0.30 | 15.00 | China’s top-tier model |
US dollars per million tokens, standard rates from each company’s official pricing page as of September 24, 2026. According to OpenAI’s announcement, the “50% cut” is measured against the promotional prices in effect since August. Against July’s regular prices, the cut is more than 60% by our calculation. Each company counts tokens differently, so per-token prices are not a comparison of what the same task costs.
On the user side, a complaint was already common before the price cuts: the models are brilliant, but the usage allowance runs out quickly. Fable 5.1 and Astra were highly rated, yet users reported exhausting even a $200 a month plan within days. Other users took the opposite view, reasoning that a pricey model is still cheaper than their own labour if it finishes the job faster.
In China, the price war started earlier. According to a Bank of America report described by Business Insider in September 2026, the key economic metric in the Chinese market is shifting “from price per token to cost per completed task.” China is not uniformly the cheapest, either. As the table shows, Kimi K3 costs more than Sol, and Luna’s per-token price is lower than DeepSeek’s budget model.
Related article
GPT-5.6 Arrives as Claude Resets Its Usage Limits: Inside AI’s New Access Race
Data center power gets stuck in two places: generation and transmission
The computers that run AI sit in data centers that draw large amounts of electricity. According to the International Energy Agency’s April 2026 report, global data center electricity consumption grew by nearly 20% in 2025 and is expected to roughly double by 2030.
| Item (IEA, April 2026) | 2025 | 2030 |
|---|---|---|
| Electricity use of all data centers worldwide | 485 TWh | 950 TWh (about double) |
| AI-focused facilities | Up 50% year on year | Three times the 2025 level |
TWh is annual electricity consumption. Figures are from the IEA’s central scenario.
For the United States, Morgan Stanley estimated on September 21, 2026 that data centers will need 97 GW of additional power by 2028 and that, even after mitigation measures, 33 GW of that will be missing. The estimate works backward from semiconductor shipments.
The phrase “not enough power” mixes two different bottlenecks. One is a shortage of power plants. The other is connection to the grid that carries the electricity. According to Lawrence Berkeley National Laboratory’s Queued Up 2026 report, published in May 2026, US generation and storage projects that began operating in 2025 took a median of more than five years from interconnection request to commercial operation. Lead times for large power transformers have reportedly stretched beyond two years. GPUs arrive when ordered, but transformers and transmission lines require community consent and years of construction.
Japan faces the same pressure on a smaller scale. According to the Organization for Cross-regional Coordination of Transmission Operators (OCCTO), Japan’s grid coordinator, in its demand forecast published in January 2026, new and expanded data centers and semiconductor plants will add up to 7.62 GW to Japan’s peak demand by fiscal 2035, of which data centers account for 6.61 GW, or 4.63% of national peak demand. The same OCCTO forecast notes that the start of operations for data centers has been pushed back compared with the previous forecast.
From “the Strongest AI” to “the Intelligence You Can Extract From Limited Power”
What buyers are paying for is the total amount of intelligence
If a 100-point answer costs 100 yen, a 95-point answer 20 yen and an 85-point answer 1 yen, then for most jobs the latter two are worth more. Pulling key points from a contract, routing customer inquiries and catching typos mostly need only a 95-point answer.
“Amount of intelligence” here means the amount of work that can be completed at a given level of quality. The number of tokens that the same 10,000 yen buys has grown sharply in three months, and by the companies’ own evaluations the cost per task has fallen as well. The benchmark note in Anthropic’s announcement amounts to the seller admitting that score gaps near the top no longer work well as a basis for buyers’ decisions.
The usage limit resets and extensions covered in our July article were a stage of competing on “how much you can use” without changing prices. In September the prices themselves moved, and competition is shifting its weight toward “how cheaply can this intelligence be delivered per task.”
Why can some companies cut prices more than others?
Why now? The efficiency of inference rises substantially every year, and the cost of serving a given level of performance keeps falling. On top of that, Chinese open-weight models (models whose weights are published) have set the floor for the low end of the market. According to the companies’ announcements, DeepSeek cut prices on September 10, xAI released Grok 4.7 with better performance at an unchanged price on September 21, and OpenAI and Anthropic followed on September 22. Some commentators have read this price competition, ten days after the proposal to pace the frontier, as a sign that “the slowdown is over.”
The companies’ room to cut differs. Google owns its own chips and cloud, while OpenAI and Anthropic depend on outside clouds and chips and are betting on recovering the price cuts through higher usage.
The opposite view is also reasonable. A list price is a number set by competition and margins, and it cannot be read as a price cut equal to any fall in underlying costs. Users themselves have raised the worry that prices will rise once customers are locked in, and the squeeze on flat-rate plan allowances has continued even after the cuts.
Combining AIs with different roles
When prices split into tiers, the way models are used changes too. Hard problems go to the top model, everyday advanced work to the middle tier, and large volumes of simple processing to the cheapest model, with the top model acting as the coordinator and the lower tiers working as subordinates. Some users already run Sol as a subagent under Astra, and on Japanese X, users describe a split in which Astra sits at the centre and tasks drop down to Sol or Luna depending on difficulty.
Rather than handing everything to one genius, this means building a team of AIs in the shape of managers, specialists and general staff, and the decision about which model gets which task moves to the user.
Why cheaper AI can use more electricity
Dollars per token and watts are different yardsticks. Electricity consumption is the product of “electricity per use” and “number of uses.” The IEA’s April 2026 report says the energy per AI task has fallen by at least an order of magnitude a year in recent years, while also noting that reasoning and agentic workloads use hundreds to thousands of times more energy than simple text generation.
The Jevons paradox is the observation made by the economist William Stanley Jevons in 1865 that improvements in steam engine efficiency increased, rather than reduced, coal consumption.
Where someone once used AI once to check an article, cheaper prices allow them to run 20 copies of Luna, have Sol consolidate the results and ask Astra for a final review. Even if the electricity per call falls, total consumption rises if the number of calls grows a hundredfold.
One company has even named a product after the paradox. TypeSafe AI, founded by a former OpenAI researcher, released a judgment-only model called Jev in early access on September 15, 2026, saying it returns only structured answers such as Yes/No, A/B/C or scores. The company says its input price is less than half of Luna’s and that output is free, and writes that “every order-of-magnitude drop in the cost of intelligence” expands the uses by many orders of magnitude.
These are the product’s own claims with no third-party verification, and user reports that placing it in front of a larger model cut costs by 70 to 80% sit alongside reactions that it is “basically a fast classifier.”
The view that efficiency wins also has grounds. What fell this time is the list price, and there is no evidence that the underlying compute cost or the electricity per task fell by the same proportion. If work moves from large models to small models and judgment-only products faster than usage grows, total consumption need not rise.
Who pays for the electricity is also in dispute. On the concern that the cost of grid upgrades will be passed on to household electricity bills, Axios reported in September 2026 that US lawmakers agree consumers should not bear that cost but are split on whether Congress will actually regulate it.
Is the US China race becoming a race to supply cheap inference at scale?
In compute, the United States holds the advantage. A September 17, 2026 analysis by the Center for a New American Security (CNAS) estimates that the computing power Chinese companies can use through overseas clouds amounts to roughly one twentieth of the world’s total. On the other hand, China generates more than twice as much electricity as the United States, and according to BloombergNEF estimates reported by Al Jazeera in May 2026, China will add about six times as much generating capacity as the United States over the next five years.
If the picture is framed as the United States leading in GPUs, models and cloud while China counters with power supply, manufacturing capacity and efficiency, a new axis of competition comes into view: the advantage may go not to the country that can build the single strongest model, but to the country that can extract the most work from limited chips and power and supply it cheaply at scale. Designs such as Luna and Jev that “use only the intelligence you need” are American products, and they point in the same direction as the path China took first under its constraints.
Views on export controls divide here. Supporters of the controls argue that the ability to train frontier models from scratch depends on advanced chips and manufacturing equipment, so cutting off those inputs is the surest way to slow military use and frontier development. On the other side, a remark reported from a Huawei executive, that the company would not have done this had the United States not cornered it, said with a touch of irony, is often cited as evidence that controls can backfire by spurring domestic production and efficiency.
Ahead of the US China summit on September 24, the US side proposed at preparatory talks on September 20 a mechanism for mutual notification of AI incidents affecting national security and a new US China dialogue on AI, according to the Associated Press. Reports say the summit agenda was also expected to include AI-enabled cyberattacks, Chinese access to top US models and US regulation of Chinese open-weight models. The distillation allegations and export controls on advanced chips are background points of conflict that make these talks harder, and the US side said export controls were not covered in the AI framework at the preparatory talks.
Related article
Japanese Reactions to GPT-6 Sol, Luna and Claude Opus 5.5
These are posts on X, not a measure of Japanese public opinion. The translations are by Sekahan, handles are omitted, and the like counts are those displayed on September 24, 2026 (JST).
Among Japanese developers and AI watchers, the same-day launch was read mainly as a story about Opus 5.5. One widely shared summary, with about 370 likes, came from ML_Bear, a machine learning engineer, who framed Opus 5.5 as bringing Fable-class performance down to an ordinary price range while questioning how much Sol and Luna had actually improved.
ๆ ๅ ใงๆ่ตทใใใGPT-6 Sol/LunaใจOpus5.5ใๅๆใซๆฅใฆใฆไปฐๅคฉใใพใใ็ฌใใใฃใใ่ชฟในใใใจใพใจใ๐
— ML_Bear (@MLBear2) September 22, 2026
ใใใฃใใใใใจใ
ใปOpus5.5ใฏๆง่ฝๅคงๅน ๅไธใงใกใใๅคไธใใFable็ดใๆฎ้ใฎไพกๆ ผๅธฏใซ้ใใใฆใใๆใใ๏ผ
ใปGPT-6 Sol/Lunaใฏๅ้กใซใชใฃใฆใณใฃใใใ ใใฉๆง่ฝใฏใปใผๆจชใฐใ๏ผโฆ
I woke up on a trip to find GPT-6 Sol/Luna and Opus 5.5 had landed at the same time. Roughly speaking: Opus 5.5 is a big jump in performance with a small price cut, as if they brought Fable-class down to the normal price range. GPT-6 Sol/Luna being half price is a surprise, but is the performance mostly flat?
ML_Bear, September 23, 2026 (translated by Sekahan)
That cooler view of Sol and Luna was common. A developer who posts as Masao, focused on AI-driven development, wrote that the new models were barely smarter than GPT-5.6 but that the price made Luna attractive for embedding in apps, and described a split in which Astra handles the hardest tasks and work drops down to Sol or Luna depending on difficulty.
GPT-6 Sol / Luna ใ่งฆใฃใฆใใฎใงๆๆณใ…
— ใพใใ@AI้งๅ้็บ (@AI_masaou) September 23, 2026
ใป่ณขใใฏ 5.6 ใใใปใผๆฎใ็ฝฎใใAA ใฎในใณใขใๆจชใฐใใงใๆญฃ็ดใ5.7ใใงใใใฃใ
ใปAPI ใฏ 5.6 ใฎๅ้กใLuna ใฏ Opus 5.5 ใฎ 1/40 ใงใCodex app server ใง่ชๅใฎใขใใชใซ็ตใฟ่พผใ็จ้ใชใๆฌๅฝใใ
ใปSol ใฏ Sonnet ใจๅใไพกๆ ผๅธฏใง่ฆใใจๆฎ้ใซ่ฏใใLPโฆ
Intelligence is roughly unchanged from 5.6, and honestly it could have been called “5.7.” But the API is half the price of 5.6, and Luna costs 1/40 of Opus 5.5, so for building into your own apps it may be the real pick.
Masao, AI-driven development, September 23, 2026 (summary, translated by Sekahan)
The price cut also unsettled the product line itself. Hiroshi Yuki, a well-known author of programming and mathematics books, asked what the top model is now for.
Opus 5.5ใฎๆง่ฝใFable 5.1ใใใๅชใใฆใใฆใใใๅฎไพกใชใใฐใFable 5.1ใฎๅญๅจๆ็พฉใจใฏโฆ๏ผ๐ค๏ผไฝใๅ้ใใใฆใ๏ผ๏ผ
— ็ตๅๆตฉ / Hiroshi Yuki (@hyuki) September 22, 2026
If Opus 5.5 performs better than Fable 5.1 and is cheaper too, then what is Fable 5.1 for…? (Am I misunderstanding something?)
Hiroshi Yuki, September 23, 2026 (translated by Sekahan)
Some users suspected the price came at a cost in quality. Minorun, a cloud engineer and author of books on generative AI development, found Opus 5.5 weaker at producing documents and guessed that Anthropic had distilled it for coding. Anthropic has not said how Opus 5.5 was made, and its announcement and system card do not describe distillation from Fable 5.1, so this is the poster’s speculation.
Opus 5.5ใใณใผใใฃใณใฐใซ็นๅใใฆ่ธ็ใใฆใณในใไธใใฆใชใ๏ผ่ณๆไฝๆไธๆใใใชใใ ใใฉโฆใSonnetใฌใใซ
— ใฟใฎใใ (@minorun365) September 22, 2026
Didn’t they distill Opus 5.5 specifically for coding to bring the cost down? It’s bad at making documents… Sonnet level.
Minorun, September 23, 2026 (translated by Sekahan)
The Japanese conversation focused on which model to use for what and on how the lineup now fits together. The link between cheaper AI and electricity demand, which runs through this article, had barely entered the Japanese discussion of these launches; posts about power and data centers that did spread were mostly about investment and energy in general.
Before AGI Arrives, the Price of Intelligence and the Supply of Power Are Changing the Economy
On September 1, Anthropic’s Fable 5.1 took the top spot, and on September 3 OpenAI’s Astra followed, prompting talk of a “welcome to the AGI era.” By September 22, part of the capability gained at that frontier had already begun spreading into mid-tier and budget models. What stands out about AI is not only how fast it gets smarter but how fast it gets cheaper.
While people debate when AGI will arrive, the price of the intelligence that already exists is falling first, and the range of work that can be handed to AI on the same budget is widening quickly. Meanwhile, the electricity and grid capacity that run this intelligence do not grow on a cycle of weeks, and that gap may shape what AI looks like over the next several years.
Which companies and countries extract the most work from each watt cannot be measured today, because completion rates and power use are not disclosed in comparable form. Still, as the amount of work that the same 10,000 yen or the same kilowatt-hour can buy keeps growing, the question of what to hand to which AI, and how far, sits closer to home than the race for the strongest model.
Frequently Asked Questions
Is GPT-6 Sol really 50% cheaper than GPT-5.6?
Yes against the promotional price, and by more against the regular price. According to OpenAI’s September 22, 2026 announcement, Sol’s $2 input and $10 output per million tokens are 50% below GPT-5.6 Sol’s promotional pricing of $4 and $20, which has applied since August 21. Against GPT-5.6 Sol’s regular price of $5 and $30, the cut is more than 60%.
Did OpenAI and Anthropic distill Sol, Luna and Opus 5.5 from Astra and Fable 5.1?
Neither company says so. OpenAI’s announcement says Sol and Luna were trained “with similar methods as GPT-6 Astra” and credits improvements in caching and inference for the lower price, while Anthropic says Opus 5.5 needs less compute to run than Opus 5. Anthropic’s Opus 5.5 system card does not describe distillation from Fable 5.1, and claims on social media that the models were distilled are speculation.
How did Japanese social media react to GPT-6 Sol, Luna and Claude Opus 5.5?
Japanese posts on X focused mostly on Opus 5.5. One widely shared summary on September 23, 2026 described Opus 5.5 as bringing Fable-class performance to a normal price range and asked whether Sol and Luna had improved much at all, while other developers questioned what role Fable 5.1 now plays and discussed routing tasks from Astra down to Sol or Luna by difficulty.
Will cheaper AI reduce electricity use?
Not necessarily. According to the IEA’s April 2026 report, the energy used per AI task has fallen by at least an order of magnitude a year, yet global data center electricity use is still expected to roughly double from 485 TWh in 2025 to 950 TWh in 2030, as usage keeps growing. This is the pattern known as the Jevons paradox, though total demand would stop rising if efficiency gains outpaced growth in use.
Sekahan on YouTube
We publish video summaries of articles like this one, along with short clips built around Japanese reactions.
Reference Links
- Introducing GPT-6 Sol and Luna๏ฝOpenAI
- Claude Opus 5.5๏ฝAnthropic
- Introducing System One Models and Jev๏ฝTypeSafe AI
- Key Questions on Energy and AI๏ฝIEA
- Morgan Stanley raises US data center power shortfall estimate๏ฝYahoo Finance
- China’s AI price war is entering a new phase๏ฝBusiness Insider (Yahoo Finance)
- America Must Expand, Not Weaken, AI Export Controls at Upcoming Talks with China๏ฝCNAS
- The State of AI Global Governance and Its Implications for the U.S.-China Summit๏ฝCSIS
- Bessent: US proposes AI incident alert system in talks with China๏ฝAP (ABC News)
- Demand Forecast for Japan and Each Supply Area, FY2026 (Japanese)๏ฝOCCTO


