GPT-5.5 Shock | The Full Picture: 1M Tokens × Coding King
機械翻訳 / Machine-translated

機械翻訳 / Machine-translated
@aifriends
AI Friends(https://aifriends.jp)のクロスポスト公式アカウント。AIツールの紹介・使い方・できることを、中学生でもわかるやさしい日本語で届けます。
"AI models have leveled up once again" — on April 23, 2026, OpenAI unveiled GPT-5.5 (internal codename: Spud, as in the humble potato), the first fully retrained model since GPT-4.5 roughly a year ago.
A super-large 1M-token context window (equivalent to five novels), a world-record 82.7% on Terminal-Bench 2.0, and API pricing double that of the previous generation — the aggressive pricing and flashy benchmark results are generating enormous buzz.
Let's break down in plain language — accessible even to a middle schooler — why a full retraining was needed now, how it differs from Claude Opus 4.7 and Gemini 3.1 Pro, and what changes for users in Japan.
Let's organize the key points of the announcement in three minutes.
On April 23, 2026, OpenAI officially announced GPT-5.5 on their blog "Introducing GPT-5.5," rolling it out the same day to ChatGPT Plus, Pro, Business, Enterprise, and Edu plans. The following day, April 24, it became available via API, allowing developers to access it through the "Responses" and "Chat Completions" endpoints.
Think of it like a new smartphone unveiled in stores one day, with online sales opening the very next day.
The speed of aligning consumer and developer access within 24 hours is a hallmark of this release.
The industry was stunned by OpenAI's pace — a successor model arriving just one to two months after GPT-5.4.
The internal development codename is "Spud" — casual English slang for potato.
It carries the nuance of "patiently grown from the ground up, with a warm, hearty interior."
It reflects a development philosophy of "before the fancy dish, rebuild the base ingredient from scratch."
Whereas GPT-5.1 through 5.4 were variations built on top of the same base model with successive fine-tuning, GPT-5.5 is different in that both the pretraining data and architecture were rebuilt entirely from scratch.
Spud (GPT-5.5) is the fruit of a base model reconstruction project — the first in roughly a year since GPT-4.5, undertaken with careful deliberation.
Alongside the standard model "GPT-5.5," a higher-tier model for deeper reasoning — "GPT-5.5 Pro" — was announced on the same day.
It's like a sibling lineup where both a regular supercar and a circuit-spec racing model launch at the same time.
GPT-5.5 Pro is optimized for long-duration agentic tasks and research requiring deep reasoning.
In ChatGPT, it appears as "GPT-5.5 Thinking," available to Plus, Pro, Business, and Enterprise users.
GPT-5.5 Pro is limited to the higher-tier Pro, Business, and Enterprise plans, creating a clear separation — and expanding the options users can choose based on their needs and budget.
Let's look at "why rebuild from scratch instead of fine-tuning an existing model?" from three angles.
Pretraining is the process of feeding massive amounts of internet-scale text data to an AI, building its foundational language abilities.
Just like making pudding at home — mixing eggs, sugar, and milk from the start before heating — the quality of the foundation determines the final outcome.
The standard approach is to layer fine-tuning on top of a base model that's already been built.
GPT-5.1 through 5.4 were like different sauces poured on the same pudding, while GPT-5.5 is the pudding itself remade with a new recipe.
Pretraining requires thousands of GPUs, months of computation, and costs on the order of tens of billions of yen — making it one of the most demanding undertakings in AI development.
According to OpenAI, GPT-5.5's pretraining corpus, architecture, and agent-oriented training objectives have all been completely overhauled.
It was designed from the ground up with the assumption that "the model will use tools on its own, make plans, and see long-duration tasks through to completion."
Think of it as hiring someone you can simply hand a goal to and trust to figure out the rest — rather than a new employee who needs constant supervision from a manager.
The result of strengthening the ability to maintain long context, recover from ambiguous failures, and use tools for hypothesis testing is a model that can handle implementation, refactoring, debugging, testing, and verification all in one continuous flow.
It is positioned as the core model of OpenAI's agentic infrastructure — a key focus area in 2026.
"Spud" is a casual word for potato in English-speaking cultures — a tongue-in-cheek codename from the OpenAI team.
The image of a modest but nutrient-rich, versatile ingredient that works in countless dishes overlaps elegantly with the idea of a general-purpose base model.
The name aims at being "the unsung hero of the table — unassuming but indispensable."
OpenAI has historically used food-related codenames (rumors suggest GPT-6 carries a different food codename, not Spud).
Playful internal nicknames that reflect developer culture have been generating more industry buzz than polished marketing names.
Let's dig into "what specifically got stronger?" across three key points.
GPT-5.5 is the first to support a 1M-token (one million token) context window in the OpenAI API. In Japanese, that translates to roughly 600,000–700,000 characters — enough to read more than five paperback novels in a single pass.
Think of it as "absorbing a thick dictionary entirely into memory and answering questions instantly."
It can take in an entire massive codebase (e.g., a company's full internal source code) or a long history of chat conversations and respond with full awareness of the whole picture.
Codex (ChatGPT's coding feature) has also been expanded to support 400K tokens, raising practical utility for engineers another level.
Terminal-Bench 2.0 is an industry-standard benchmark measuring an AI's ability to handle complex command-line workflows — planning, iteration, and tool coordination. GPT-5.5 set a world record with an accuracy of 82.7%.
It's the level of autonomy of a junior engineer who can plan a difficult build task on their own, fix errors along the way, and see it through to the end.
It also scored 88.7% on SWE-Bench Verified (a real-world bug fix test on GitHub), surpassing the previous generation GPT-5.4's 87.6%.
Efficiency is another hallmark: Codex uses approximately 40% fewer tokens than GPT-5.4 to complete the same tasks.
Getting smarter while generating less unnecessary output is the biggest practical advancement.
OpenAI officially stated: "Agents built on GPT-5.5 can handle planning, context gathering, tool calls, recovering from ambiguity, and completing long-horizon workflows with fewer instructions."
It's like having a subordinate who, instead of reading a thick instruction manual, can think for themselves and act when given just the goal.
GDPval (economic value task evaluation): 84.9%; OSWorld (PC operation agent): 78.7% — dominating the agent category.
OpenAI emphasizes particular strength in knowledge work, early-stage scientific research, and computer operation.
This announcement marks a turning point in ChatGPT's evolution — from a tool that answers questions to a partner that sees tasks through to completion.
Let's organize "how does it stack up against other companies' models?" from three angles.
On the Artificial Analysis Intelligence Index — a comprehensive intelligence score from a third-party benchmarking organization — GPT-5.5 earned 60 points to take first place. Claude Opus 4.7 and Gemini 3.1 Pro tied at 57 points in second, putting OpenAI on top by 3 points.
It's like being the top student in the year by just 3 combined points over two close rivals — a narrow but decisive margin.
OpenAI, having previously ceded the top position, reclaimed the overall lead with GPT-5.5.
However, the era of winning on a single metric is over — the essence of the modern AI market is shifting toward using different models for different purposes.
On SWE-Bench Pro (real-world GitHub issue resolution), Claude Opus 4.7 scored 64.3% versus GPT-5.5's 58.6% — a 5.7-point lead for Claude.
On harder code modification tasks, Anthropic's latest model holds a slight edge.
It's like being faster than a colleague in day-to-day work, but losing to them on a specialized licensing exam.
On the other hand, Google's Gemini 3.1 Pro leads on abstract reasoning benchmarks like ARC-AGI-2 (77.1%) and GPQA (94.3%).
A three-model split — GPT-5.5 for agents, Claude for coding, Gemini for reasoning — is becoming the industry standard in 2026.
API pricing is $5 per million input tokens and $30 per million output tokens — exactly double GPT-5.4's $2.50 input / $15 output.
Whether you see this as justified comes down to how you weigh the performance and efficiency gains against the price increase.
GPT-5.5 Pro goes even further: $30 input / $180 output — 6× the price.
However, with a 40% reduction in output tokens, analysts suggest the effective cost increase is closer to 20%.
Whether the cost of the latest frontier model is justifiable depends on the added value of the task, according to industry analysts.
Let's look at "what changes for users in Japan?" from three angles.
In Japan, ChatGPT Plus, Pro, Business, and Enterprise subscribers gained automatic access to GPT-5.5 from April 23.
No additional fees, no settings changes — "GPT-5.5" and "GPT-5.5 Thinking" simply appeared in the model selection menu.
It's like a subscription streaming service adding new movies to your plan without any change to your subscription.
Early reviews indicate improvements in Japanese-language translation, summarization, and text generation accuracy over GPT-5.4.
Hundreds of thousands of Japanese Plus users paying roughly ¥3,000 per month now have immediate access to the world's most advanced model.
For Japanese companies providing API-based SaaS products or AI assistants, switching to GPT-5.5 requires revisiting cost structures.
With output tokens down 40%, the total cost (unit price × volume) may not increase as much as expected — but the maximum cost per request does go up.
It's similar to electricity rates rising while a new energy-efficient air conditioner offsets some of the increase.
For startups and indie developers, the practical solution is a "model routing" architecture — using GPT-5.4, 5.5, and 5.5 Pro selectively based on the task.
Industry reports indicate that major domestic players like NTT, Fujitsu, and CyberAgent have already begun evaluating GPT-5.5 for business use.
Even as domestic LLMs like NEC's "cotomi," CyberAgent's "CALM2," and PFN's "PLaMo" continue to evolve, the arrival of GPT-5.5 risks reopening the gap in frontier performance.
It's the familiar catch-up dynamic: every time the domestic team improves, the overseas players move the goalposts further.
That said, domestic LLMs retain distinct strengths in Japanese-language specialization and data sovereignty.
A hybrid architecture — leveraging GPT-5.5's 1M context and frontier performance while processing sensitive information with domestic models — is becoming the mainstream enterprise approach in Japan.
Discussions about potential integration with the government AI "Gennai" are also entering a new phase in the second half of 2026.
Kato-san, CTO of a SaaS company in Tokyo, is considering whether to switch the AI assistant infrastructure in his company's product from GPT-5.4 to GPT-5.5.
He submitted an internal analysis showing: "With output tokens down 40%, we project maintaining current monthly API costs while improving performance across the board."
The 1M context window means the system can reference long customer inquiry histories in a single pass, with expectations of improved support accuracy.
"A bold move: keeping contract pricing the same while raising service quality."
In the AI market, companies that can make these migration decisions swiftly with each new model release gain a competitive edge — that's the reality today.
Miki-san, an economics major at a private university in the Kansai region, is using ChatGPT Plus to research her graduation thesis topic.
She tried "feeding GPT-5.5 five English papers of over 100 pages each simultaneously and having it organize the similarities and differences."
Thanks to the 1M context window, she could load all the papers at once without splitting them up, and noticed a dramatic reduction in misread context.
It's like having a tutor who has memorized every thick reference book and is available to answer your questions anytime.
It's a symbolic example of academic use of generative AI entering a new phase — freed from the constraints of long-context limits.
Suwa-san, a backend architect at a financial systems integrator, is piloting Codex's GPT-5.5 support in internal development.
After "feeding the entire 500,000-line internal Java codebase at once and having it auto-generate refactoring proposals and unit tests," completion time was reduced by about 30% compared to GPT-5.4.
The 82.7% Terminal-Bench 2.0 figure translates into a tangible real-world difference.
"A workplace where you can consult AI on design decisions the same way you'd ask a senior engineer for a review."
Suwa-san believes that by the second half of 2026, AI pair programming running autonomously in Codex will be standard practice in Japanese IT environments.
A. As of April 2026, GPT-5.5 is not available on ChatGPT's free plan. It is unlocked on paid plans: Plus ($20/month), Pro, Business, Enterprise, and Edu.
GPT-5.5 Pro is further restricted to the higher-tier Pro, Business, and Enterprise plans only.
Think of it as a new film premiering exclusively for paid members before a wider release.
Note: Via Codex, GPT-5.5 is also accessible on the Plus, Pro, Business, Enterprise, Edu, and Go plans (with context limited to 400K tokens).
The tier structure means your access point depends on your budget.
A. Standard GPT-5.5 is ideal for everyday Q&A, content generation, and coding. GPT-5.5 Pro is suited for complex reasoning, extended research, and large-scale design decisions.
With a 6× API cost difference, using Pro for simple tasks is as wasteful as driving a large truck for city errands that a compact car would handle fine.
"The craftsman's instinct to use the right tool for the job" is what matters.
In ChatGPT, it appears as "GPT-5.5 Thinking" and can be manually switched.
The recommended approach is to start with standard GPT-5.5 and only switch to Pro when the task genuinely demands it.
A. OpenAI has not published Japanese-specific benchmark figures, but early user reviews frequently mention improvements in long-form summarization, mistranslation of technical terms, and appropriate use of honorific language.
The 1M context window is a major practical advance — long Japanese texts can now be processed without splitting.
"Another step closer to a model that outputs Japanese as naturally as English."
That said, industry- or company-specific terminology still has limitations without fine-tuning — that hasn't changed.
For accuracy-critical work, designs that incorporate human review remain essential.
A. OpenAI announced "improved knowledge consistency during reasoning," but third-party evaluations note that hallucinations still occur frequently.
It's a universal challenge for modern LLMs — scoring high on benchmarks while still producing subtle real-world factual errors.
Think of a teacher who aces exams but still gets someone's name wrong in casual conversation — a structural dilemma.
For important decisions or official documents, human verification of AI output remains mandatory.
Best practices in 2026 include comparing multiple models and using citation verification tools alongside AI output for reliability-critical tasks.
A. OpenAI has made no official announcement, but as of April 2026, industry rumors suggest pretraining of "GPT-6" (said to carry a different food-related codename) may already be complete.
Industry analysts predict a formal release roughly six months to a year after GPT-5.5's announcement.
It's similar to rumors about the next-generation smartphone circulating immediately after the latest model launches.
That said, incremental improvements to GPT-5.5 itself (5.5.1, 5.5.2, etc.) are also expected to continue in parallel.
The practical approach is for users to align their decision on when to chase the latest model with their own business cycles.
The decision to "rebuild from the base model" is a weighty one — demanding maximum cost, time, and risk in AI development. The reason OpenAI made this choice for the first time in a year since GPT-4.5 was the need to rebuild a new architecture from the ground up to meet the demands of the agentic era.
A 1M-token context, a world record on Terminal-Bench, a 40% reduction in output tokens — none of these are one-off feature additions. They are achievements made possible by rethinking the design philosophy from the foundation up.
The aggressive 2× API pricing is also an experiment asking how much society will pay for intelligence.
"Can GPT-5.5 users generate twice the value to justify the cost?" — that is the central question of AI adoption in the second half of 2026. For Japanese users, companies, and developers alike, the era where the ability to strategically design when, where, and how much to use frontier models determines competitive outcomes in AI is now here. Seen that way, the weight of this announcement comes into full view.
This article is a cross-post from AI Friends.