ChatGPT Images 2.0 | Why Garbled Text Is Now a Thing of the Past
機械翻訳 / Machine-translated

機械翻訳 / Machine-translated
@aifriends
AI Friends(https://aifriends.jp)のクロスポスト公式アカウント。AIツールの紹介・使い方・できることを、中学生でもわかるやさしい日本語で届けます。
"I tried adding Japanese text to an AI-generated image and ended up with a mess of mysterious symbols." — Sound familiar?
On April 22, 2026, OpenAI announced ChatGPT Images 2.0 (gpt-image-2) — a model designed to solve that problem at its root. This is not a simple update; it represents a fundamental reimagining of the architecture (design philosophy) behind image-generation AI.
OpenAI officially released its new image-generation model, "ChatGPT Images 2.0," on April 22, 2026.
ChatGPT users access it as "ChatGPT Images 2.0," while developers can use it via API under the name "gpt-image-2."
The rollout to ChatGPT Plus, Team, and Enterprise users began on the same day as the announcement, with API access also opening immediately.
The previous-generation model was "GPT Image 1.5," based on DALL-E 3.
gpt-image-2 is an entirely different model — rebuilt from scratch with a new architecture (design philosophy).
The biggest change is the integration of "Thinking capabilities." The ability to "think carefully before answering," a hallmark of OpenAI's O-series models (o3, o4-mini, etc.), has been incorporated into the pre-generation stage of image creation.
So why were traditional AIs so poor at handling Japanese text?
Most previous image-generation AIs used a mechanism called a "diffusion model." This approach reconstructs an image gradually from noise (static).
Diffusion models have a fundamental weakness: text occupies only a tiny fraction of the pixels in an image. In models trained predominantly on English and Latin-script data, accurately reproducing the complex letterforms of Japanese was simply out of reach.
A single kanji character contains multiple strokes and subtle variations in shape. Diffusion models would interpolate (fill in) those details as "roughly this shape," producing characters that couldn't be read.
Non-Latin scripts in particular — Japanese, Chinese, Korean, Arabic, and others — had far less training data than English and far more complex letterforms, making garbled output a frequent occurrence.
The solution gpt-image-2 adopts is adding a "blueprint creation" step before any drawing begins.
When the model receives a prompt, it first reasons carefully — deciding where to place text, how large to make characters, which direction they should run — before it starts generating the image.
To use a designer analogy: previously it was "jumping straight into the final product without even a rough sketch." Now it's more like "building a wireframe before moving into design."
OpenAI's O-series models are embedded in this reasoning step, pre-planning the layout, composition, and spatial arrangement of text.
OpenAI announced in its official blog that gpt-image-2 achieves over 95% accuracy in rendering text in non-Latin scripts, including Japanese, Chinese, Korean, Arabic, and Bengali.
In practical terms, the following types of output are now viable:
Additionally, because web search can be executed during generation, accurate renderings of things like the latest product names and correct spellings of real place names can be embedded directly into images.
Beyond eliminating garbled text, gpt-image-2 brings a host of other capabilities.
① Up to 2K resolution: The previous 1024px ceiling rises to 2048px — sufficient quality for print and A3 posters.
② Flexible aspect ratios: From 3:1 (ultra-wide banner) to 1:3 (tall smartphone portrait), with step-by-step control. Optimized for the exact dimensions required by each social media platform.
③ Consistent batch generation of up to 8 images from a single prompt: Multiple variations are produced in one go while maintaining the same character or visual world. Ideal for laying out manga pages or creating A/B test variants.
④ Web search integration: The internet can be searched during generation to pull in up-to-date information — useful for visuals reflecting current events or images referencing real logos and product appearances.
⑤ Self-check function: After generation, the output is self-evaluated; if quality falls short, the model attempts to regenerate. This dramatically reduces the need to repeatedly hit "generate again."
As of April 2026, several major image-generation AIs are available. Here's a breakdown by use case:
For use cases involving Japanese text in images, multiple reviews agree that gpt-image-2 offers the highest accuracy as of April 2026.
That said, Midjourney v8 remains the stronger choice for photorealistic portrait-style generation or distinctive artistic aesthetics, and Adobe Firefly is the go-to when licensing risk must be completely eliminated for commercial use.
Speaking with designers at advertising agencies, it's increasingly common to hear of workflows like "gpt-image-2 for Japanese banners, Midjourney for art-style visuals targeting overseas audiences" — using each tool where it shines.
Until now, the standard workflow for creating banners with Japanese text was: generate a base image with AI, then add text in Photoshop afterward.
With gpt-image-2, finished output including Japanese text can be generated directly, dramatically reducing post-processing effort.
Consider a sales banner for an EC site. Previously, the process required at least four steps: ① AI generation → ② text added in Photoshop → ③ size adjustment → ④ review. With gpt-image-2, this compresses to three: ① single prompt → ② output with specified dimensions → ③ review.
The combination of generating up to 8 consistent images from a single prompt and rendering Japanese dialogue opens up revolutionary possibilities for manga thumbnail (rough layout) creation.
Among manga artists and webtoon creators, a hybrid production style is already spreading: "use gpt-image-2 to prototype panel layouts and dialogue → draw line art by hand."
Publishers and game companies are also increasingly adopting it as an effective way to cut costs at the concept art stage.
This is especially significant news for small businesses that cannot afford dedicated designers.
The ability to produce Japanese-language promotional materials with AI means higher update frequency without sacrificing budget to outside design work.
Possibilities include social media menu images for restaurants, flyers for cram schools and hobby classes, and product banners for online shops — all previously outsourced, now potentially brought in-house.
A. As of April 22, 2026, access is limited to ChatGPT Plus, Team, and Enterprise users, as well as API users.
No official announcement has been made regarding a rollout to the free plan, but there is a possibility of gradual expansion, as was the case with the previous-generation model (GPT Image 1.5).
A. The gpt-image-2 API is priced at approximately $0.04–$0.08 per image at standard resolution (roughly ¥6–¥12).
Selecting 2K resolution or high-quality mode increases the cost. Generating 10,000 images per month would run approximately $400–$800 (around ¥60,000–¥120,000). Compared to Midjourney's monthly subscription plans, the API can be more cost-efficient for high-volume generation.
A. Under OpenAI's terms of service, commercial use of images generated with gpt-image-2 is permitted.
However, use involving likenesses of real individuals, reproduction of copyrighted material, or misleading content is prohibited. Unlike Adobe Firefly, OpenAI does not explicitly disclose commercial licensing of training data, so companies with a low risk tolerance may opt for Firefly instead.
A. Basic vertical text is supported, but complex layouts should be verified.
Explicitly stating "in vertical text" (縦書きで) in the prompt improves accuracy. For punctuation placement and line-break rules (禁則処理), it is recommended to generate multiple variations and compare. For business use, always perform a visual check after generation.
A. The architecture (design) is fundamentally different. The biggest distinction is the presence or absence of a reasoning step.
GPT Image 1.5 was a diffusion model based on DALL-E 3. gpt-image-2 is an autoregressive model that includes a thinking step before drawing. Text accuracy, compositional precision, and the ability to follow complex instructions have all improved substantially.
The first people who should try this are every creator and marketer who regularly produces images with Japanese text. If you're a ChatGPT Plus user, you can start right now. Begin with a simple prompt like "Create a banner that includes Japanese text."
This article is a cross-post from AI Friends.