Alibaba Launches Qwen-Image-3.0: Supports 4.5K Ultra-Long Prompt Input, Delivers Complex Text-and-Image Layouts in One Shot

Alibaba launched its next-generation image generation model, Qwen-Image-3.0, on July 21. The core breakthrough lies in its support for ultra-long text inputs of up to 4,500 tokens, significantly enhancing its understanding of complex prompts and the precision of text-and-image layouts. The model natively supports rendering in 12 languages and over 20 fonts, enabling one-click generation of multilingual posters, multi-layered nested UI interfaces, film storyboards, and long-form knowledge infographics with clear text—effectively solving the common AI image generation problems of garbled text and layout chaos. Evaluations indicate its performance ranks among the top in China. API trials are now open, with the potential to lower the production costs of commercial materials in advertising, creative design, and film. However, unlike its two predecessors, this release did not include benchmark scores, model weights, or a technical report, leaving its claimed performance improvements without a basis for external verification.
Alibaba Launches Qwen-Image-3.0: Supports 4.5K Ultra-Long Prompt Input, Delivers Complex Text-and-Image Layouts in One Shot

Image generation models are evolving from mere "image output tools" into the core foundation for professional content production. On July 21, Alibaba officially launched its new-generation foundational image generation model, Qwen-Image-3.0. As the third generation in the Qwen image series, the model's biggest breakthrough is expanding the text input length to a maximum of 4,500 tokens, achieving precise understanding of ultra-long, complex prompts and high-precision text-and-image layouts. This directly addresses the long-standing pain points of text errors and long-text control failures that have plagued AI image generation in professional application scenarios.

According to a first-hand evaluation by Zhixi, Qwen-Image-3.0 has achieved significant upgrades in text alignment, detail restoration, and complex composition rendering. Currently, in text-to-image machine evaluations, the model is considered the top-ranked model in China, just below GPT Image 2. Drive House reported that the new model natively supports the precise rendering of 12 languages, including Chinese, English, and Korean, and over 20 fonts, which can substantially reduce the production costs of commercial materials like multilingual posters and film storyboards.

From a generational evolution perspective, the 4,500-token figure is the most specific and easily verifiable number in this upgrade. Alibaba clearly stated in the Qwen-Image-2.0 technical report that the previous generation model supported a maximum prompt length of about 1,000 tokens. Based on this, the input window for version 3.0 has expanded by roughly 4.5 times compared to its predecessor.

VersionRelease DateMax Prompt LengthMultilingual Text RenderingParameters
Qwen-Image (1st Gen)Aug 4, 2025Not officially specifiedPrimarily Chinese and English20B
Qwen-Image-2.0May 11, 2026~1K tokensMultilingual layout, number of languages not listedNot disclosed
Qwen-Image-3.0Jul 21, 20264.5K tokens12 languages, 20+ fontsNot disclosed

Note: The date for Qwen-Image-2.0 is the submission date of its technical report (arXiv 2605.10730); the first-generation parameter count and licensing information are from its HuggingFace model repository.

Conquering Complex Layouts: From Concert Tickets to Multilingual Posters

In terms of text generation accuracy, Qwen-Image-3.0 demonstrates extremely strong control. In Zhixi's testing, the model easily generated materials containing massive amounts of text with complex layouts. For example, when generating an "Eason Chan World Tour Concert" ticket, the model not only accurately presented core information like the theme, date, and address but also maintained a consistent font style, with the text complementing background elements like the singer and microphone without the common issue of text distortion.

For long-form text generation, Qwen-Image-3.0 also performed exceptionally well. When given a prompt to generate the nearly 800-character Preface to the Prince of Teng's Pavilion, the model produced the full text without errors or omissions, cleverly differentiated the font and size for the title, author, and body, and even reproduced the author's red seal script stamp and the printed texture of a Chinese textbook.

The model also achieved a breakthrough in the traditionally difficult area of multilingual mixed layouts. In an art exhibition poster blending watercolor and cyberpunk styles, the model accurately handled a bilingual layout of Chinese and Korean, with the font colors perfectly complementing the two distinct visual styles on either side. Even when superimposing multilingual floating lyrics in Chinese, English, and Korean onto a live concert photo, the text remained clear and complete, without missing characters, garbled text, or misaligned deformation.

Mastering Multi-Layered Nested Logic: From UI Interfaces to Film Storyboards

Another core highlight of Qwen-Image-3.0 is its support for complex logical nesting and industrialized batch image generation. Drive House reported on an extreme test: the model needed to understand the complex prompt, "In a VS Code programming interface, generate a poster for pour-over coffee posted in a WeChat chat interface via the Qwen App." The final generated image clearly reflected the multi-layered UI nesting relationship from VS Code to the Qwen App to the social media chat interface, with complete layouts and text content at each level, and even tiny 10-pixel text appearing sharp and clear.

In the realm of professional content creation, the model also demonstrated its potential as a productivity tool. According to Zhixi's experience, Qwen-Image-3.0 can directly output a vertical short comic strip containing 20 storyboard panels in one go. The entire set of panels features a smooth narrative flow, clear dialogue bubble text without garbled characters, and consistent character expressions and scene atmosphere. Furthermore, in generating hand-drawn style advertising storyboards, the model can present six panels at once, complete with camera parameters, lighting tones, and camera movement speed, achieving a level of completion sufficient for initial creative proposals.

In terms of complex information graphics, Drive House, citing official information, stated that the model can generate a 9-grid knowledge infographic covering 9 different fields in one go, with every character crystal clear. In Zhixi's testing, when faced with a group portrait featuring over twenty characters spanning various styles like historical costume, xianxia fantasy, and cyberpunk, the model maintained a harmonious overall atmosphere while ensuring the text on core shop signs was accurate and smooth. Even when generating a structural diagram of chloroplast photosynthesis principles—content containing numerous chemical formulas and molecular structures—the model accurately output process annotations and equations.

It is worth adding that Alibaba's official release page summarizes this generation's capabilities into three main themes: ultra-long prompts, fine details, and "world knowledge." The "world knowledge" aspect received less attention in most evaluation reports—beyond natively rendering 12 languages, official examples also include replicating web pages, live streaming interfaces, and even generating weather forecast maps for specified cities and dates. This implies the model can call upon online data to participate in image creation, blurring the line between image models and data-driven design tools. Officials also noted that the aforementioned 9-grid knowledge infographic was generated in one go from a single prompt of approximately 3,700 tokens, rather than being stitched together from multiple images afterward.

A Shift in Open Strategy: No Weights, No Technical Report This Time

Beyond the touted capabilities, the release method for Qwen-Image-3.0 shows a clear departure from its two predecessors. Tech media outlet Unite.AI pointed out on the day of release that the official announcement lacked benchmark scores, parameter counts, licensing terms, and downloadable weights or a technical report. Users were directed straight to Qwen Chat for a trial.

This differs from the series' previous conventions. The first-generation Qwen-Image, released on August 4, 2025, opened the 20B parameter model weights under an Apache 2.0 license and published a technical report on the same day; Qwen-Image-2.0 also had a public technical report. As of press time, no official model repository for Qwen-Image-3.0 has appeared on HuggingFace.

For image models, the impact of this gap is more direct than for text models: image quality is inherently subjective, and text rendering is precisely the dimension where generative models shine in curated examples but are prone to faltering under systematic testing. Without weights and a public evaluation set, the evidence available to developers is essentially limited to the vendor's self-selected sample images. Unite.AI also noted that the practices of competing products in the same period show that "opening only a chat interface" is a choice, not a limitation—Tencent's HunyuanImage 3.0 was released with open weights.

A reference data point comes from Alibaba itself. The Qwen team's evaluation set, Qwen-Image-Bench, published this May, uses a Q-Judger evaluation model fine-tuned from Qwen3.6-27B, covering 1,000 prompts and 56 verifiable sub-metrics. In its published leaderboard, the previous flagship, Qwen Image 2.0 Pro, scored 57.84 overall, ranking fifth.

ModelQualityAestheticsAlignmentReal-World RestorationCreative GenerationTotal Score
GPT Image 258.6567.5365.8557.3875.2364.69
Nano Banana 2.054.7761.0862.4054.2867.0559.82
GPT Image 1.555.1460.8861.7253.9566.3559.65
Nano Banana Pro55.6760.2661.2554.0766.2359.45
Qwen Image 2.0 Pro54.3958.6759.2851.8364.9457.84

Note: This leaderboard is the top five of the Qwen team's self-built evaluation, scoring models from the Qwen-Image-2.0 generation and does not include the newly released 3.0. The earlier mention by Zhixi of "top-ranked in China" refers to the relative ranking among Chinese models; the two metrics have different scopes.

Industry Impact: Lowering the Trial-and-Error Cost of Professional Creation

Currently, AI image generation is at a critical juncture, transitioning from a personal creative tool to a pillar supporting professional content production. Early image models, limited by input text length and complex semantic understanding, struggled to deliver substantive value in professional fields like advertising, film, and science communication. The release of Qwen-Image-3.0 marks a significant improvement in text-image alignment precision and industrialized batch image generation capabilities.

Drive House's analysis suggests that with its extreme mastery of semantic juxtaposition and spatial control, users no longer need to compress complex requirements into a single sentence. Instead, they can fully describe the image structure like writing a requirements document, significantly reducing the cost of repeated generation and manual adjustments. Beyond generation capabilities, the model's accumulated text rendering and detail texture skills have also been transferred to editing scenarios, allowing complex editing tasks like ancient painting restoration and converting hand-drawn sketches to PPT to be completed in one click.

Alibaba Cloud Bailian and the Qwen AI Platform have reportedly opened API trials, with Qwen Studio and the Qwen App soon to be available for free user experience. As such tools further lower the barrier to realizing complex creative ideas, the potential of AI image generation as a foundational technology for professional content production across all industries is rapidly being amplified. As for whether this generation's actual capabilities can live up to the promotional claims, two signals to watch for are: whether Alibaba will supplement the weights and model card as it did with previous generations, and whether independent third-party evaluations can verify its performance on ultra-long prompts and small-text rendering.

Add to Google Preferred Sources

Once added, BigGo Finance appears first in Google Search Top Stories, so you get the broadest, most up-to-the-minute, and most comprehensive global financial news first.







More Related News