Alibaba Releases Qwen-Image-3.0 with Advanced Text Rendering and 4.5k Token Support
Alibaba has officially released Qwen-Image-3.0, a new image generation foundation model that supports ultra-long inputs of up to 4.5k tokens. The model enables native rendering of 12 languages and over 20 fonts, allowing for the creation of complex layouts like academic papers and multi-layered user interfaces. This release shifts the focus of AI image generation from purely aesthetic outputs to highly practical, production-ready business assets. By solving the long-standing challenge of precise text rendering and complex spatial layouts, it significantly reduces production costs for multilingual posters, educational materials, and storyboards. Qwen-Image-3.0 can render tiny text down to 10px clearly and supports complex LaTeX mathematical formulas, nested UI designs, and 9-grid infographics in a single generation. The model's API is currently available for testing on Alibaba Cloud's Bailian and Qwen AI platforms.
## BACKGROUND
Historically, text-to-image models like Midjourney and Stable Diffusion have struggled with text rendering, often producing misspelled or unreadable text. Qwen-Image is a series of image generation and editing models developed by Alibaba's Qwen team, with previous versions focusing on accuracy and aesthetics, while this third generation prioritizes practical utility.