SenseTime Open-Sources SenseNova U1.5-Lite-Preview with Native 4K Image Generation
SenseTime has open-sourced SenseNova U1.5-Lite-Preview, a lightweight 8B-parameter multimodal model that natively supports 4K image generation. This preview version utilizes a Mixture of Transformers (MoT) architecture and outperforms its predecessor, U1, across multiple generation and editing benchmarks. By achieving image generation and editing capabilities comparable to commercial closed-source models at a lightweight scale, this release lowers the barrier for high-resolution image synthesis. It provides developers and researchers with a highly capable, open-source alternative for advanced multimodal tasks. The model features improved fine-grained textures, more accurate English and Chinese text rendering within images, and better visual instruction following. It is currently hosted on GitHub, Hugging Face, and ModelScope for community access.
## BACKGROUND
SenseNova is SenseTime's family of large-scale foundation models designed for multimodal integration and high-performance AI tasks. The Mixture of Transformers (MoT) architecture is a sparse machine learning design that routes multimodal inputs through modality-specific parameters, optimizing computational efficiency for large models.