Microsoft Releases AesCode 8B and 32B for Visual Artifact Generation
Microsoft has released AesCode 8B and 32B, two multimodal code models designed to generate structured, editable HTML/CSS visual artifacts such as slides and dashboards. The models use visual reference images combined with decoupled cross-modal supervision to ensure accurate layouts and content. This release bridges the gap between pure code models that struggle with visual aesthetics and image generators that often misrender text and numbers. By outputting editable HTML/CSS, it enables users to easily modify generated dashboards and slides for practical applications. The larger AesCode-32B checkpoint builds upon Qwen3-VL-32B-Instruct, undergoing cold-start SFT followed by GDPO across seven distinct reward channels. Quantized GGUF versions by bartowski and raw checkpoints are both available directly on Hugging Face.
## BACKGROUND
Large language models often struggle to balance the precise structural requirements of code with the artistic layout demands of visual design. Multimodal models address this by processing both text and visual inputs, though coordinating semantic text requirements with visual aesthetic references remains a significant technical hurdle.