~/MULTIMODAL A/ant-group-open-sources-ling-3-0-flash-vl-multimodal-model-with

Ant Group Open-Sources Ling-3.0-flash-VL Multimodal Model with Visual Feedback Loop

Ant Group has open-sourced Ling-3.0-flash-VL, a native multimodal Mixture-of-Experts (MoE) model featuring 124 billion total parameters with 5.5 billion active per token. The model natively supports text, image, and video inputs across a 256K token context window and introduces a visual feedback closed-loop mechanism for iterative task execution. By implementing a continuous 'observe-act-verify-correct' feedback loop, Ling-3.0-flash-VL enables more reliable execution in agentic tasks like GUI automation, code rendering, and medical image interpretation. Open-sourcing an efficient MoE architecture provides the AI developer community with a competitive open alternative to proprietary multimodal models. The model incorporates an arbitrary-resolution visual encoder, VideoRoPE, and a 42-layer language backbone combining KDA and Gated MLA in a 5:1 ratio. Ant Group released BF16 and FP8 weights on Hugging Face and ModelScope, with lower-bit FP4 and INT4 versions planned for release soon.

## BACKGROUND

Multimodal Large Language Models (MLLMs) usually process visual inputs in a single forward pass, which can lead to errors when performing multi-step visual tasks. Mixture-of-Experts (MoE) architectures help models maintain broad knowledge capacity by training a large pool of parameters while activating only a small subset during inference to keep computing costs low.

## REFERENCES

## KEYWORDS

#Multimodal AI#MoE#Open Source#Computer Vision#LLM

$ subscribe --daily

Ant Group Open-Sources Ling-3.0-flash-VL Multimodal Model with Visual Feedback Loop | Daily News