Tencent Open-Sources Youtu-Parsing-Omni Multimodal 5B Model
Tencent has open-sourced Youtu-Parsing-Omni, a 5-billion parameter multimodal model that parses documents, images, charts, audio, and video into unified structured JSON formats. Guided by task prompts, the single omni-encoder extracts layout, OCR, ASR, visual captions, bounding boxes, and camera motion across seven task families. By consolidating disparate perception tasks like OCR, speech recognition, layout analysis, and visual description into a compact 5B model, Youtu-Parsing-Omni greatly simplifies multimodal data ingestion pipelines. Achieving the highest score among open-weight models on OmniParsingBench (75.08 Avg.), it offers an efficient, locally deployable alternative to proprietary models like Gemini-3-Pro. The model achieved state-of-the-art results on OmniDocBench v1.6 with an overall score of 96.96 and shows strong performance on domain-specific benchmarks like ChemOCR and music score parsing. It comes ready for deployment with an included vLLM plugin, serving configurations, and prompt templates for supported modalities.
## BACKGROUND
Multimodal document parsing involves converting unstructured content—such as scanned PDF pages, charts, audio recordings, or videos—into structured formats like Markdown, HTML, or specialized formats such as OTSL (One-dimensional Table Structure Language). Historically, developers had to chain together separate optical character recognition (OCR), automatic speech recognition (ASR), and visual object detection models, making unified omni-modal architectures much more efficient for automated data pipelines.