~/DEEPSEEK/deepseek-releases-experimental-multimodal-model-deepseek-v4-flash-vision-exp

DeepSeek Releases Experimental Multimodal Model DeepSeek-V4-Flash-Vision-Exp

DeepSeek has launched `deepseek-v4-flash-vision-exp`, an experimental vision-language model, on its API platform. The model supports text and image inputs in formats like JPEG, PNG, GIF, and WebP for tasks such as image description, text recognition, and chart analysis. This release introduces multimodal capabilities to DeepSeek's next-generation V4 architecture, significantly boosting its performance on agentic benchmarks to near Opus-4.8 levels. It allows developers to build more capable visual AI agents without sacrificing the speed and low cost of the Flash model. The model charges for images based on tokens (up to 384 tokens per image) at the same rate as the text-only V4-Flash. Additionally, DeepSeek introduced a free Files API that lets users upload images once and reference them via `file_id` to save request bandwidth.

## BACKGROUND

Multimodal large language models (LLMs) integrate visual processing with text generation, enabling them to understand screenshots, documents, and real-world scenes. Agent benchmarks assess a model's ability to execute complex, multi-step workflows and use external tools to solve problems.

## REFERENCES

## KEYWORDS

#DeepSeek#Multimodal LLMs#Artificial Intelligence#Open-Source AI

$ subscribe --daily

DeepSeek Releases Experimental Multimodal Model DeepSeek-V4-Flash-Vision-Exp | Daily News