DeepSeek Unifies Interaction Modes and Rolls Out V4.1 Flash Model
DeepSeek has merged its fast, expert, and vision interaction modes into a single unified intelligent mode that automatically detects prompt complexity and activates visual capabilities when needed. Concurrently, DeepSeek is releasing the V4.1 Flash model to replace V4 Pro, delivering faster speeds and superior performance across all benchmarks. This update streamlines the user experience by eliminating the need for manual mode switching while making high-performance multimodal AI more accessible and cost-effective. By deprecating V4 Pro in favor of the faster, cheaper, and natively multimodal V4.1 Flash, DeepSeek continues to push the boundaries of efficient AI inference. DeepSeek V4.1 Flash introduces a new architecture with native multimodal capabilities, and all V4 Pro traffic will automatically route to V4.1 Flash starting September 14, 2026. DeepSeek also updated Flash API pricing, setting off-peak rates at 0.02 RMB for cache hits, 1 RMB for cache misses, and 4 RMB for output per million tokens, with peak-hour pricing doubled.
## BACKGROUND
Historically, AI platforms required users to manually switch between different specialized model tiers or interaction modes depending on whether they needed quick answers, deep reasoning, or image inputs. Developers integrate these capabilities into applications using standard API design patterns such as Chat Completions or Messages endpoints to pass prompt context and manage multimodal inputs.