AIResearchAIResearch
Machine Learning

DeepSeek Debuts V4 Flash Vision Exp Multimodal Model

DeepSeek releases V4 Flash Vision Exp, a multimodal AI model trained on 32T tokens, excelling in image analysis and surpassing Anthropic's Opus 4.8.

2 min read
DeepSeek Debuts V4 Flash Vision Exp Multimodal Model

TL;DR

DeepSeek releases V4 Flash Vision Exp, a multimodal AI model trained on 32T tokens, excelling in image analysis and surpassing Anthropic's Opus 4.8.

DeepSeek has launched V4 Flash Vision Exp, a multimodal extension of its V4 Flash series, trained on 32 trillion tokens. The model currently is available only through DeepSeek's paid developer platform, though the company has open-sourced earlier models and may follow suit. V4 Flash Vision Exp builds on V4 Flash, a mixture-of-experts model with 284 billion parameters, where only relevant subnetworks activate per prompt to reduce compute demands.

Across seven text-based benchmarks, V4 Flash Vision Exp improved over its predecessor in all but one: Cybergym, which tests software vulnerability discovery. The largest gains came in image analysis, where the model scored over 10% higher on two of four visual benchmarks. It also surpassed Anthropic's Opus 4.8 on ALE and ZeroBench, evaluations covering multi-step tasks and challenging image reasoning.

DeepSeek has not disclosed V4 Flash Vision Exp's architecture, but its Hugging Face page details V4 Flash, the base model. That model uses a mixture-of-experts design with 284 billion parameters split into 13-billion-parameter subnetworks. A KV cache stores context for prompt responses, and the sparse activation approach cuts hardware requirements compared to dense models.

The release comes amid broader shifts in AI leadership. Google DeepMind's Demis Hassabis stepped down as CEO to become chairman, with CTO Koray Kavukcuoglu taking over operations. The move follows delays in Gemini 3.5 Pro and departures including Jeff Dean, raising questions about Google's AI direction.

Meanwhile, DeepMind outlined 15 years of game AI research, from AlphaGo to SIMA 2, now partnering with EVE Online developer Fenris Creations. The lab is shifting toward generalist agents that operate across environments without game code access.

For practitioners, V4 Flash Vision Exp signals that Chinese labs are matching or exceeding Western multimodal benchmarks with efficient architectures. Its mixture-of-experts approach offers a path to strong performance without massive compute overhead, relevant for teams constrained by hardware budgets.

The model's success on visual reasoning tasks suggests multimodal systems are approaching practical utility in domains requiring both text and image understanding. However, its narrow miss on Cybergym highlights that security-related reasoning remains a challenge even for advanced models.

As AI development accelerates globally, models like V4 Flash Vision Exp demonstrate that innovation is no longer centralized in a few Western labs. The competition is driving faster iteration and more diverse architectural choices, benefiting the broader research community.

FAQ

What is V4 Flash Vision Exp?
It's DeepSeek's new multimodal model trained on 32T tokens, extending the V4 Flash series with image analysis capabilities.

How does it compare to Opus 4.8?
V4 Flash Vision Exp outperformed Anthropic's Opus 4.8 on two visual benchmarks: ALE and ZeroBench.

Is it open source?
V4 Flash Vision Exp is currently only available via DeepSeek's paid developer platform, though earlier models were open-sourced.

What architecture does it use?
It builds on V4 Flash, a mixture-of-experts model with 284 billion parameters, using sparse activation to reduce compute needs.

About the Author

Guilherme A.

Guilherme A.

Former dentist (MD) from Brazil, 41 years old, husband, and AI enthusiast. In 2020, he transitioned from a decade-long career in dentistry to pursue his passion for technology, entrepreneurship, and helping others grow.

Connect on LinkedIn