Key Points
- DeepSeek unveiled an experimental multimodal version of its flagship V4 Flash model that can read images
- The company says its performance nears Anthropic's Opus 4.8 on multimodal, autonomous agent tasks
- The release intensifies a US-China race over cheaper AI models matching premium Western performance
The Latest:
DeepSeek unveiled an experimental multimodal version of its flagship V4 Flash model, called V4-Flash-Vision-Exp, that can interpret images and act on visual prompts alongside text. The Hangzhou-based company made the model available to developers through its API on Friday.
Details:
- Performance: DeepSeek said the new model’s performance on multimodal agentic tasks is close to Anthropic’s Opus 4.8.
- Text skills: The company said on X the experimental model matches V4 Flash’s text performance in agentic tasks, reasoning and world knowledge.
- Background: DeepSeek has released visual models before under its DeepSeek-VL family, but this marks the first time the capability reached its flagship line.
- Context: Chinese AI firms are racing US developers, often matching their performance at a fraction of the price, Bloomberg reported.
- Access: The experimental model is available now through DeepSeek’s API.
What to Watch:
DeepSeek hasn’t set a timeline for a full, non-experimental release, and its Opus 4.8 comparison rests on the company’s own claim rather than independent benchmarks.