Story thread · 2 reports / 2 sources
DeepSeek's top-ranked V4 Flash stumbles on real agent tasks as its prices surge
venturebeat.com · 2h · first report

How the coverage leans
Across 2 sources · syndicated copies counted once
DeepSeek's V4 Flash has topped model leaderboards and been hailed by developers as a "total monster" since its rollout. But in real-world testing, it completed just 53.8% of a batch of complex agent tasks. Composio ran the model through eight different agent harnesses , including Claude Code, Codex, and OpenCode, on 30 deliberately difficult, multi-step tasks spanning live tools like Gmail, GitHub, Slack, and Google Sheets. Of 240 total runs, 129 passed — and only six of the 30 workflows were completed successfully by every harness tested. The gap illustrates why orchestration, not raw model capability, may decide whether the model succeeds in enterprise settings: the same model produced substantially different results depending on the harness, tool configuration, caching behavior, retries, and provider stack it ran on. DeepSeek said it will be hiking the prices for V4 Flash and Pro , models that have quickly become favorites among developers building coding assistants and agents. Whil
The coverage
- DeepSeek’s V4 Flash struggles with real-world tasks despite topping AI leaderboards
cryptobriefing.com · 2h
The conversation · 0
Sign in to join the conversation.
No comments yet — start the thread.