Story thread · 2 reports / 2 sources

DeepSeek's top-ranked V4 Flash stumbles on real agent tasks as its prices surge

venturebeat.com · 2h · first report

DeepSeek's top-ranked V4 Flash stumbles on real agent tasks as its prices surge

How the coverage leans

Across 2 sources · syndicated copies counted once

DeepSeek's V4 Flash has topped model leaderboards and been hailed by developers as a "total monster" since its rollout. But in real-world testing, it completed just 53.8% of a batch of complex agent tasks. Composio ran the model through eight different agent harnesses , including Claude Code, Codex, and OpenCode, on 30 deliberately difficult, multi-step tasks spanning live tools like Gmail, GitHub, Slack, and Google Sheets. Of 240 total runs, 129 passed — and only six of the 30 workflows were completed successfully by every harness tested. The gap illustrates why orchestration, not raw model capability, may decide whether the model succeeds in enterprise settings: the same model produced substantially different results depending on the harness, tool configuration, caching behavior, retries, and provider stack it ran on. DeepSeek said it will be hiking the prices for V4 Flash and Pro , models that have quickly become favorites among developers building coding assistants and agents. Whil

The coverage

  1. DeepSeek’s V4 Flash struggles with real-world tasks despite topping AI leaderboards

    cryptobriefing.com · 2h

The conversation · 0

Sign in to join the conversation.

No comments yet — start the thread.