Technology · 32 views
DeepSeek's top-ranked V4 Flash stumbles on real agent tasks as its prices surge
DeepSeek's V4 Flash has topped model leaderboards and been hailed by developers as a "total monster" since its rollout. But in real-world testing, it completed just 53.8% of a batch of complex agent tasks.
AI Summary
DeepSeek’s V4 Flash, which has led model leaderboards and been praised by developers, was evaluated on real‑world agent tasks. In testing by Composio, the model was run through eight different agent harnesses on 30 complex, multi‑step workflows involving live tools such as Gmail, GitHub, Slack and Google Sheets. Across 240 total runs, only 129 succeeded, meaning the model completed 53.8% of the tasks, and just six of the workflows were completed successfully by every harness. This performance gap highlights a discrepancy between benchmark rankings and practical task execution.
AI summaries can be wrong sometimes—always verify important details using the source article.
How AI & Automation are usedCoverage Context
More from Technology
Continue reading recent Technology coverage
- Hackers obtain counterfeit TLS certificates for Google and other large servicesContinue reading
- Two mystery PlayStation products have leaked — one might be a PlayStation Portal OLEDContinue reading
- Oracle Health Hack Exposes Data of Nearly 20 Million PeopleContinue reading
- Everybody’s Favorite Art TV Is Nearly Half Off for Prime Day (2026)Continue reading
Support HappeningNow
Independent AI-powered news analysis is reader-supported. Your contribution helps cover infrastructure, summaries, and continued platform development.
Support HappeningNow