跳到正文
原文
The Decoder· Matthias Bastian·· 4 小时前AI 评分66

Vals AI 研究:AI 智能体团队成本高数倍,质量提升却难以测量

AI agent teams waste massive tokens for barely measurable quality gains, research finds

AI 导读

评测公司 Vals AI 在 Vibe Code Bench 上测试 GPT-6 Sol 和 Claude Opus 5.5,分别以单智能体和智能体团队形式运行,并设置中等与最高两档推理强度。

来源:The Decoder · the-decoder.com