Assertion 1
On DeepSWE, a software-engineering test, Claude Fable 5 solves 69.7% of tasks on the first try against 62.8% for DeepSeek's new V4 Pro 0813 Together AI Blog.
Assertion status: No spot-check verdict is published for this assertion.
Running DeepSeek V4 Pro 0813 first and escalating to Claude Fable 5 only when DeepSeek fails solves 82.7% of DeepSWE tasks at a cost of $8.28 per task.
No stored spot-check names this claim in this edition.
Claude Fable 5 alone solves 69.7% of DeepSWE tasks at a cost of $21.63 per task.
No stored spot-check names this claim in this edition.
Claude Fable 5 achieves 69.7% pass@1 compared to DeepSeek V4 Pro 0813's 62.8% pass@1 on DeepSWE, a 7-point lead for Fable on the first try.
No stored spot-check names this claim in this edition.