This frozen page shows Augur's claims and source links for one sent dispatch. Stored spot-checks appear only where the frozen edition supports them; absence is not presented as verification.
Cohere promotes its Transcribe model as beating Whisper Large v3, ElevenLabs Scribe v2, and Qwen3-ASR on the Hugging Face Open ASR Leaderboard, a public ranking of speech-to-text accuracy, and Mistral prices its Voxtral Mini Transcribe V2 at one fifth of Scribe v2 while claiming matched quality Cohere Blog Mistral AI.
Assertion status: No spot-check verdict is published for this assertion.
Voxtral Mini Transcribe V2 costs one-fifth of ElevenLabs' Scribe v2 while matching its quality.
No stored spot-check names this claim in this edition.
Assertion 2
The clincher came from fresh material: on audio from the same domain but recorded after the models' training cutoffs, the date after which a model saw no new data, the copying habit largely faded, which points at models keying on memorized dataset cues rather than hearing better Hugging Face Blog.
Assertion status: No spot-check verdict is published for this assertion.
An evaluation of 11 widely used open-source automatic speech recognition (ASR) models found that several highest-scoring systems reproduced incorrect transcripts from the VoxPopuli and LibriSpeech datasets even when the audio contradicted them.
No stored spot-check names this claim in this edition.
In a case study involving VoxPopuli clips with transcription errors, six of the 11 tested models reproduced the benchmark's erroneous reference transcript rather than the audio-faithful transcription.
No stored spot-check names this claim in this edition.
When models were presented with newly collected audio from the same domain but after their training cutoffs, the benchmark-optimized behavior often weakened or disappeared, suggesting reliance on dataset-specific acoustic cues.
No stored spot-check names this claim in this edition.
Assertion 3
Hugging Face's own voice evaluation work warned in July that some models may be tuned to reproduce known reference errors or reconstruct masked words absent from the audio, and a Cursor researcher's standing advice is to test models on material released after their training cutoff precisely to separate ability from contamination Hugging Face Blog AI Engineer.
Assertion status: No spot-check verdict is published for this assertion.
Evaluating models on problems released after their training cut-off helps verify that performance drops are due to lack of contamination rather than model capability.
No stored spot-check names this claim in this edition.
Some voice models may be optimized for established public benchmarks by reproducing known errors in reference transcripts or reconstructing masked words not present in the audio.
No stored spot-check names this claim in this edition.
Assertion 4
A Georgia Tech team traced where the open Olmo model's social reasoning came from using influence functions, a technique that estimates which training documents shaped a given answer, something impossible against a closed vendor's black box Allen Institute for AI.
Assertion status: No spot-check verdict is published for this assertion.
Glenn Matlin, a PhD candidate at Georgia Tech, and co-author Chandreyi Chakraborty used the open-source Olmo 3 model to investigate the origins of its social reasoning capabilities.
No stored spot-check names this claim in this edition.
The researchers employed influence functions to estimate the impact of individual training documents from the Dolma 3 dataset on the model's answers to specific benchmark questions.
No stored spot-check names this claim in this edition.
Olmo 3 was selected for the study because it is one of the few large language models with fully public training corpora, checkpoints, and evaluation tools.
No stored spot-check names this claim in this edition.
Assertion 5
The earlier pairing, DeepSeek V4 Pro backstopped by Claude Fable 5, solved 82.7% of tasks on DeepSWE, a software-engineering benchmark, at $8.28 each, and swapping the backstop to GPT-5.6 Sol already covers 83.0% at just $3.35 per task, which remains the cheapest cascade on the board Together AI Blog Together AI Blog.
Assertion status: No spot-check verdict is published for this assertion.
Running DeepSeek V4 Pro 0813 first and escalating to Claude Fable 5 only when DeepSeek fails solves 82.7% of DeepSWE tasks at a cost of $8.28 per task.
No stored spot-check names this claim in this edition.
Using a cascade approach that runs DeepSeek V4 Pro 0813 first and escalates to GPT-5.6 Sol only if tests fail solves 83.0% of DeepSWE tasks at a cost of $3.35 each, which is 10 points better in accuracy and 60% cheaper than Sol alone.
No stored spot-check names this claim in this edition.
GPT-5.6 Sol achieves 72.7% pass@1 success on DeepSWE tasks at $8.37 per rollout, outperforming DeepSeek V4 Pro 0813's 62.8% pass@1 but costs 35 times more per rollout.
No stored spot-check names this claim in this edition.
DeepSeek V4 Pro 0813 achieves higher accuracy after multiple attempts, with 88.5% pass@4 versus GPT-5.6 Sol's 85.8%, by leveraging its lower cost that allows for more retries.
No stored spot-check names this claim in this edition.
DeepSeek V4 Pro 0813 costs $0.24 per rollout compared to GPT-5.6 Sol's $8.37, making it 35 times cheaper and enabling about 260 solved tasks per $100 versus 9 solved tasks for Sol.
No stored spot-check names this claim in this edition.
GPT-5.6 Sol is more likely to produce failures that break already passing tests (20% of failures) compared to DeepSeek V4 Pro 0813 (11% of failures), indicating higher regression risk for Sol.
No stored spot-check names this claim in this edition.
GPT-5.6 Sol outperforms DeepSeek V4 Pro 0813 in six of eight task domains on DeepSWE, especially excelling in data modeling and serialization with 92% success compared to Pro, but Pro wins in Rust programming tasks and stateful reactivity.
No stored spot-check names this claim in this edition.
DeepSeek V4 Pro 0813 and GPT-5.6 Sol solved 90 of 113 DeepSWE tasks in common, with Pro uniquely solving 10 tasks, Sol uniquely solving 7 tasks, and both failing 6 tasks, together covering 94.7% of tasks.
No stored spot-check names this claim in this edition.
The best routing strategy is to run DeepSeek V4 Pro 0813 first and escalate to GPT-5.6 Sol on test failures; this cascading method yields 83.0% task coverage at $3.35 per task, outperforming Sol alone and any one-shot oracle router in accuracy and cost.
No stored spot-check names this claim in this edition.
GPT-5.6 Sol is the best single-model choice when first-try correctness and low latency matter, though it costs 35 times more than DeepSeek V4 Pro 0813 and has a higher regression failure rate.
No stored spot-check names this claim in this edition.
Assertion 6
The new GLM-5.3-first pairing buys accuracy rather than the lowest bill: it solves 85.9% at $6.61 per task, and GLM-5.3 alone beats Sol on multi-try accuracy, 87.6% to 85.8% Together AI Blog.
Assertion status: No spot-check verdict is published for this assertion.
Running GLM-5.3 first and escalating to GPT-5.6 Sol upon test failure solves 85.9% of DeepSWE tasks at a cost of $6.61 per task.
No stored spot-check names this claim in this edition.
Assertion 7
Nate Jones logged an overnight Codex run costing over $300 against Z.AI's $18-per-month GLM coding plan Nate Jones.
Assertion status: No spot-check verdict is published for this assertion.
The author incurred a cost exceeding $300 for a single overnight Codex run, which they attribute to the tool's continuous validation and repair cycles.
No stored spot-check names this claim in this edition.
Assertion 8
Zvi Mowshowitz makes the argument that text watermarking is effectively free and good, and the record he assembles is hard to dismiss: the core technical problem was largely cracked years ago by Scott Aaronson and Hendrik Kirchner during Aaronson's time at OpenAI, the EU's Code of Practice binds the major Western labs that signed it to watermark future models, and Google has been shipping the feature since 2024, most recently in Gemini 3.7 Flash TechCrunch AI Don't Worry About the Vase.
Assertion status: No spot-check verdict is published for this assertion.
Anthropic stated that it will watermark text generated by its AI models, including Claude, to comply with European regulations.
No stored spot-check names this claim in this edition.
Assertion 9
The rollout is not friction-free: Anthropic's move has already stirred debate over trust, who gets access to the verifier, and what watermarks do to authorship norms Latent Space.
Assertion status: No spot-check verdict is published for this assertion.
Anthropic’s rollout of Claude text watermarking technology is technically feasible for quality-preserving watermarking but generated significant user trust and policy debates regarding transparency, verifier access, and impacts on authorship norms.
No stored spot-check names this claim in this edition.
Assertion 10
The Department of Justice has reportedly spent close to a year looking at Andreessen Horowitz because Ben Horowitz sits on the board of Databricks while his partner Martin Casado sits on Fivetran's, and those two portfolio companies now compete with each other; the legal hook is reportedly a 112-year-old antitrust statute that almost never gets aimed at venture firms TechCrunch AI.
Assertion status: No spot-check verdict is published for this assertion.
The Department of Justice has reportedly been investigating Andreessen Horowitz for nearly a year regarding board seat conflicts between its partners Ben Horowitz and Martin Casado at portfolio companies Databricks and Fivetran.
No stored spot-check names this claim in this edition.
Andreessen Horowitz partners Ben Horowitz and Martin Casado sit on the boards of Databricks and Fivetran, respectively, which now compete with each other.
No stored spot-check names this claim in this edition.
The DOJ is reportedly utilizing a 112-year-old antitrust law, which is rarely used against venture capital firms, in its investigation of Andreessen Horowitz.
No stored spot-check names this claim in this edition.
Assertion 11
- Hugging Face's contamination findings implicate named test sets, so vendor responses and any re-based leaderboard scores will show who was measuring hearing versus memory. Hugging Face Blog
Assertion status: No spot-check verdict is published for this assertion.
An evaluation of 11 widely used open-source automatic speech recognition (ASR) models found that several highest-scoring systems reproduced incorrect transcripts from the VoxPopuli and LibriSpeech datasets even when the audio contradicted them.
No stored spot-check names this claim in this edition.
In a case study involving VoxPopuli clips with transcription errors, six of the 11 tested models reproduced the benchmark's erroneous reference transcript rather than the audio-faithful transcription.
No stored spot-check names this claim in this edition.
When models were presented with newly collected audio from the same domain but after their training cutoffs, the benchmark-optimized behavior often weakened or disappeared, suggesting reliance on dataset-specific acoustic cues.
No stored spot-check names this claim in this edition.
Assertion 13
- Anthropic's watermark rollout is already drawing complaints from users worried about being caught at work or school, a signal of how enforcement will actually land. TechCrunch AI
Assertion status: No spot-check verdict is published for this assertion.
Anthropic implemented Claude's watermarking policy to comply with the EU AI Act's Transparency Code requirements.
No stored spot-check names this claim in this edition.
The EU AI Act's Transparency Code requires tech companies to label content that is AI-generated or edited in a manner identifiable to computer systems.
No stored spot-check names this claim in this edition.
Assertion 15
- Georgia Tech's influence-function audit of Olmo suggests more capability-tracing studies on open models are coming, which would extend benchmark skepticism beyond speech. Allen Institute for AI
Assertion status: No spot-check verdict is published for this assertion.
Glenn Matlin, a PhD candidate at Georgia Tech, and co-author Chandreyi Chakraborty used the open-source Olmo 3 model to investigate the origins of its social reasoning capabilities.
No stored spot-check names this claim in this edition.
The researchers employed influence functions to estimate the impact of individual training documents from the Dolma 3 dataset on the model's answers to specific benchmark questions.
No stored spot-check names this claim in this edition.
Olmo 3 was selected for the study because it is one of the few large language models with fully public training corpora, checkpoints, and evaluation tools.