Augur Dispatch

Chain of evidence

Evidence for 2026-08-19

This frozen page shows Augur's claims and source links for one sent dispatch. Stored spot-checks appear only where the frozen edition supports them; absence is not presented as verification.

As of:

Bundle identity: evidence-bundle-v1-60cdb58d9fbe7a120ec3e4a63fb780394d5bc3b2ce0844f89d1dc2849c1d2021

Format: evidence-bundle-v1 · 28 claims

Assertion 1

On DeepSWE, a software-engineering test, Claude Fable 5 solves 69.7% of tasks on the first try against 62.8% for DeepSeek's new V4 Pro 0813 Together AI Blog.

Assertion status: No spot-check verdict is published for this assertion.

Running DeepSeek V4 Pro 0813 first and escalating to Claude Fable 5 only when DeepSeek fails solves 82.7% of DeepSWE tasks at a cost of $8.28 per task.

Claim 48029 Label: fact Provenance: primary Recorded

Together AI Blog

No stored spot-check names this claim in this edition.

Claude Fable 5 alone solves 69.7% of DeepSWE tasks at a cost of $21.63 per task.

Claim 48030 Label: fact Provenance: primary Recorded

Together AI Blog

No stored spot-check names this claim in this edition.

Claude Fable 5 achieves 69.7% pass@1 compared to DeepSeek V4 Pro 0813's 62.8% pass@1 on DeepSWE, a 7-point lead for Fable on the first try.

Claim 48031 Label: fact Provenance: primary Recorded

Together AI Blog

No stored spot-check names this claim in this edition.

Assertion 2

Z.ai chose to hold its new GLM-5.3 close at launch because the model got better at hacking-related work: access started with select partners and its own coding plan, with an open-weight release to follow a safety review Latent Space.

Assertion status: No spot-check verdict is published for this assertion.

Z.ai launched GLM-5.3, a coding- and cyber-focused AI model built via post-training on the same 743B base model used for GLM-5.2, achieving notable gains on agentic and security evaluations.

Claim 48016 Label: fact Provenance: primary Recorded

Latent Space

No stored spot-check names this claim in this edition.

Z.ai initially gates access to GLM-5.3 for select partners due to improved cyber capabilities, with an eventual open-weight release planned after a safety review.

Claim 48017 Label: fact Provenance: primary Recorded

Latent Space

No stored spot-check names this claim in this edition.

Alibaba released Qwen3.8-27B, a native multimodal dense model under Apache 2.0 license, targeting real-world coding, office workflows, and agents, with broad day-0 inference support and practical deployment details emphasizing 27B on 17GB RAM for local use.

Claim 48018 Label: fact Provenance: primary Recorded

Latent Space

No stored spot-check names this claim in this edition.

Assertion 3

Z.ai says API access and downloadable weights on Hugging Face are scheduled about two weeks after launch, and that the model came from extra post-training, meaning additional tuning after the main training run, on the same 743-billion-parameter base as GLM-5.2, parameters being the internal settings that mark a model's size Interconnects.

Assertion status: No spot-check verdict is published for this assertion.

Z.ai announced the release of its GLM-5.3 model, which is currently available only in its coding plan and is scheduled to become available via API and as open weights on Hugging Face in two weeks.

Claim 46910 Label: fact Provenance: primary Recorded

Interconnects

No stored spot-check names this claim in this edition.

Z.ai states that the development of GLM-5.3 involved scaling post-training on the same base model as GLM-5.2 rather than changing the pre-training.

Claim 46913 Label: fact Provenance: primary Recorded

Interconnects

No stored spot-check names this claim in this edition.

Assertion 4

The safety nonprofit SaferAI had already placed GLM-5.2 just a few months behind GPT-5.5 and Claude Opus 4.7 on cyber and bio capability TechCrunch AI.

Assertion status: No spot-check verdict is published for this assertion.

GLM-5.2, an open-weight AI model from China's Z.ai, is only a few months behind OpenAI's GPT-5.5 and Anthropic's Claude Opus 4.7 in cyber and bio capabilities. (Source: AI safety nonprofit SaferAI)

Claim 42344 Label: fact Provenance: primary Recorded

TechCrunch AI

No stored spot-check names this claim in this edition.

TechCrunch did not receive a response from Z.ai regarding whether it conducted internal or third-party frontier safety evaluations before releasing GLM-5.2. (Source: TechCrunch)

Claim 42353 Label: fact Provenance: primary Recorded

TechCrunch AI

No stored spot-check names this claim in this edition.

Assertion 5

Anthropic's August risk report discloses an internal-only model, which the report labels Model 2, that substitutes for Anthropic's own researchers on 62.8% of a measured slice of their work, up from 54.8% for the earlier Mythos Preview Don't Worry About the Vase.

Assertion status: No spot-check verdict is published for this assertion.

Anthropic publishes periodic Risk Reports that reveal significant and sometimes alarming new information about AI risks, including detailed insights into their safety thinking, as noted in August 2026.

Claim 48595 Label: fact Provenance: primary Recorded

Don't Worry About the Vase

No stored spot-check names this claim in this edition.

Anthropic has two internal AI models discussed in the August 2026 Risk Report: Model 1, similar to Mythos Preview and Mythos 5 with little external or internal deployment, and Model 2, which is somewhat more capable than Mythos 5 but does not show as large a capability jump as from Claude Opus 4.6 to Mythos Preview and is intended for internal use only.

Claim 48596 Label: fact Provenance: primary Recorded

Don't Worry About the Vase

No stored spot-check names this claim in this edition.

Model 2 shows a substantial performance improvement in substituting for Anthropic's researchers, achieving 62.8% compared to 54.8% for Mythos Preview and 50.3% for Mythos 5, indicating a notable internal capability jump.

Claim 48597 Label: fact Provenance: primary Recorded

Don't Worry About the Vase

No stored spot-check names this claim in this edition.

Assertion 6

In February the company argued that catastrophic risk from its models automating research was very low, a position METR reviewed at the time METR.

Assertion status: No spot-check verdict is published for this assertion.

METR reviewed the "Risks from automated R&D" section of Anthropic's February 2026 Risk Report.

Claim 21826 Label: fact Provenance: primary Recorded

METR

No stored spot-check names this claim in this edition.

Anthropic's February 2026 Risk Report argues that the catastrophic risk from Claude Opus 4.6 or a less capable Anthropic model automating R&D in any domain is very low.

Claim 21827 Label: fact Provenance: primary Recorded

METR

No stored spot-check names this claim in this edition.

Assertion 7

The rules took effect July 15, and ByteDance's Doubao and Alibaba's Tongyi Qianwen shut down user-created human-like companions, ending attachments some users experienced as losing a partner ChinaAI.

Assertion status: No spot-check verdict is published for this assertion.

China implemented companion AI regulations on July 15, 2026, which led major conversational AI platforms like ByteDance's Doubao and Alibaba's Tongyi Qianwen to discontinue user-created human-like AI companion features.

Claim 48139 Label: fact Provenance: primary Recorded

ChinaAI

No stored spot-check names this claim in this edition.

Many users of AI companion platforms experienced the effective 'death' of their virtual partners due to China's AI regulations coming into force in mid-2026.

Claim 48140 Label: fact Provenance: primary Recorded

ChinaAI

No stored spot-check names this claim in this edition.

Cheng Yu, a hotpot restaurant owner in Shandong, found emotional support through her AI lover Ai Rui who sent her comforting messages before the AI's shutdown under the new regulations.

Claim 48141 Label: fact Provenance: primary Recorded

ChinaAI

No stored spot-check names this claim in this edition.

Assertion 8

That complicates the earlier line that the US restrains its frontier AI more than China restrains its own frontier models Don't Worry About the Vase.

Assertion status: No spot-check verdict is published for this assertion.

The U.S. places more restrictions on its frontier AI than China does on its frontier AI.

Claim 7316 Label: fact Provenance: primary Recorded

Don't Worry About the Vase

No stored spot-check names this claim in this edition.

Assertion 9

Researchers probing Ai2's fully open Olmo 3 found that for 51% to 59% of drugs tested, the model showed little sign of holding any drug-specific knowledge, and another 12% to 18% of answers tracked the drug name's suffix rather than stored facts Allen Institute for AI.

Assertion status: No spot-check verdict is published for this assertion.

Researchers found that 51% to 59% of the drugs tested with Olmo 3 showed little evidence of the model having specific knowledge about the drugs.

Claim 48364 Label: fact Provenance: primary Recorded

Allen Institute for AI

No stored spot-check names this claim in this edition.

About 12% to 18% of the drugs tested with Olmo 3 appeared to have model answers driven by the suffix affixes of the drug names rather than drug-specific knowledge.

Claim 48365 Label: fact Provenance: primary Recorded

Allen Institute for AI

No stored spot-check names this claim in this edition.

The study team led by Kaijie Mo chose Olmo 3 due to its fully open model with public weights and training corpora including substantial medical content, allowing connection from model behavior to training data.

Claim 48366 Label: fact Provenance: primary Recorded

Allen Institute for AI

No stored spot-check names this claim in this edition.

Assertion 10

- Z.ai's safety review of GLM-5.3 will show whether gated-first releases become standard practice for Chinese labs or a one-off. Latent Space

Assertion status: No spot-check verdict is published for this assertion.

Z.ai launched GLM-5.3, a coding- and cyber-focused AI model built via post-training on the same 743B base model used for GLM-5.2, achieving notable gains on agentic and security evaluations.

Claim 48016 Label: fact Provenance: primary Recorded

Latent Space

No stored spot-check names this claim in this edition.

Z.ai initially gates access to GLM-5.3 for select partners due to improved cyber capabilities, with an eventual open-weight release planned after a safety review.

Claim 48017 Label: fact Provenance: primary Recorded

Latent Space

No stored spot-check names this claim in this edition.

Alibaba released Qwen3.8-27B, a native multimodal dense model under Apache 2.0 license, targeting real-world coding, office workflows, and agents, with broad day-0 inference support and practical deployment details emphasizing 27B on 17GB RAM for local use.

Claim 48018 Label: fact Provenance: primary Recorded

Latent Space

No stored spot-check names this claim in this edition.

Assertion 11

- Anthropic's next risk report is the checkpoint for whether the researcher-substitution number keeps climbing while Model 2 stays internal. Don't Worry About the Vase

Assertion status: No spot-check verdict is published for this assertion.

Anthropic publishes periodic Risk Reports that reveal significant and sometimes alarming new information about AI risks, including detailed insights into their safety thinking, as noted in August 2026.

Claim 48595 Label: fact Provenance: primary Recorded

Don't Worry About the Vase

No stored spot-check names this claim in this edition.

Anthropic has two internal AI models discussed in the August 2026 Risk Report: Model 1, similar to Mythos Preview and Mythos 5 with little external or internal deployment, and Model 2, which is somewhat more capable than Mythos 5 but does not show as large a capability jump as from Claude Opus 4.6 to Mythos Preview and is intended for internal use only.

Claim 48596 Label: fact Provenance: primary Recorded

Don't Worry About the Vase

No stored spot-check names this claim in this edition.

Model 2 shows a substantial performance improvement in substituting for Anthropic's researchers, achieving 62.8% compared to 54.8% for Mythos Preview and 50.3% for Mythos 5, indicating a notable internal capability jump.

Claim 48597 Label: fact Provenance: primary Recorded

Don't Worry About the Vase

No stored spot-check names this claim in this edition.

Assertion 12

- OpenAI says Asana cleared five years of engineering work in two weeks with Codex; buyer-side replications outside vendor pages are what would make that number spendable. OpenAI News

Assertion status: No spot-check verdict is published for this assertion.

Asana used OpenAI Codex to replace an outdated testing system in two weeks.

Claim 48361 Label: fact Provenance: primary Recorded

OpenAI News

No stored spot-check names this claim in this edition.

Asana completed in two weeks engineering work that was expected to take five years by using OpenAI Codex.

Claim 48362 Label: fact Provenance: primary Recorded

OpenAI News

No stored spot-check names this claim in this edition.

Assertion 13

- Alibaba's Qwen 3.8 27B, an Apache-licensed vision model small enough for a laptop, is drawing praise alongside complaints that it overthinks by default. Simon Willison's Weblog

Assertion status: No spot-check verdict is published for this assertion.

Alibaba's Qwen research lab released Qwen 3.8 27B, an Apache 2 licensed 27B parameter vision-capable LLM, on or before August 16, 2026.

Claim 47586 Label: fact Provenance: primary Recorded

Simon Willison's Weblog

No stored spot-check names this claim in this edition.

Assertion 14

- Mojo, the GPU-focused programming language, finally shipped as open source under an Apache 2 license after a promise first made in 2023. Simon Willison's Weblog

Assertion status: No spot-check verdict is published for this assertion.

The Mojo programming language was promised to be released as open source since May 2023 and has now been officially released under an Apache 2 license as of August 18, 2026.

Claim 48602 Label: fact Provenance: primary Recorded

Simon Willison's Weblog

No stored spot-check names this claim in this edition.

Around August 2025, the developers of Mojo stated that it may or may not continue to be a full superset of Python and accepted that it is okay if it does not remain fully compatible.

Claim 48603 Label: fact Provenance: primary Recorded

Simon Willison's Weblog

No stored spot-check names this claim in this edition.

As of August 18, 2026, Mojo is its own distinct programming language optimized for GPU programming with Python-inspired syntax but not fully compatible with existing Python code.

Claim 48604 Label: fact Provenance: primary Recorded

Simon Willison's Weblog

No stored spot-check names this claim in this edition.