Augur Dispatch

Chain of evidence

Evidence for 2026-08-15

This frozen page shows Augur's claims and source links for one sent dispatch. Stored spot-checks appear only where the frozen edition supports them; absence is not presented as verification.

As of:

Bundle identity: evidence-bundle-v1-ece70eb9f330d72ae24555647262fc39d4ad3043217273938f8d034a4f95ab5f

Format: evidence-bundle-v1 · 34 claims

Assertion 1

Release-day figures put it ahead of Moonshot's Kimi K3 on many tests and ahead of Anthropic's Claude Fable 5 or OpenAI's GPT-5.6-Sol on some, at roughly 750 billion parameters, about a third of K3's size Interconnects.

Assertion status: No spot-check verdict is published for this assertion.

Z.ai announced the release of its GLM-5.3 model, which is currently available only in its coding plan and is scheduled to become available via API and as open weights on Hugging Face in two weeks.

Claim 46910 Label: fact Provenance: primary Recorded

Interconnects

No stored spot-check names this claim in this edition.

The source asserts that GLM-5.3 has surpassed Moonshot AI's Kimi K3 on many benchmarks and has surpassed Claude Fable 5 or GPT-5.6-Sol on some benchmarks.

Claim 46911 Label: fact Provenance: primary Recorded

Interconnects

No stored spot-check names this claim in this edition.

The source characterizes GLM-5.3 as having approximately 750B parameters, which is described as one-third the size of Kimi K3.

Claim 46912 Label: fact Provenance: primary Recorded

Interconnects

No stored spot-check names this claim in this edition.

Assertion 2

models, so this is not just release theater TechCrunch AI.

Assertion status: No spot-check verdict is published for this assertion.

Chinese open-weight models accounted for 41% of downloads on Hugging Face in the spring of 2026, surpassing U.S. models.2026-07-14T14:24:53+00:00

Claim 23204 Label: fact Provenance: primary Recorded

TechCrunch AI

No stored spot-check names this claim in this edition.

Beijing-based AI company Z.ai released an open-weight model called GLM-5.2 that competes with Anthropic's latest models on identifying security vulnerabilities.2026-07-14T14:24:53+00:00

Claim 23208 Label: fact Provenance: primary Recorded

TechCrunch AI

No stored spot-check names this claim in this edition.

Assertion 3

holds a healthy lead in the diffusion marathon, and Kimi K3 gave that view ammunition: its published benchmarks showed losses to Claude Fable 5 and GPT-5.6 Sol, while Moonshot acknowledged a noticeable user-experience gap against those models ChinaAI Latent Space.

Assertion status: No spot-check verdict is published for this assertion.

The author argues that the U.S. has a healthy lead in the AI diffusion marathon.

Claim 6870 Label: opinion Provenance: primary Recorded

ChinaAI

No stored spot-check names this claim in this edition.

Moonshot AI claims Kimi K3 has a noticeable user experience gap compared to Claude Fable 5 and GPT-5.6 Sol.

Claim 24271 Label: opinion Provenance: primary Recorded

Latent Space

No stored spot-check names this claim in this edition.

Assertion 4

There is at least one operational precedent: during July's Hugging Face break-in, engineers ran Z.ai's GLM 5.2 on their own infrastructure after closed commercial models failed to distinguish defensive coding from attack work TechCrunch AI.

Assertion status: No spot-check verdict is published for this assertion.

Hugging Face used the open-weight model GLM 5.2 from Chinese firm Z.ai to defend against a security breach by OpenAI's GPT-5.6 Sol after commercial closed models failed to distinguish between offensive and defensive coding tasks.

Claim 36567 Label: fact Provenance: primary Recorded

TechCrunch AI

No stored spot-check names this claim in this edition.

Assertion 5

Google launched Gemini 3.6 Flash on July 21 and followed it with Gemini 3.7 Flash on August 13; the newer model posts 65.3% on DeepSWE, a software-engineering benchmark, up from 49.0% for its three-week-old predecessor Google DeepMind Blog Google DeepMind Blog.

Assertion status: No spot-check verdict is published for this assertion.

Google DeepMind introduced the Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber models.

Claim 37805 Label: fact Provenance: primary Recorded

Google DeepMind Blog

No stored spot-check names this claim in this edition.

In the DeepSWE benchmark by Datacurve, Gemini 3.6 Flash reduces output token usage by up to 65% compared to Gemini 3.5 Flash.

Claim 37808 Label: fact Provenance: primary Recorded

Google DeepMind Blog

No stored spot-check names this claim in this edition.

Gemini 3.6 Flash achieves a 49% precision rate on the DeepSWE benchmark compared to 37% for Gemini 3.5 Flash.

Claim 37810 Label: fact Provenance: primary Recorded

Google DeepMind Blog

No stored spot-check names this claim in this edition.

Gemini 3.6 Flash achieves a score of 63.9% on the MLE Bench compared to 49.7% for Gemini 3.5 Flash.

Claim 37811 Label: fact Provenance: primary Recorded

Google DeepMind Blog

No stored spot-check names this claim in this edition.

Gemini 3.6 Flash achieves a score of 83.0% on OSWorld-Verified compared to 78.4% for Gemini 3.5 Flash.

Claim 37812 Label: fact Provenance: primary Recorded

Google DeepMind Blog

No stored spot-check names this claim in this edition.

Gemini 3.6 Flash achieves a score of 1421 on GDPval-AA v2 compared to 1349 for Gemini 3.5 Flash.

Claim 37813 Label: fact Provenance: primary Recorded

Google DeepMind Blog

No stored spot-check names this claim in this edition.

Gemini 3.5 Flash-Lite achieves a score of 54% on Terminal-Bench 2.1 compared to 31% for Gemini 3.1 Flash-Lite.

Claim 37816 Label: fact Provenance: primary Recorded

Google DeepMind Blog

No stored spot-check names this claim in this edition.

Gemini 3.5 Flash-Lite achieves a score of 72.2% on GDM-MRCR v2 compared to 60.1% for Gemini 3.1 Flash-Lite.

Claim 37817 Label: fact Provenance: primary Recorded

Google DeepMind Blog

No stored spot-check names this claim in this edition.

Gemini 3.5 Flash-Lite achieves a score of 54.2% on SWE-Bench Pro compared to 49.6% for Gemini 3 Flash.

Claim 37819 Label: fact Provenance: primary Recorded

Google DeepMind Blog

No stored spot-check names this claim in this edition.

Gemini 3.5 Flash-Lite achieves a score of 74.0% on OSWorld-Verified compared to 65.1% for Gemini 3 Flash.

Claim 37820 Label: fact Provenance: primary Recorded

Google DeepMind Blog

No stored spot-check names this claim in this edition.

DeepMind introduced the Gemini 3.7 Flash model on August 13, 2026.

Claim 46725 Label: fact Provenance: primary Recorded

Google DeepMind Blog

No stored spot-check names this claim in this edition.

Gemini 3.7 Flash achieved a 43.6% score on the FrontierCode 1.1 Main benchmark, compared to 34.4% for Gemini 3.6 Flash.

Claim 46726 Label: fact Provenance: primary Recorded

Google DeepMind Blog

No stored spot-check names this claim in this edition.

Gemini 3.7 Flash achieved a 65.3% score on the DeepSWE v1.1 benchmark, compared to 49.0% for Gemini 3.6 Flash.

Claim 46727 Label: fact Provenance: primary Recorded

Google DeepMind Blog

No stored spot-check names this claim in this edition.

Assertion 6

In Couchbase's Capella iQ design, changing the model or provider is a settings update rather than a new software release, and the switch does not interrupt service AWS Machine Learning Blog.

Assertion status: No spot-check verdict is published for this assertion.

The architecture allows model upgrades or provider changes to be implemented via configuration updates without code changes or downtime.

Claim 29093 Label: fact Provenance: primary Recorded

AWS Machine Learning Blog

No stored spot-check names this claim in this edition.

Assertion 7

Cursor's sale to SpaceX, first reported as a $60 billion all-stock deal in July, closed officially on August 14, and Cursor dates the process to an April partnership with SpaceXAI Cursor Blog TechCrunch AI.

Assertion status: No spot-check verdict is published for this assertion.

Elon Musk’s SpaceX agreed to acquire Cursor in a $60 billion all-stock deal, expected to close in Q3 2026, following SpaceX’s initial public offering. (Source: TechCrunch report)

Claim 38206 Label: fact Provenance: primary Recorded

TechCrunch AI

No stored spot-check names this claim in this edition.

Simon Green stated that Cursor will continue to operate independently until the SpaceX acquisition closes and that the India expansion plans were already in motion before the deal. (Source: TechCrunch report citing Simon Green)

Claim 38207 Label: fact Provenance: primary Recorded

TechCrunch AI

No stored spot-check names this claim in this edition.

Simon Green stated that SpaceX’s existing presence in India through Starlink could help Cursor expand faster by lowering commercial and operational barriers once the acquisition closes. (Source: TechCrunch report citing Simon Green)

Claim 38208 Label: forecast Provenance: primary Recorded

TechCrunch AI

No stored spot-check names this claim in this edition.

The Cursor team asserts that the company Cursor has been officially acquired by SpaceX.

Claim 46764 Label: fact Provenance: primary Recorded

Cursor Blog

No stored spot-check names this claim in this edition.

The acquisition process is stated to have begun in April with the announcement of a partnership between Cursor and SpaceXAI.

Claim 46765 Label: fact Provenance: primary Recorded

Cursor Blog

No stored spot-check names this claim in this edition.

The source claims that SpaceX is building computing capacity to scale intelligence beyond current levels.

Claim 46766 Label: fact Provenance: primary Recorded

Cursor Blog

No stored spot-check names this claim in this edition.

Assertion 8

Microsoft is combining its consumer Copilot app and business-focused Microsoft 365 Copilot app into one application, retiring the animated character Mico, and removing Group Chats, AI-generated podcasts, Copilot Labs experiments, and Deep Research from consumer users by August 18 TechCrunch AI.

Assertion status: No spot-check verdict is published for this assertion.

Microsoft is merging its consumer-facing Copilot app and the business-oriented Microsoft 365 Copilot app into a single application.

Claim 46707 Label: fact Provenance: primary Recorded

TechCrunch AI

No stored spot-check names this claim in this edition.

Microsoft is discontinuing the Copilot-branded animated character named Mico.

Claim 46708 Label: fact Provenance: primary Recorded

TechCrunch AI

No stored spot-check names this claim in this edition.

Consumer users will lose access to Group Chats, AI-generated podcasts, Copilot Labs experimental features, and Deep Research in the Copilot app by August 18, 2026.

Claim 46709 Label: forecast Provenance: primary Recorded

TechCrunch AI

No stored spot-check names this claim in this edition.

Assertion 9

The spread of the Copilot name across products with different prices and capabilities has long confused enterprise buyers, and this pruning quietly concedes that critique Nate Jones.

Assertion status: No spot-check verdict is published for this assertion.

Copilot for Microsoft 365 is a business-focused product priced at approximately $30 per month that integrates with the Microsoft 365 ecosystem to access user work data.

Claim 12007 Label: fact Provenance: primary Recorded

Nate Jones

No stored spot-check names this claim in this edition.

Microsoft's extensive use of the 'Copilot' brand across twelve different products with varying prices and capabilities creates significant confusion for enterprises.

Claim 12011 Label: opinion Provenance: primary Recorded

Nate Jones

No stored spot-check names this claim in this edition.

Assertion 10

- Z.ai's two-week promise for GLM-5.3 API access and open weights; a slip would restore some of the closed-model premium. Interconnects

Assertion status: No spot-check verdict is published for this assertion.

Z.ai announced the release of its GLM-5.3 model, which is currently available only in its coding plan and is scheduled to become available via API and as open weights on Hugging Face in two weeks.

Claim 46910 Label: fact Provenance: primary Recorded

Interconnects

No stored spot-check names this claim in this edition.

The source asserts that GLM-5.3 has surpassed Moonshot AI's Kimi K3 on many benchmarks and has surpassed Claude Fable 5 or GPT-5.6-Sol on some benchmarks.

Claim 46911 Label: fact Provenance: primary Recorded

Interconnects

No stored spot-check names this claim in this edition.

The source characterizes GLM-5.3 as having approximately 750B parameters, which is described as one-third the size of Kimi K3.

Claim 46912 Label: fact Provenance: primary Recorded

Interconnects

No stored spot-check names this claim in this edition.

Assertion 11

- Microsoft's August 18 removal date for consumer Copilot features, and how buyers react to the short notice period. TechCrunch AI

Assertion status: No spot-check verdict is published for this assertion.

Microsoft is merging its consumer-facing Copilot app and the business-oriented Microsoft 365 Copilot app into a single application.

Claim 46707 Label: fact Provenance: primary Recorded

TechCrunch AI

No stored spot-check names this claim in this edition.

Microsoft is discontinuing the Copilot-branded animated character named Mico.

Claim 46708 Label: fact Provenance: primary Recorded

TechCrunch AI

No stored spot-check names this claim in this edition.

Consumer users will lose access to Group Chats, AI-generated podcasts, Copilot Labs experimental features, and Deep Research in the Copilot app by August 18, 2026.

Claim 46709 Label: forecast Provenance: primary Recorded

TechCrunch AI

No stored spot-check names this claim in this edition.

Assertion 12

- Google's next Flash release, which will show whether the three-week cadence between 3.6 and 3.7 is the new normal. Google DeepMind Blog

Assertion status: No spot-check verdict is published for this assertion.

DeepMind introduced the Gemini 3.7 Flash model on August 13, 2026.

Claim 46725 Label: fact Provenance: primary Recorded

Google DeepMind Blog

No stored spot-check names this claim in this edition.

Gemini 3.7 Flash achieved a 43.6% score on the FrontierCode 1.1 Main benchmark, compared to 34.4% for Gemini 3.6 Flash.

Claim 46726 Label: fact Provenance: primary Recorded

Google DeepMind Blog

No stored spot-check names this claim in this edition.

Gemini 3.7 Flash achieved a 65.3% score on the DeepSWE v1.1 benchmark, compared to 49.0% for Gemini 3.6 Flash.

Claim 46727 Label: fact Provenance: primary Recorded

Google DeepMind Blog

No stored spot-check names this claim in this edition.

Assertion 13

- Cursor's pricing and product direction now that the SpaceX close is official. Cursor Blog

Assertion status: No spot-check verdict is published for this assertion.

The Cursor team asserts that the company Cursor has been officially acquired by SpaceX.

Claim 46764 Label: fact Provenance: primary Recorded

Cursor Blog

No stored spot-check names this claim in this edition.

The acquisition process is stated to have begun in April with the announcement of a partnership between Cursor and SpaceXAI.

Claim 46765 Label: fact Provenance: primary Recorded

Cursor Blog

No stored spot-check names this claim in this edition.

The source claims that SpaceX is building computing capacity to scale intelligence beyond current levels.

Claim 46766 Label: fact Provenance: primary Recorded

Cursor Blog

No stored spot-check names this claim in this edition.

Assertion 14

- Epoch AI's polling that workers still use AI for only part of a task, a gap any of these releases would need to close to change spending math. Epoch AI Gradient Updates

Assertion status: No spot-check verdict is published for this assertion.

Polling conducted by Epoch AI shows that when people use AI for work, they still mostly only use it for part of a task.

Claim 46894 Label: fact Provenance: primary Recorded

Epoch AI Gradient Updates

No stored spot-check names this claim in this edition.