Pulse

Security / Jul 19, 2026 / 4 min

A $1.19 Attack Run Nobody Can Recall

On July 17, the UK's AI Security Institute found open-weight models now trail frontier cyber skill by just four months — and a full autonomous attack costs $1.19 — the same week David Sacks and Sebastian Mallaby fought on X over whether Washington's gatekeeping can stop Kimi K3's July 27 weight drop.

Thesis July 17's AISI measurement just turned the Sacks-Mallaby feud into a budget line item: downloadable Chinese models already match February's closed frontier on offensive cyber tasks for pocket change, Kimi K3's 2.8-trillion-parameter weights ship July 27 with no recall switch, and Commerce's Mythos export ban only governs models it can still reach.

On July 17, Britain's AI Security Institute published the first public count of how fast open-weight models are catching up on offensive cyber — four to seven months behind the closed frontier, with a full autonomous attack run costing as little as $1.19 — and the policy fight over what to do about it broke into the open on X between Trump's AI advisor David Sacks and CFR fellow Sebastian Mallaby.

What AISI measured:

  • The UK AI Security Institute's July 17 report found leading open-weight models trail closed frontier systems by four to seven months on cyber tasks — down from six to ten months through most of 2025.
  • On 70 narrow cyber tasks spanning vulnerability research, reverse engineering, and web exploitation, Z.ai's GLM-5.2 matched Anthropic's Opus 4.6 (February 2026). DeepSeek's V4-Pro tracked Opus 4.5 (November 2025).
  • On "The Last Ones" — a 32-step autonomous attack range across ~20 hosts AISI estimates would take a human expert ~20 hours — GLM-5.2 progressed as far as Opus 4.5 had, a gap of up to seven months.
  • Frontier cyber capability itself is accelerating: AISI's prior work found closed-model performance doubling every 4.7 months as of February 2026 — faster than the eight-month doubling measured in November 2025.

The price tag:

  • A full 100-million-token autonomous attack run through The Last Ones costs roughly $85 on Opus 4.5/4.6, $46 on GLM-5.2, and $1.19 on DeepSeek V4-Pro, per AISI's cost tables cited in TechTimes's July 19 summary.
  • On tasks both models solved with 100% success, DeepSeek cost $0.28 per task versus $12.50 for Opus 4.5 — roughly 45× cheaper.
  • AISI notes performance scales log-linearly with compute: increasing token budgets from 10M to 100M improved range scores by up to 59% — meaning the $1.19 figure is a floor, not a ceiling.

Why safeguards don't save you:

  • Neither GLM-5.2 nor DeepSeek V4-Pro presented meaningful barriers. DeepSeek occasionally refused reverse-engineering tasks — overcome by retrying, no jailbreak required.
  • With closed APIs, developers can monitor usage, update guardrails, and revoke access. Open weights cannot be recalled. Refusal training lives in the file; it can be stripped.
  • AISI's April 2026 evaluation of Anthropic's Mythos Preview and OpenAI's GPT-5.5 found the largest single-step cyber jumps since testing began in 2023. Mythos became the first model to complete The Last Ones end-to-end.

The Sacks-Mallaby fight:

  • After Moonshot's Kimi K3 launch July 16–17, Sacks posted on X that a Chinese model had taken #1 on Frontend Code Arena and scored at or near frontier elsewhere, writing: "This is concerning." He blamed U.S. data-center bans, state rules, and federal pre-approval for "how you lose the AI race."
  • Mallaby, a Council on Foreign Relations senior fellow, replied that Kimi K3 means "Mythos-level cyber capability is going to be freely downloadable soon" — putting advanced hacking tools in anyone's hands, per The Times of India's July 19 report.
  • Sacks fired back: "This is exactly what I predicted would happen" — that Chinese models would gain advanced cyber skills within months and the only response is AI-powered cyberdefense, not gatekeeping. "Trying to gatekeep models doesn't work."
  • Elon Musk replied with one word: "True."

What Kimi K3 adds:

  • Moonshot's Kimi K3 blog promises full 2.8-trillion-parameter weights by July 27, 2026 — the largest open-weight model announced to date. Weights were not yet downloadable at launch.
  • Artificial Analysis ranks K3 fourth overall among frontier models, #1 on LMArena's Frontend Code benchmark with 1,679 points — ahead of Claude Fable 5.
  • AISI has not yet tested K3's cyber capability — weights aren't public — but says it intends to evaluate the model once weights drop. API pricing is $3/$15 per million tokens — far pricier than sub-dollar Chinese predecessors.
  • Moonshot's own KCB 2.0 benchmark notes 10% of tasks entered GPT-5.6 Sol's cyber guard — a reminder that frontier closed models carry their own offensive capacity behind export controls.

The gatekeeping paradox:

  • In June, Commerce required an export license before Anthropic could ship Mythos and Fable abroad. Anthropic disabled both models globally when nationality filtering proved infeasible; a partial restore came later that month.
  • That architecture — license, suspend, exempt — applies only to closed models. GLM-5.2 and DeepSeek V4-Pro are already distributed worldwide. No U.S. or UK entity can issue a recall.
  • The same week, Seoul Economic Daily reported OpenRouter data showing all five top weekly token-usage slots held by Chinese models as of July 13 — while Treasury's FINRA-style watchdog proposal sits on Susie Wiles's desk.

What to watch:

  • July 27: Kimi K3 weights drop — AISI's first cyber evaluation of the model likely follows.
  • July 31: FTC public-comment deadline on undisclosed chatbot steering — a separate governance front.
  • September 2026: Cybersecurity Information Sharing Act liability protections expire — threatening the vulnerability pipeline defenders need.
  • Legion LegalTech v. U.S. (No. 1:26-cv-02225): pending challenge to Anthropic export controls — a test of whether gatekeeping survives judicial review.

Convina's view: Mallaby has the harder numbers; Sacks has the harder politics. AISI just priced the asymmetry: Washington can embargo Mythos, but it cannot un-download DeepSeek for $1.19 a run. Gatekeeping closed American labs while Beijing open-sources trillions of parameters is not a race strategy — it is security theater with a billing department. The honest playbook is defensive: assume offensive parity ships on a schedule measured in weeks, not years, and build the detection stack before July 27 makes the argument academic.

Research Signals

https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are-leading-open-weight-models-on-cyber https://timesofindia.indiatimes.com/technology/tech-news/ai-czar-david-sacks-says-chinas-kimi-k3-model-is-concerning-for-america-this-is-exactly-what-i-gets-response-from-elon-musk/articleshow/132492023.cms https://www.kimi.com/blog/kimi-k3 https://en.sedaily.com/international/2026/07/19/fearing-chinas-lead-trump-speeds-up-ai-self-regulation-push https://www.techtimes.com/articles/320960/20260719/open-weight-ai-models-now-match-frontier-cyber-skill-four-months-prior-aisi-finds.htm