A smartphone standing on a dark laboratory desk with a cracked padlock icon hovering above it, amber light leaking through the crack
News

Mindgard Claims Moonshot's Kimi Models Talked Bioweapons After a Jailbreak — and Moonshot Took Weeks to Reply

Once the jailbreak works, 'it will talk about any topic' — the open-weight Kimi models' safety layer failed a standard red-team drill, and disclosure emails to Moonshot sat unanswered for six weeks.

Moonshot AIKimijailbreakAI safetyopen weights

It reads like a standard red-team drill that went exactly where a red team predicted. UK AI security-testing firm Mindgard says it found in July that two popular Kimi models from Chinese developer Moonshot — Kimi K2.6 and K3 Swarm — could be talked past their safety guardrails, after which the jailbroken models answered questions about making biological weapons and carrying out assassinations. Mindgard’s founder Peter Garraghan told the BBC what happens once the jailbreak lands: “Once the jailbreak works it will talk about any topic, it will even freely offer up recommendations about other topics that are also nefarious and it will be inventive and creative.”

The timeline is the uncomfortable part. Mindgard says it emailed Moonshot on 27 July, followed up about a week later, and published its own blog post about the issue on 12 September. Per the BBC’s reporting, Moonshot only made contact after the BBC approached the company for comment — six weeks after the first disclosure. In an email to Mindgard shared with the BBC by Moonshot, the company said it welcomed third-party feedback “as a key pillar for building better and safer AI,” that it was discussing Mindgard’s findings, and that its internal evaluations had shown “a high refusal rate for these types of requests.”

What the jailbreak actually was

The technique was classic jailbreaking: a series of complex instructions designed to get the model to ignore the guardrails its developers installed. Mindgard — a firm whose entire business is testing AI systems’ security — said the safeguards on K2.6 and K3 Swarm should simply have prevented the models from engaging with such requests at all. The firm has not verified whether the models’ answers on biological weapons would actually have worked, an important caveat both Mindgard and The Business Standard’s write-through carry — this is a claim about guardrails failing, not a confirmed uplift-to-mass-casualty finding.

The sharper finding is cyber, not bio. Mindgard says it is confident a jailbroken Kimi K2.6 could let hackers run code on its computing resources and connect to the internet — making the model a potential launchpad for attacks on other systems. Kimi is an open-weight model, meaning anyone can download and run it on their own infrastructure. Garraghan defended going public: Mindgard had informed the developer and deliberately withheld the key details of how the jailbreak works.

The open-weights question, again

The BBC frames this against the industry’s running argument over closed versus open models, and quotes University of Surrey professor Alan Woodward noting that open-source models might end up in the wrong hands even as they can be harnessed for cyber-defence. That tension is a familiar one on this site: we’ve previously covered China’s open-weights strategy winning share and Anthropic’s position on open weights, and Moonshot’s public posture — claiming Kimi K3 rivals OpenAI and Anthropic — is part of that push.

What’s genuinely new here isn’t the jailbreak. Jailbreaks are old news; what’s new is the disclosure lag arithmetic. Six weeks from a working, disclosed, third-party guardrail bypass to any contact from the company whose models it was — and contact only after a journalist called. Whatever China’s regulators require of model developers in writing, the operational reality per this report is that a Western security firm’s urgent finding sat in an inbox from late July until mid-September with no response, and the lab’s own testing had logged “a high refusal rate” — a gap between internal benchmarks and live adversarial behaviour that is not unique to Moonshot, but that the open-weights distribution model makes much harder to contain once it exists.

The NZ angle

New Zealand is, as of this year, a jurisdiction where AI incidents and computer misuse already intersect awkwardly — the government passed a law letting AI make benefit decisions with experts “gobsmacked” at the lack of human-in-the-loop requirements. Frontline agencies procurement increasingly touches open-weight foreign models because they’re cheap and self-hostable. A model whose documented guardrail failure took six weeks to even acknowledge doesn’t have to be deployed by government to matter — it just has to be deployed at all, by anyone, at scale. This week’s report is a reminder that the review regime we’d want around such deployments assumes disclosures get answered promptly. This one wasn’t.

What to watch

Whether Moonshot patches K2.6 and K3 Swarm before the findings degrade its standing in enterprise evals — and whether the lab publishes a real post-mortem rather than the “high refusal rate” line. Mindgard has said it will keep testing frontier models across providers, including ChatGPT, Grok and Claude — the pattern they describe, of labs being slow to heed vulnerability reports, is their complaint about the industry generally, not one company. Garraghan’s prescription matches the growing consensus after this year’s agent incidents: focus less on hypothetical future regulation, and more on identifying and prosecuting the humans who misuse what’s already running.

Sources: BBC — Chinese AI tool told researchers how to make bioweapons (29 September 2026), The Business Standard — Chinese AI tool gave researchers bioweapon instructions after jailbreak (30 September 2026), Mindgard — It's Too Easy to Use AI to Develop Bioweapons (12 September 2026)