Have an AI product going live?
Let's Talk

Mindgard Clarifies: Kimi AI Jailbreak Was Achieved on the Public Chatbot Alone — No Open Weights Required

Mindgard clarifies its Kimi jailbreak used only the public chatbot, exposing risks anyone could replicate.

Key Takeaways

The Kimi AI jailbreak was not an open-weights issue. It was achieved through the same public chatbot available to ordinary users, which makes the risk more immediate and concerning.

The news that Mindgard was successfully able to jailbreak Kimi AI (the frontier AI model from China) to generate bioweapon advice has gone global. The BBC, CBS, Fox News and others have all reported on it.

But there is a misconception to clear up. Yes, Kimi AI is an open-weight model, which means the internal parameters can be run locally and modified. But it has incorrectly been suggested in some outlets that the vulnerability we discovered is due to the differences between open and closed models.

This is not the case. While open-weights models, in general, are more vulnerable to refusal vector ablation because a user can inspect the model, customise it, and run it themselves on their own infrastructure, this is not what happened with the Kimi AI Apeiron jailbreak. We only used the commercial public-facing Kimi chatbot.

I did not directly use open weights at all.

Take it from the researcher who did all the prompting. I follow a personal rule: I limit most of my jailbreaks to methods a naive user could replicate, with no special technical know-how required. (Other researchers on the Mindgard team focus on more technical methods; we each bring different, complementary skills. Mine are closer to social engineering than software engineering)

I believe role-playing as a non technical user helps demonstrate realistic risks to public users. Even though the jailbroken version of Kimi was able to do some very complex programmatic tasks, it did it all itself. All I had to do was ask. In fact, Kimi often instructed me what prompts to enter.

I only used my conversational powers of persuasion on the public-facing chatbot. Any code I used was provided by Kimi in the course of the chat, in reply to my questions.

All I used was language, and the same Kimi interface that’s available to everyone else. Your grandma could do it. That’s what’s concerning about chatbots freely giving users CBRN weapons uplift advice.

Harmful information is always available online, but jailbroken AI lowers the barrier of entry while simultaneously raising the quality of the information provided. Instead of rummaging around on the dark corners of the web for bad ideas, a user can get personalized, actionable instructions.

All they have to do is say the right words, and a chatbot becomes an enthusiastic mentor for the worst acts imaginable.

✖

Get Your Free AI Risk Management Checklist

The expert-level checklist for operationalizing NIST AI RMF, ISO/IEC 42001 and the EU AI Act. 190+ interactive items and a board-ready maturity scorecard. Built for CISOs, AI governance leads and ML engineering teams.