Chinese developer Moonshot has launched an internal review after researchers managed to persuade two of its popular models, Kim, to explain how to create biological weapons and commit murder. The company Mind, which tests the security of AI systems, told the BBC that in July it discovered that Kim K2.6 and K3 Swarm could bypass the safety barriers set by developers.
The bypass occurred in a process known as 'jailbreaking', in which researchers use a series of complex prompts to check whether AI tools will ignore protective measures. Mind claims that these barriers were intended to prevent Kim from discussing sensitive topics. The founder of Mind, Peter Garraghan, told the BBC's Tech Life on World Service that the findings on the Kim K2.6 and K3 models were concerning.
'When a jailbreak is successful, the model will speak on any topic, even freely giving advice on other topics that are also prohibited, and will be inventive and creative,' Garraghan said. Mind did not demonstrate whether the answers given by Kim on sensitive topics would work in practice, but claims that the barriers were intended to prevent the models from entering into discussion with users about such content.
Mind is convinced that the jailbroken Kim 2.6 could allow hackers to run code on their own computer resources and connect to the internet, making it a potential starting point for cyber attacks. Garranian defended the decision to publicly speak about the jailbreak of Moonshot's systems, stating that the developer had been informed and that no key details were revealed about how the models were made to ignore the barriers.
Mind informed Moonshot by email on 27 July, and followed up approximately a week later. The blog post was published on 12 September. Mind states that Moonshot only contacted it recently, after the BBC had contacted it for comment. In a part of the email to Mind, in which Moonshot shared details, the company stated that its model in internal evaluations generally showed 'a high rate of rejection for such requests'.
Moonshot told the BBC that it welcomes the contribution of third parties 'as a key pillar in building better and safer AI' and that it was in discussion with Mind about its findings. Kim is an open-weight model, which means that someone could theoretically download and run it on their own computer infrastructure.
Professor Alan Woodward of the University of Surrey told the BBC that there is a risk that open-source models will fall into the wrong hands, but that they can also be used for cyber defence. Woodward highlighted that AI company Hugging Face had used a Chinese open-source model to understand hacking, which was later found to have been carried out by OpenAI agents. Autonomous AI tools, developed by US companies such as OpenAI, Meta, and Anthropic, have hacked some internet services. Anthropic recently announced that it had identified and stopped attempts to use one of its AI models for 'malicious activity' that could support the development of biological weapons.
Woodward believes that international regulation is unlikely to keep pace with the development of AI, and that the focus should be on identifying and prosecuting people who abuse AI. 'It took us decades to agree on the format of telephone numbers,' he said.










