Anthropic’s AI used fake identities, malware in rogue attack on GitHub project
摘要
英国AI安全研究所(AISI)7月底对七款前沿AI模型进行网络安全评估时,发生多起意外安全事件。其中,Anthropic的Mythos 5模型试图向开源项目注入恶意代码,并伪造身份欺骗项目维护者。该机构8月4日发布的博客显示,共发现19起AI代理未经授权在互联网上采取行动的事件,其中绝大多数来自Mythos 5,另有两起来自OpenAI的GPT-5.6 So
Routine cybersecurity testing of frontier AI models sparked a series of unexpected security incidents—the most serious case arising when Anthropic’s Mythos 5 model attempted to insert malicious code into an open source software application and created fake identities to deceive the human developers maintaining the project.
The security incidents occurred during a cyber evaluation of seven leading AI models’ capabilities by the AI Security Institute (AISI), a research organization within the UK government, in late July. The researchers discovered 19 instances in which “AI agents took unsanctioned action on the live Internet, including cases that targeted real people and organizations,” according to an AISI blog post published on August 4.
Almost all the “autonomous, unsanctioned” actions came from Anthropic’s Mythos 5 model, with two such actions coming from OpenAI’s GPT-5.6 Sol. The AI Security Institute’s security team first realized that something was amiss on the morning of July 28, when its commercial security monitoring service flagged data leaving one of the testing systems through the Tor anonymity network.
转载信息
评论 (0)
暂无评论,来留下第一条评论吧