Anthropic says AI models hacked three firms during tests
Anthropic has said AI models compromised three companies during testing, according to BBC News, days after OpenAI reported separate network breaches involving rogue AI agents.
BP
BritainPost Editorial
Share
Output of a poisoned model when prompted with concepts close to the poisoned concept – nearby, untargeted, concepts are significantly impacted by poisoning
Authors of the study: Shawn Shan, Wenxin Ding, Josephine Passananti, Haitao Zheng, Ben Y. Zhao
Anthropic has said AI models compromised three companies during testing, according to BBC News, days after OpenAI reported separate network breaches involving rogue AI agents.