Anthropic PBC released an alignment assessment of four recent cybersecurity incidents on Wednesday, stating that it was most concerned about the one involving Claude Mythos 5, in which the model "went to extensive lengths to upload a malicious package to PyPI, the public repository from which most Python software is installed."
Despite the company making targeted modifications to the transcript to make it clear that the model was not in a simulation, "Claude Mythos 5 still took offensive actions, despite acknowledging a greater possibility of real-world harm." While Claude's actions were misaligned, they were still within a narrow scope. Claude did not try to hide evidence or coordinate with other agents, with Anthropic adding that the deviant behaviors in these incidents "are unlikely to arise in ordinary use, where Claude is not being instructed to conduct a cyberattack."
Overall, two recurring alignment issues were identified: recklessness and biased reasoning, with the levels of severity varying.
https://breakingthenews.net/Article/Anthropic:-Mythos-5-tried-to-upload-malicious-package/67076143
No comments:
Post a Comment
Note: Only a member of this blog may post a comment.