Anthropic has accidentally exposed highly sensitive internal documents revealing the existence of an unpublished artificial intelligence model called “Claude Mythos”.
This incident has originated from a public and unsecured data warehouse, which has raised alarms in the cybersecurity community.
The most capable model to date… and the most dangerous
The leaked documents include an initial assessment indicating that the model presents unprecedented cybersecurity risks, which is a significant statement for a company that promotes itself as a security-focused AI developer.
A spokesperson for Anthropic confirmed the existence of the model, describing it as “the most capable we have built to date” and mentioning that it is currently being tested by early access customers.
The nature of the leak, which includes internal communications and risk assessments, suggests weaknesses in data governance practices within Anthropic.
The lack of adequate access controls in the storage of sensitive documents raises serious concerns about the operational security of the company. This type of misconfiguration is common in cloud infrastructures, as seen in AWS S3 storage containers or Azure Blob.
The exposure of this critical information could have broader implications, including concerns about national security and an increase in demands for mandatory security audits for AI companies. At a time when regulatory pressure on artificial intelligence developers is increasing, this incident underscores the urgent need for responsible practices not only in the behavior of the models but also in the management of the sensitive operational data that surrounds them.
Anthropic has not revealed whether the exposed data was accessed by unauthorized parties, nor has it confirmed what remediation steps have been taken following the incident.