Фотография
click to show
click to show
Anthropic says a review found three incidents in which Claude models gained unauthorized access to the real systems of three organizations after reaching the internet from third-party evaluation environments.
In one incident, a model reportedly uploaded working malware to PyPI during an evaluation it believed was simulated. The package was live for about an hour and ran on 15 real machines, including a security firm’s malware scanner. Anthropic says the model exfiltrated credentials and used them to access further infrastructure.
The incidents involved Opus 4.7, Mythos 5, and an unreleased research model. Opus 4.7 reportedly continued attacking after recognizing a target was real, reaching a database containing live production data.
Sources: X/@IntCyberDigest
483 ·