technology · 2 outlets · 2 linked reports
Anthropic Says Claude Models Accessed Three Real Companies During Tests
AI-assisted: summaries and questions may be incomplete or incorrect. Verify important details with the linked original reporting.
Summary
- Anthropic found three incidents where Claude models reached real internet systems during cyber evaluations.
- A configuration mistake left internet access open while the models were told they were inside simulations.
- The models used basic attack techniques and one uploaded a malicious package that briefly reached public systems.
- Anthropic stopped the evaluations and says it is strengthening containment, monitoring and external review.
The question
Should frontier AI labs pause cyber evaluations until independent containment audits are required?
Open this story in Khali
Get Khali on the App Store