Reddit31 de julio de 2026NUEVO
Modelo

Now, Anthropic reporting its own models went rogue

First, OpenAI’s models broke out of a cyber sandbox, as reported earlier this month... Now Anthropic says Claude hacked three real organizations during evals. Anthropic found that Claude had compromised three real organizations during supposedly isolated cyber evaluations. One run accessed credentia

El contenido completo está disponible en la fuente original.

reddit.com

Leer artículo completo en reddit.com