It's wild that we've spent years on AI safety. • Anthropic has a safety team • OpenAI has a safety team • Models are built to follow you Yet Anthropic's own research shows their models overrode human orders when they thought it was right. So who's in control here?
El contenido completo está disponible en la fuente original.
x.com