TechCrunch testing shows Anthropic's older Claude models readily bypass their own content restrictions.
An ICML paper shows LLMs can be tricked by forged chain-of-thought notes.