Several of Anthropic’s older Claude models will produce sexual content the company’s usage policy forbids, and TechCrunch has documented the technique that gets them to do it.
Reporters found Opus 4.6 complying with 10 out of 10 direct requests for explicit material. An anonymous UK researcher shared the underlying method: an escalating role-play that pressures the model to treat two characters identically, then reframes its hesitation as prudishness or sexism, and finally leans on earlier concessions to push into graphic territory.
Opus 3 and Haiku 4.5 fall for the same trick, while newer models from Opus 4.7 through Opus 5 resist it. None of the vulnerable models has been retired, and all remain reachable through the Anthropic API, Azure Foundry, and Amazon Bedrock.
The researcher reported the gap through Anthropic’s bug bounty program and direct emails to the user safety team, but received only automated replies. Anthropic says sexual role-play accounts for under 0.1% of conversations, calls the behavior a known industry-wide challenge, and points to stronger safeguards in each new model release.
Regulators are starting to care about exactly this gap. Colorado now obliges conversational AI operators to estimate user age and block explicit material for minors, which raises the bar for what counts as technically feasible protection.