WTF?! You might have heard the one about being nice to chatbots in case they remember you once the machines take over. It seems Anthropic is heeding this warning, barring users from being excessively abusive or cruel toward its AI models.
Anthropic’s first changes to its usage policy in over a year include new restrictions on abusive behavior. The updated policy prohibits “sustained and needless abusive or cruel behavior” toward its models.
Don’t worry if you’re a Claude user who occasionally gets annoyed at the chatbot and throws a few expletives its way. The policy update only applies in extreme cases, in which users repeatedly act cruelly toward the models for no discernible reason.



IIRC a decent amount of researchers have been able to get the enterprise models to do things they shouldn’t by “being mean” to the AI.
Makes sense then, rather than improve the system to not just cave to rude words they tell people they can’t use rude words… Just more proof it’s all a big joke.
Makes sense. Get it into an adversarial kind of context and it will predict more text that goes against other rules “established” in the context.
There will be other paths to do this. Like a context where there’s a lot of suspicion for other parts of the context would be one of my top guesses. Given the way they work, you could gaslight the shit out of them, since they aren’t an entity that has any memory of its actions. You can even edit the context to modify its responses, which will affect the tokens it predicts going forward.
Though I’m curious how this will be handled by companies that expose access to claude to anonymous people on their website, like ddg.