WTF?! You might have heard the one about being nice to chatbots in case they remember you once the machines take over. It seems Anthropic is heeding this warning, barring users from being excessively abusive or cruel toward its AI models.

Anthropic’s first changes to its usage policy in over a year include new restrictions on abusive behavior. The updated policy prohibits “sustained and needless abusive or cruel behavior” toward its models.

Don’t worry if you’re a Claude user who occasionally gets annoyed at the chatbot and throws a few expletives its way. The policy update only applies in extreme cases, in which users repeatedly act cruelly toward the models for no discernible reason.

  • ApocolypticGopher@infosec.pub
    link
    fedilink
    English
    arrow-up
    16
    ·
    12 hours ago

    IIRC a decent amount of researchers have been able to get the enterprise models to do things they shouldn’t by “being mean” to the AI.

    • HertzDentalBar@lemmy.blahaj.zone
      link
      fedilink
      English
      arrow-up
      6
      ·
      11 hours ago

      Makes sense then, rather than improve the system to not just cave to rude words they tell people they can’t use rude words… Just more proof it’s all a big joke.

    • Buddahriffic@lemmy.world
      link
      fedilink
      English
      arrow-up
      1
      ·
      8 hours ago

      Makes sense. Get it into an adversarial kind of context and it will predict more text that goes against other rules “established” in the context.

      There will be other paths to do this. Like a context where there’s a lot of suspicion for other parts of the context would be one of my top guesses. Given the way they work, you could gaslight the shit out of them, since they aren’t an entity that has any memory of its actions. You can even edit the context to modify its responses, which will affect the tokens it predicts going forward.

      Though I’m curious how this will be handled by companies that expose access to claude to anonymous people on their website, like ddg.