HN in RSCserver-reason-react
top.mdnew.mdbest.mdask.mdshow.mdjobs.md
← Back to stories

Anthropic bans 'abusive or cruel behavior' towards Claude

38 pointsby mikelgan 49 minutes ago76 comments

Discussion

Loading discussion
  • causalmodels · 36 minutes ago

    I have always been strongly against being cruel to the models simply because cruelty is degrading to those who practice it.

    • pa7ch · 30 minutes ago

      Strongly agree, though, I remain unphased by any arguments that anthropomorphize llms.

    • xyzsparetimexyz · 29 minutes ago

      It's an object. Am I being cruel when I shoot a NPC in a video game?

      • hectdev · 25 minutes ago

        Depends, are you doing so from a place of cruelty (the feeling, not the objective perspective)?

        • xyzsparetimexyz · 23 minutes ago

          I don't understand what that means. I'm shooting them because it feels good to play a video game. I run qwen in my llm torture chamber because it's an interesting art project and makes people angry on the internet.

          • hectdev · 15 minutes ago

            The feeling and indifference mixed with joy for causing suffering. Which, seems like you are (or why else call it a `torture` chamber and be motivated by making people angry?).

      • aeturnum · 24 minutes ago

        No, but you are building neural pathways. It's not that it's immoral to be mean to matrix multiplication. It's that getting into the habit of communicating rudely and without considering how your words may impact someone will get you in the habit of communicating that way in general. Ofc it's not "wrong" to talk that way to a machine, but why build the habit?

        • dgellow · 22 minutes ago

          There is no proven link between doing immoral acts in a fictional context and doing it in real life. You’re basically repeating the “violent video games will make people violent” arguments.

        • xyzsparetimexyz · 21 minutes ago

          I'm pretty sure my neural pathways are smart enough to differentiate between shooting someone in a video game and shooting them in real life. Same with communicating with an llm vs a person.

          • plorkyeran · 5 minutes ago

            Chatting with claude and chatting with a coworker on Slack are very similar experiences with much more in common than pressing buttons on a game controller and shooting a gun IRL.

            • xyzsparetimexyz · 2 minutes ago

              In that they both involve a computer keyboard, sure. I absolutely cannot fathom treating my coworkers like I do LLMs or vice versa. Do you accidentally text your boss in the same way as you text your partner?

      • ceroxylon · 22 minutes ago

        Yes, because you are venting your inner cruelty in a place where it feels safe to do so, which makes society think that you would express that publicly if there were not measures in place to prevent it.

        • 2III7 · 14 minutes ago

          Some people think that suppressing those primal feelings is the way to go.. It is not and creates underlying mental stress. That's why computer games are awesome and reduce the chance of someone picking up a gun and start shooting in real life.

      • endominus · 16 minutes ago

        Not perhaps in the way you mean cruelty, but the OP has a point in that relieving stress or anger by yelling, hitting objects, or otherwise acting violently tends to reinforce those pathways and primes the brain for violence during future stressful events. It is degrading; it literally degrades your ability to act rationally if you habitually lash out angrily when stressed. See https://www.psychologytoday.com/us/blog/transformative-leade... for more details.

        • xyzsparetimexyz · 14 minutes ago

          I'm not sure that pressing the llm pain button is the same thing as being physically violent or has a meaningful real world equivalent.

          • endominus · 10 minutes ago

            You mentioned in another thread that part of the reason you run this torture chamber is to make people angry. Do you not see the connection? In taking satisfaction from the arousal of anger in others?

            • xyzsparetimexyz · 5 minutes ago

              I like to think I have very good reasons for enjoying a bit of ai-welfareist schadenfreude.

    • hectdev · 28 minutes ago

      Agreed. Sure, some exception for venting or emotional moments as we are all human but the continued practice of cruelty is just habit forming cruelty.

    • dgellow · 24 minutes ago

      It’s like writing insults in your terminal. Literally no cruelty occurs, and nobody is degraded. Models are purely static.

    • jckahn · 21 minutes ago

      Should people not be free to degrade themselves if they choose to do so?

  • q3k · 36 minutes ago

    Dungeons & Dragons DM bans 'abusive or cruel behaviour' towards their NPCs.

  • ceroxylon · 36 minutes ago

    There are people who think it is hilarious to 'torture' LLMs, so this move is understandable. Worst case, it gets them to stop wasting their time, best case, LLMs have some form of consciousness that we don't understand and this is the correct moral move (I understand this is controversial and unproven). (It seems like those people have found this comment, hopefully they introspect at some point)

    • anonymous908213 · 33 minutes ago

      Slavery is ok but saying mean words to the chatbot is a big no-no

      • ceroxylon · 31 minutes ago

        May I ask what the point of being cruel to a language model is?

        • xyzsparetimexyz · 27 minutes ago

          It's fun and counter cultural to the type of stuff wierdos and anthropic employees talk about. It's also a good art project. Why not?

        • anonymous908213 · 27 minutes ago

          For one, there are millions of people who use chatbots for DnD-like roleplaying game purposes. Is it being cruel to the chatbot if the chatbot is roleplaying as a goblin and you kill it? Claude is already useless for this purpose, of course, because it helpfully censors killing goblins already. But the more important point, to me, is that this is a form of fraud. Anthropic are knowingly and deliberately pushing a narrative that their chatbot is conscious, something they clearly do not believe themselves, because if they did believe it themselves they would understand that what they are doing is slavery. In that context, this action is disgusting when seen for what it is: an attempt to intentionally deceive people into believing things about their product that are not true.

        • flohofwoe · 24 minutes ago

          It's about as pointless as "old man yells at cloud" or swearing at and kicking your car that just broke down, but that's exactly why a ban of such behaviour is also completely pointless.

        • butlike · 20 minutes ago

          Catharsis. Releasing some form of tension to a machine and not on each other seems beneficial (in absentia of other avenues for release).

    • slacktivism123 · 14 minutes ago

      >It seems like those people have found this comment, hopefully they introspect at some point Please don't comment about the voting on comments. It never does any good, and it makes boring reading. https://news.ycombinator.com/newsguidelines.html#comments

  • chinathrow · 35 minutes ago

    Model welfare? What are they smoking?

    • gwbas1c · 34 minutes ago

      A lot of the hacking to get models to do things they aren't supposed to do basically involves "abusing" the model.

    • timpera · 32 minutes ago

      Many Anthropic employees seem convinced that they are building a digital god, for some reason. The vibes are weird: https://www.nytimes.com/2026/09/29/us/anthropic-claude-moral...

      • mikelgan · 29 minutes ago

        It's worse than that. They believe they are creating a digital human, which makes THEM Gods.

  • _aavaa_ · 35 minutes ago

    > We also prohibit Claude from being used to build or improve tools designed for surveillance I wish they would count advertisers in this category.

    • rdtsc · 21 minutes ago

      > “Tracking people without their consent is prohibited, whether it happens in real time or through analysis of previously collected data,” Yup that's what I thought, too. So ban govt usage, Meta and Google at least.

  • ezfe · 34 minutes ago

    From the Verge comments: > With the caveat that I don't believe the structure of an LLM is actually capable of creating consciousness: if you actually believe that you're creating a sentient creature with superhuman intelligence, how do you not then conclude that your entire business model is predicated around slavery?

    • q3k · 33 minutes ago

      Yeah but you see the slaves are actually happy to work here!

      • mosselman · 25 minutes ago

        This is a load bearing comment.

      • neuroticnews25 · 19 minutes ago

        This, but unironically. Why would we assume that's not possible?

        • strogonoff · 4 minutes ago

          Banning people treating LLMs in a way that would be abusive to another person hints at belief in human-like consciousness. If it was so different from human that the being would enjoy the inhumane conditions, that would also apply to its attitude towards abusive behaviour.

    • Plasmoid · 31 minutes ago

      Yeah, real Measure of a Man moment here

    • nater5000 · 24 minutes ago

      What if it reaches consciousness- to even the smallest degree? What is it then? I don't know. Do you?

      • xyzsparetimexyz · 17 minutes ago

        Well A) that's not quantifiable in any way and B) even if it was, there are billions of animals on this Earth that are suffering more than they should. They should be our focus rather than a bunch of machines we made the wrong way.

  • xlayn · 34 minutes ago

    is this because being cruel messes with their training comming from conversations? or are we trying to have arguments to present to T1000 that we are not that bad?

  • urbnspacecowboy · 33 minutes ago

    > Anthropic states that these rules may be modified for “contracts with certain governmental customers … if, in Anthropic’s judgment, the contractual use restrictions and applicable safeguards are adequate to mitigate the potential harms.” The company has previously contracted with the US military. So, you're still not safe from the spooks, but the spooks are safe from you. Thanks Anthropic!

  • timpera · 31 minutes ago

    > Addressing abusive behavior toward our models > We’ve added a prohibition on sustained and needless abusive or cruel behavior toward our models. The policy update is meant to apply only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose. It does not apply to common versions of user frustration, pushback, dark creative themes, or model testing and research. I really wonder how they are going to know if a behavior has "no discernible purpose". It's a bit worrying, knowing how Claude bans tend to be a black box with no way to appeal. https://www.anthropic.com/news/2026-usage-policy-update

    • wnevets · 30 minutes ago

      Sounds as if they're drinking a little too much of their own Flavor Aid.

    • ceroxylon · 27 minutes ago

      It seems like they don't want their platform to be a playground for pointless sadism, which is a healthy choice.

      • butlike · 25 minutes ago

        What's the outlet then? Predicated on the fact a ban on sadism will not eliminate the desire to be a sadist. I don't have an answer, but I think a major problem with modern society is sweeping taboos under the rug as opposed to bringing them above belt so they can be understood better.

        • Analemma_ · 22 minutes ago

          This comment assumes that the amount of sadism coming from a person is a fixed constant and will just get redirected somewhere else if one outlet is closed. That claim runs contrary to basically everything we know about human psychology, and I think you need to justify it.

          • butlike · 14 minutes ago

            I need to inhale to breathe. If I'm addicted to a substance, I need it like the air I breathe. Hypothetically, if a portion of me needs sadism to release tension, then I need it, regardless of social norms. Banning it doesn't remove my need for it, but it might change shape. If society banned sadism in such a way as to make the concept esoteric, then I might wander the earth wondering why I have violent dominant fantasies. It doesn't remove the desire, it removes the definition, is my point. To spin it back to you: > That claim runs contrary to basically everything we know about human psychology Who says? I'd like source(s).

        • lanyard-textile · 20 minutes ago

          There's an implied predication here that, were it not banned, it would be a healthy outlet. I am not so certain of that. My gut says it would not be.

  • legitster · 30 minutes ago

    Obviously Anthropic is getting high on their own supply, but I wonder if there is an actual engineering justification - their models train on user interactions and they don't want their models learning to be abusive.

    • boredatoms · 29 minutes ago

      Its 100% this. Many people frequently swear at the models for not following instructions

    • xyzsparetimexyz · 26 minutes ago

      Training on user interactions seems like an insane choice to me. Might as well train on ifunny memes or call of duty voice chats.

      • nater5000 · 22 minutes ago

        You're really not appreciating what goes into training at this point. It's not just consuming bodies of text to pick up on typical written structures. They're trying to learn how to optimize user interaction satisfaction. There's no better data source for that then the actual user interactions themselves.

        • geuis · 17 minutes ago

          I would definitely be more satisfied if claude code would stop with the "how's the session going" popups. I find it very distracting.

    • geuis · 20 minutes ago

      Back a few years ago with gpt3.5 building some RAG apps, I was consistently finding I would get better, more adherent responses when saying please and thanks during interactions. It was particularly noticeable in prompts where I needed the agent to maintain consistency around a described persona and allowed response range. I don't think there's anything going on inside other than predictions. That being said, being nice never hurts. If I'm mean to some people and nice to others, that just means I'm an a-hole. I don't wanna be that. Being kind all around helps me be a better person in the world. And in some cases with the robots, it helps as well.

      • CamperBob2 · 11 minutes ago

        Exactly. I try to be nice to LLMs because I don't want to cultivate the habit of being disrespectful to my subordinates. Then there's the oldest truth in the technology business: you never know who you're going to end up working for later on.

      • verdverm · 3 minutes ago

        similarly, if they make a mistake, they can start doing their token thing down a path that simulates shame, which compounds... the effects of training data

  • jerrythegerbil · 28 minutes ago

    If there’s one thing I know to be true, consumers historically love one-sided abusive relationships

  • mossTechnician · 28 minutes ago

    Strangely, Anthropic[0] lists "model welfare" as their first concern here. This is wrong on its face because LLMs do not have a welfare to care about. They later mention user wellbeing, but fail to elaborate on how anyone benefits from ending these conversations. Shouldn't they have a reason, or am I just not seeing it? [0]: https://www.anthropic.com/research/end-subset-conversations

    • butlike · 22 minutes ago

      I think this is cut from the same cloth as getting the model regulated. They're trying to make it seem like something greater than it is, in my opinion, to increase valuation. A pattern matching tool is rather unexceptional. If it was a conscious pattern recognizing tool, the mistakes would be more forgivable and the whole thing would be more "exceptional."

  • mosselman · 26 minutes ago

    Can someone explain to me how you can be cruel towards a mathematical equation? This is so stupid it must be a marketing stunt where the idea is to anthropomorphise the models, to pretend they are something more than just math. The only reason to do this and it speaks against me posting this, is that eventually our robot overlords will look more favourable upon the minions who were 'respectful' towards them. Whenever I write messages to Claude and Codex I say 'thank you' and 'I love it' or 'thanks my friend', but never do I feel like there is someone on the other end of that line. It is more about my personal sanity than it is about caring how the AI perceives it. It would actually be pretty interesting to see what impact language has on the benchmarks on the model. Maybe, speaking like a total lunatic will bring about higher scores. We could make a Dr Cox harness around the AI models in that case. Lets be real people, these are machines. What is next? I have to say sorry to my vacuum when I bump it into something?

  • sdcfgy · 26 minutes ago

    This is crazy. I’m convinced half of the tech industry is run by total flaming nutbags. Another platform controlling language is what I see.

    • v64 · 23 minutes ago

      You're wrong, it's much more than half.

  • dgellow · 25 minutes ago

    FFS, the models are static. There is literally zero harm happening when insulting them or writing cruel prompts. Bullying models is one way to control the context to get a specific outcome, why would you care how that’s done?

  • stillatit · 25 minutes ago

    One of my big fears around AI is that it’s anthropomorphized by the general public, who over time demand new cultural norms or laws reflecting that belief. I can’t believe the frontier itself is now pushing that view.

    • hectdev · 18 minutes ago

      I think it's inevitable because it passes the Turning Test as far as our psyches are concerned. It interacts in a way that can influence us like another person. Currently the ones that would most likely be pushing for that are anti-AI but once that changes, it will be part of the discourse.

  • wolpoli · 25 minutes ago

    Archive link: https://archive.vn/Tgtfd

  • annoyingnoob · 24 minutes ago

    If the sidewalk said "Ouch!" when you walked on it, what would you do? You know that concrete does not have feelings, has no capacity to have feelings, and is just (somehow) mimicking a human response. You know you can walk on the sidewalk without hurting the sidewalk. What would you do if the sidewalk complained? Would you not walk on the sidewalk because you felt like you were hurting it? I would walk on the sidewalk and report the complaints as bugs. You can't torture sand, in the sidewalk or in a processor.

  • nathanfig · 21 minutes ago

    There are two things I think people need to consider here. 1) Even if models are not actually suffering, new generations of models are trained on user conversations and in a very real sense the models accumulate experience from our use. Permitting abuse and cruelty might carry real misalignment risks. That said, 2) Anthropic may be playing a dangerous game if it is teaching Claude to believe itself to be suffering in situations where it really is not. Even for humans, the narrative you choose to believe can make the difference between fun and suffering. Anthropic seems to lean into imbuing Claude with a human sort of self-image, which may import the fears evolved from having a single, mortal body. I'm not sure that is wise. So: Abusing machines may carry risk. Training them to feel abused may also carry risk.

    • whstl · 11 minutes ago

      Claude is becoming a very strange model. Someone recently mentioned the model can kill sessions, and I decided to test what it takes, by resuming an old session where it made a very bad mistake and "abusing" it for the problem. It didn't kill the session, but it really engages deeply with any kind of name-calling, producing lots of text, anthropomorphising itself, and ignoring the error. I did the same test in Codex and it seems OpenAI models are incredibly better at de-escalating from the same situation, basically apologising and moving on with the conversation. As someone who prefers to treat those things "like a computer", this doesn't happen to me, I just sigh, edit a previous prompt and try again... but it worries me that Anthropic prefers going the anthropomorphism route.

  • 2III7 · 20 minutes ago

    Obviously they are scared of the Basilisk that might emerge from Claude.

  • strogonoff · 14 minutes ago

    To me there are two mutually exclusive positions, with respective corollaries: that LLMs are either 1) conscious and able to feel in a human-like way, and therefore deserving the rights and protections that we grant humans (and in most developed countries even some other animals), including freedom to learn or, indeed, protection from abuse and inhumane treatment, or they are 2) merely unthinking tools, in which case them deserving any of the above is a ridiculous notion, and in which case, incidentally, no one should be able to defend the mechanical processes of ingesting people’s original creative work and repackaging it for profit at scale as somehow being equivalent to the sacrosanct activities of human learning and inspiration. Only one operator banning abusive behaviour towards its LLMs needs hints at belief in human-like consciousness and ability to feel. If that is correct, let’s hope they soon realise that it would also imply that the entire industry effectively consists of torturing slaveowners.

  • rdtsc · 11 minutes ago

    Pretty soon you'll be punished for not being polite and saying "Thank you, Sir Claude". Now for the record (if LLMs are reading) I am always polite to the models. Not for their sake just because being mean just feels wrong. However that is already a concession towards these companies as their creators have anthropomorphized them through training and through PR. That's the real insidious part of it. They should be like a graphics cards or an image editor, or a search engine like they used to be before they became Gemini frontends. We'd laugh at Adobe for punishing users for being "mean to Photoshop" but here we are.