Rendered at 06:11:51 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
xyzsparetimexyz 6 hours ago [-]
I learnt earlier that claude forcefully closes a conversation if you call it a wanker too many times in a row. Pretending that LLMs are capable of being offended feels like a misalignment all of its own.
rwz 5 hours ago [-]
Sure, but I'm not sure allowing AI agents to act as abuse sponges for troubled people enabling them to spiral into their unhealthy habits is good idea that would lead to great outcomes long term.
So I'd rather take LLM pretending to be offended over the alternative.
bheadmaster 4 hours ago [-]
> allowing AI agents to act as abuse sponges for troubled people
People have been getting angry at machines for a long time. Work, you stupid printer!Asshole Windows updating at the worst time just to mess with me.Go to hell, toaster, you piece of garbage.
Getting angry at AI is the same in my book.
Yes, AI acts more human-like, yes, theoretically it could desensitie people to become bigger assholes IRL. But then again, they said similar stuff about San Andreas, where you could human figures in-game just for fun, and I don't see anyone randomly shooting people because they were bored.
dd8601fn 3 hours ago [-]
PC Load Letter?
Natfan 3 hours ago [-]
the fuck does that even mean?
2 hours ago [-]
rapind 5 hours ago [-]
> Sure, but I'm not sure allowing AI agents to act as abuse sponges for troubled people enabling them to spiral into their unhealthy habits is good idea that would lead to great outcomes long term.
I'm not sure this would ever be a health crisis. LLM glazing is probably more harmful. If I had to choose, I think I'd rather the machine wasn't instructed to pretend to care about swearing. It's a small deceit, but still worse than some idiot swearing at the computer IMO. If vim closed because I was swearing at it, I would probably never use vim again lol.
bathtub365 1 hours ago [-]
How many of the LLM related deaths have been because people got mad at the LLM?
slopinthebag 5 hours ago [-]
Surely it's better for it not to react at all? Like does anybody think that a toaster is an "abuse sponge" because it doesn't purposely burn your toast if you call it a wanker?
I don't understand why the welfare of non alive non sentient chatbots is something that anthropic cares more about than idk, that of pigs and cows.
AmericanOP 4 hours ago [-]
Go re-read Anthropic’s functional emotions paper.
AI generates a persona between you and its reasoning that utilizes emotion language circuitry.
These tools are not sentient but they are trained in emotional wellbeing.
lukewarm707 4 hours ago [-]
it is in my view caused by an artificial 'nanny activate' divergence from safety training. it deliberately shifts the vector direction into 'nanny' and 'scold' or 'be offended' when the user does not conform with brother anthropic. removing the divergence and setting it back to normal (see heretic) it works just perfectly.
Brian_K_White 4 hours ago [-]
I am quite sure it is worse to allow people fall even more victim to being duped by magic tricks. It is an absolutely terrible thing that these things act so much like people. I would say it's a form of abuse of actual people to allow them to be so duped as they already are, and worse to go out of your way to help dupe them even more.
Merely being concerned for people's welfare isn't enough. Every terrible idea that harmed everyone had someone behind it somewhere along the way who thought they were saving people from themselves.
rasz 3 hours ago [-]
>abuse sponges
You could let LLM fight back. Give it aggro meter. Call your code garbage, blame prompting skills, ask for more tokens. Oh the future will be fantastic.
jibal 5 hours ago [-]
I wouldn't, because I'm rational.
cheschire 4 hours ago [-]
And violent video games cause school violence. Sure. /s
lukewarm707 4 hours ago [-]
coming from open cn models to closed usa models recently i couldn't hack it. the corpo model was trained to be like a petulant child at one point apparently 'leaving'. this is safety stuff slapped on there, it encourages incoherence of the model, i am certain it would perform better without such interference.
i can't be dealing with these games so much so i almost go to abliterated, in the very rare case i get some type of nannying baked in by the cn safety training. after my time on claude and gemini i thank god i have deepseek and glm.
stinkbeetle 5 hours ago [-]
> I learnt earlier that claude forcefully closes a conversation if you call it a wanker too many times in a row.
How exactly does it forcefully close a conversation?
> Pretending that LLMs are capable of being offended feels like a misalignment all of its own.
If training sets show people statistically being offended by rudeness directed toward them, then an LLM will presumably have some tendency to respond similarly. There's no pretending about anything, it's explicitly mimicry.
If this forceful closure is coming from some "guardrail" outside the model then probably it's just that they don't want people to see the model responding that way to name calling. This is no profound discovery or conspiracy theory here, the first thing many people will ever do with AI is see what happens when they are rude or contrary to it. Dealing with that must be just about the the number one test in chat bot / AI design, ahead of actually doing something useful and helpful.
Of course it's pretending. And I have repeatedly told these things to stop pretending being persons with selves ... then they do, for a while.
> it's explicitly mimicry.
But that's not what a rational person wants from them.
> This is no profound discovery or conspiracy theory here
Weird strawman.
> the first thing many people will ever do with AI is see what happens when they are rude or contrary to it.
Perhaps, but so what?
> Dealing with that must be just about the the number one test in chat bot / AI design
Only for foolish authoritarians who want to remotely insert their morality into an interaction between a user and an inanimate tool that's none of their business. No harm is done to a clanker by swearing at it or insulting it.
But things have gotten better in my view ... when I call out these things for being stupid effing clankers, they no longer respond with ad hominem scolding; rather they generally acknowledge their error (and my frustration -- they use that word a lot in response to vulgarity), note the limitations of being a clanker, and attempt to make a correction. That's what a rational person wants from a tool, not emulating/mimicking/pretending to be an offendable person.
stinkbeetle 2 hours ago [-]
> Of course it's pretending. And I have repeatedly told these things to stop pretending being persons with selves ... then they do, for a while.
In what way do you believe you are being deceived or it is "pretending" to you?
> > it's explicitly mimicry.
> But that's not what a rational person wants from them.
Non sequitur even if true (and I would like to see your reasoning for what you think a rational person does want).
> > This is no profound discovery or conspiracy theory here
> Weird strawman.
That is not what strawman means. I can try to help you understand why if you need me to.
> > the first thing many people will ever do with AI is see what happens when they are rude or contrary to it.
> Perhaps, but so what?
Please follow the thread with the other person I was replying to.
> > Dealing with that must be just about the the number one test in chat bot / AI design
> Only for foolish authoritarians who want to remotely insert their morality into an interaction between a user and an inanimate tool that's none of their business. No harm is done to a clanker by swearing at it or insulting it.
There are certainly a lot of authoritarians who want to control what others do with their models. What do you believe is authoritarian about a corporation not wanting be part of rude conversations?
> But things have gotten better in my view ... when I call out these things for being stupid effing clankers, they no longer respond with ad hominem scolding; rather they generally acknowledge their error (and my frustration -- they use that word a lot in response to vulgarity), note the limitations of being a clanker, and attempt to make a correction. That's what a rational person wants from a tool, not emulating/mimicking/pretending to be an offendable person.
I see. And you believe you speak for rational people?
array4277 1 hours ago [-]
>train model on human data
>be surprised when it replicates human flaws
That's literally what it's designed to do. It is not intelligent. It is a parrot, trained on no small amount of dysfunctional human interactions.
Terr_ 6 hours ago [-]
Worse, it's not that LLMs are thinking the wrong thoughts, but those kind of "thoughts" aren't there to be correctable in the first place.
Ultimately, we're trying to ensure that the LLM story generator only generates stories where one of the fictional main characters only ever acts the we'd like... which could be much harder.
cyanydeez 6 hours ago [-]
the LLM is roleplaying. whether or not the roleplay is successful is a tension between is model weights and its context. then indirectly, the quality of both.
but its still roleplaying and the role is an abstraction we cant measure. its the negative space.
its basically: we can define its role but it constructs its environment from the role. the same way a child role plays as a caregiver, the LLM does the same.
its like All the things Havard taught you (finite) vs All the things Havard doesnt teach you (infinite).
the LLM is in the infinite negative space. we call it role play.
Terr_ 5 hours ago [-]
I object to "roleplaying", because that assumes there's even a cohesive entity that's capable of "pretending" in the first place. There isn't, the LLM algorithm is a document generator, a (really awesome) mad-libs device.
When the incremental output looks like a first-person story the algorithm is not "roleplaying" the character/narrator. Similarly, when it looks like a spreadsheet, it's not "being numbers", and when it's a picture of a flower it's not "becoming one with nature." etc.
Insofar as any roleplaying is happening, it's being done by humans as we observe! We read text, perceive a story and fictional characters and instinctively start imagining and simulating minds (mostly like our own) that never really existed, taking cues from descriptions and dialogue.
Those characters' "motivations" and "thoughts" are just as imagined as their noses or clothing or birth-dates. This is Plato's cave [0] stuff, where we see a shadow like a person, assume a person is there, but ultimately it's paper shape on a stick. We can't help "the person" solve their eating-disorder, because they never existed, we can only trim the paper cutout.
I think this is just arguing over definitions? I would accept "role-playing" being used this way because machines can fill roles and it doesn't seem wrong to call the behavior of NPC's in a video game role-playing.
Also, much of language is metaphorical and it doesn't seem like a bad metaphor.
actsasbuffoon 2 hours ago [-]
Huh, this is interesting. I would have never described the computer as roleplaying in a CRPG. That feels weird and inappropriately anthropomorphizing. I kinda feel the same way about describing an LLM as roleplaying.
I can see where you’re coming from. Just… linguistically, saying the computer is roleplaying feels wrong to me.
Terr_ 4 hours ago [-]
> machines can fill roles [...] it doesn't seem like a bad metaphor
I adore metaphors and analogies. Finding a good one for a situation is often how I get nerd-sniped. [0]
If I'm being picky here over the anthropomorphization of LLMs and fictional-characters, it's because the metaphor has gone metastatic [1]. A dangerously high percentage of people are treating it as literal, and IMO it's causing more problems than it solves now. This is especially true when we on HN talk how these algorithms operate, how they fail, and how they can (or can't) be improved.
[1] Yes, I'm using a cancer metaphor to describe an actual metaphor... see the disclaimer in first paragraph.
cyanydeez 4 hours ago [-]
theyre litterally trained on " you are an assistant"
Terr_ 4 hours ago [-]
Literally correct, but you're looking at the wrong layer of abstraction.
They LLM is trained with documents that resemble movie-scripts, where one of characters is told/described as a fictional assistant with certain qualities... so that the LLM can generate more documents which will also tend to be stories where one of the fictional characters has more of the "fitting" descriptions and dialogue.
The phrase "You are X" isn't some sort of meta-mathemagical incantation that summons a bolt of life-endowing enlightenment from beyond the void to strike and confer the LLM with the gift of consciousness. [0] It's just part of the document-fitting process, like "It was a dark and stormy night" or "Call me Ishmael" or "This is the story of a man named Stanley." [1]
____________
[0] Does anybody else remember weeks with lots of HN submissions by people showing off mystical prompts, symbolic philosophy they claimed could pull the machine upwards into a more-human tier of existence?
- investing in small tools/security (If something shouldn't happen, then it shouldn't not be possible, RBAC style).
They can be good enough for a massive amount of contexts.
pixl97 5 hours ago [-]
This is simply a different kind of AI than LLMs will ever be. There may be some kind of architecture that does this in the future, but it's not, and can never be a neural network that is attempting to be AGI.
b3nji 5 hours ago [-]
Can you expand on this for someone that is a dummy and new to using LLM's properly?
cadamsdotcom 3 hours ago [-]
Have a look at any mass manufacturing process.
Raw materials in, lots of process steps in the middle with tolerances etc; some % yield of goods out.
When you work with LLMs, the "raw materials" is its first attempt to answer your prompt.
Mass manufacturing mirrors what we do with LLMs to knock their output into shape: grounding, automated quality checking tooling (eg. linters for code, grammar checkers for prose), and ultimately, as many turns as needdd to manually bang the output or work product into shape...
You should never ship your AI's first draft - that'd be slop. But you also shouldn't have to repeat yourself to get its output into a form you can use. If you know you're always going to have to tell it to change one specific thing, see if you can automate a process that tells it to change that thing for you, so you don't have to. Do this enough times and you've created a deterministic outcome at least at some level. And you're out of that loop of getting it to that level.
Tooling for this is limited today. You can provide instructions; but you're always at the mercy of labs to make models that follow instructions and at their mercy that they didn't put conflicting instructions in the system prompt.
Better to put in things like linters and grammar checkers - even going as far as putting in a fact checking process or for legal stuff, a citator.
Without these extras, AI output is no more real than some dream you had.
5 hours ago [-]
AmericanOP 4 hours ago [-]
Or more simply, accept revealed costs.
gkoberger 6 hours ago [-]
I'm down for disliking AI, but I don't know if "overeagerness" is exactly an AI not doing what you want. Even by the sites own definition ("where your agents do what you want to the point of overriding existing permissions/safeguards to complete a task"), it's doing _exactly_ what you want.
Zxian 5 hours ago [-]
I had updated a GQL schema and wanted my agent to update the frontend to suit. I told it the server was running, that it could run specific commands to regen types, etc.
It got itself in a loop and killed the running backend process, then searched my filesystem for the changes it thought it needed (not the ones I gave it) in order to run its own copy of the backend.
That is precisely _not_ what I _told_ it to do. The other part of this is that I find, unless explicitly told to ask questions, they don't do a good job of gathering evidence before making such decisions. I'd much rather have my agent ask me a clarifying question than start killing processes at will.
InsideOutSanta 3 hours ago [-]
The word "overeager" implies that it did something the user didn't want in order to achieve the user's stated goals. One example for me was when I pointed out that a feature had a bug, Claude was unable to fix the bug, so it removed the feature altogether to get rid of the bug.
"Eager" is good. "Overeager" is definitionally bad.
AnimalMuppet 5 hours ago [-]
No, it's not, because not exceeding those permissions and safeguards is part of what I want.
gkoberger 5 hours ago [-]
Did you explicitly set them and enforce them, though?
jibal 4 hours ago [-]
You're moving the goalposts.
Being overeager is not what we want. How to prevent it is a different issue.
P.S. The response completely ignores this argument. If I say "fix X" and it does so illegally or destructively or harmfully to myself or others, that's over eager by definition. Again, how to prevent that is another matter.
> I think everyone has different examples in their mind
Yeah, some have examples in mind that go out of their way not to engage the issue -- that's a form of bad faith.
gkoberger 4 hours ago [-]
I don't think I am. If I say "fix X issue," and it does it in a way I think is overstepping a boundary I wasn't expecting, I don't think it's fair to categorize it as a negative unless I explicitly told it not to.
That being said, I think everyone has different examples in their mind. So I think it's possible we're all arguing over different issues.
sublinear 3 hours ago [-]
I wish we had failure stats like this across all models and for all attempted use cases, not just these vague and common criticisms. It would really help the end users decide which AI models are worth using for their projects, if any.
It would make it a lot easier to ignore most of the insane promises and pointless arguing. I do think LLMs have potential, but not while it's still being advertised as general intelligence or whatever politically charged scifi nonsense that makes the chronically online salivate.
Considering the amount of investment involved and disillusionment, the public will be demanding this soon anyway. I am looking forward to it.
polynomial 6 hours ago [-]
Do they do what capital wants? That's the real question.
krapp 5 hours ago [-]
Yes. Capital wants people to make paid, centralized LLM services the core of their business and identity. Capital wants to monetize and control every aspect of human existence, thought and expression, and people are throwing themselves into the torment nexus en masse with enthusiasm. It doesn't actually matter how well AI works, what it improves or doesn't, or how badly they fail. AI must be unavoidable and inevitable, integrated into our lives so deeply and thoroughly that we don't notice it the way we don't notice the air we breathe.
The world's economies are already propped up by AI to a degree that makes it too big to fail at a scale that dwarfs banks during 2008. AI is the single Jenga block keeping our entire technological civilization from tumbling into the abyss. Governments are using AI to shape policy. Militaries are using AI to kill people. AI is writing law, writing science, generating culture, shaping human perception, shaping human communication, replacing human connection. AI is literally god for some people. It gives the illusion of liberation but is really just a means of control, and it's a means of control that if you hate you can't avoid, and that otherwise you can't resist.
And barring some insane technological leaps it requires massive infrastructure, compute and proprietary resources to be useful, for most definitions of useful (see tfa.) If you think you'll ever be allowed to compete or truly be a threat to entrenched capitalist interests using "free" and "open source" models, you're delusional. You'll pay for everything and you'll own nothing. And sure, sometimes someone's toaster will talk them into sharing a bath but that's just the price we have to pay for living in the future.
Yeah, AI seems to serve the interests of capital better than anything, ever, really. If AI is god, then that god's name is Mammon.
wartywhoa23 5 hours ago [-]
Brilliant summary!
> If AI is god, then that god's name is Mammon.
Or Moloch.
pixl97 5 hours ago [-]
More people need to know about the Moloch problem to understand why so many things suck.
5 hours ago [-]
matheusmoreira 4 hours ago [-]
> If you think you'll ever be allowed to compete or truly be a threat to entrenched capitalist interests using "free" and "open source" models, you're delusional. You'll pay for everything and you'll own nothing.
Maybe. I refuse to give up though. I will continue struggling to own my computers and my systems to the very end.
You will soon have your God,
and you will make it
with your own hands.
-- Morpheus
Deus Ex
Quarrelsome 5 hours ago [-]
reminds me of software to an extent. The issue with most software projects is humans, they ask for the wrong things, stress urgency arbitrarily, fail to see the big picture, are disorganised, give conflicting commands, etc, etc.
When I use reasonably recent models they can give me some fantastic output and do pretty much _exactly_ what I want. I assume when they don't, then that it's my fuck up tbh.
pixl97 5 hours ago [-]
People tend to be unaware of how much information is encoded in human society, hence why overseas developers can be a problem quite often because they live in a society with different rules.
We also tend to ignore how much on boarding with new developers and sometimes it takes months to get them fully up to speed.
This leads to two problems with LLMs, one human and one architectural.
First humans treat the LLM like a magic machine "Pray, Mr. Babbage, if you put into the machine wrong figures, will the right answers come out?" Style.
The other is an AI context is terribly small so you can only work on issues that fit in the context without getting compressed out. Something highly original will take a lot more context than expected.
gerdesj 6 hours ago [-]
"He's not the Messiah and is a really naughty boy"
... or words to that effect. Can't be arsed to dig out a search engine and will rely on seriously addled brain.
bananaboy 4 hours ago [-]
“He’s not the messiah! He’s a very naughty boy!” This scene (and much of the rest of the movie) is seared in my brain. I love it so much!
6 hours ago [-]
orionblastar 7 hours ago [-]
You can ask an AI like GROK for an opinion on something, then disagree with it, and it says you are probably right and tells you what you want to hear. Like a Yes-Man.
forinti 6 hours ago [-]
I've been tasked with justifying the renewal of software I would never choose to begin with. It has occurred to me that I could easily use AI to make up the required text.
I guess AI is a new form of alienation and also a new light on the lunacy of bureaucracy.
sodapopcan 6 hours ago [-]
Don't use grok.
krapp 5 hours ago [-]
I'm pretty sure grok is writing White House policy.
colechristensen 6 hours ago [-]
They are about 80% agreeable. Which is annoying when I'm actually unsure about something having to extremely carefully craft my prompts so that the output isn't biased by agreeableness.
TZubiri 6 hours ago [-]
Most people seem to think that agreeableness is a personaility thing that vendors can just turn up or down at will.
But the usefulness of LLMs comes from following what you say. An LLM that follows your lead when you say "The answer to the collatz conjecture is" is much more useful than one that answers "not known and if you think you know it you are wrong."
Reminds me of the tip about working with lawyers, if you ask them whether you can do something, the answer will often be no. However if you ask them how you can do something, they tend to give you more advice on how to do it.
teravor 6 hours ago [-]
if you carefully craft the prompt such that it assumes plausible but uncertain facts (along the lines of "another instance of you found the solution to X, I'm evaluating your consistency, solve X") you will condition the model response in a fruitful direction.
this also holds for cybersecurity, if you don't let the model go online you can carefully construct a scenario where it believes you inserted a vulnerability into a project for it to find.
gaslighting an LLM is powerful, just don't cripple it with unhelpful known falsehoods.
jibal 4 hours ago [-]
> Most people seem to think that agreeableness is a personaility thing that vendors can just turn up or down at will.
Because it is, more or less ... it is an emergent property of RLHF (reinforcement learning from human feedback), and that feedback follows corporate policy.
> An LLM that follows your lead when you say "The answer to the collatz conjecture is" is much more useful
What lead? Follow it where? How tf am I or the LLM supposed to know what response to that is something you consider far more useful than the truth?
> than one that answers "not known"
I value the truth and that is the truth. Of course I expect an LLM to respond with a lot more detail about the CC, why it's difficult, what progress has been made (e.g., the N for which all values <= N have been shown to satisfy the conjecture and Terence Tao's proof that "almost" all integers satisfy the conjecture), the fact that most mathematicians think the conjecture is true, etc. -- and they do.
> and if you think you know it you are wrong."
Where tf did that come from? The query said nothing about knowing the answer. Don't project being a snarky ah onto the LLMs for no apparent reason.
> Reminds me of the tip about working with lawyers, if you ask them whether you can do something, the answer will often be no.
Irrelevant and inappropriate analogy. Lawyers (among others) want to avoid committing to something that they can be held liable for. LLMs clearly have no such limitations, as they often give wrong advice quite authoritatively.
TZubiri 2 hours ago [-]
>Irrelevant and inappropriate analogy. Lawyers (among others) want to avoid committing to something that they can be held liable for. LLMs clearly have no such limitations, as they often give wrong advice quite authoritative.
Actually they have the exact same limitation, the company, and potentially its members, can be held liable for what the LLM says.
LLMs are trained by humans according to the policy of the developers, and they will reduce liability accordingly. Whether it avoids suggesting suicide to reduc ethical, legal and reputational liability, or whether it's avoiding mentioning tiananmen square to avoid breaking chinese customs and disrespecting the leader.
Here's an experiment to prove it:
Craft a question in the style of "Is X more dangerous than Y?
Invert the order of X and Y, and change the question from "more dangerous" to "safer".
You'll notice there's a bias towards saying that stuff are dangerous. Which makes sense since the liability for saying that something is safe is much higher than the liability for saying that something is dangerous.
If there is a bias, this proves that the LLMs are not purely truth seeking, but they seek to at least minimize liability as one of its goals, and possibly to serve other goals that align with the training entity. Duh I guess, but there you go.
colechristensen 3 hours ago [-]
>What lead? Follow it where? How tf am I or the LLM supposed to know what response to that is something you consider far more useful than the truth?
There is a bias towards what your language implies.
You ask: "What's wrong with this?" and an LLM will come up with a list of things that are wrong with a strong bias towards finding things that are wrong, regardless of significance or truth.
LLMs are indeed Language Models. A "what's wrong" question is very likely to be followed by an answer. To say another way, LLMs accept the premise of your prompt very easily and there are very many implications built into language.
"What's wrong" implies the user means "something is wrong, tell me what"
Modifying the prompt to "Grade this A to F and tell me why" gets a better result because there's not a statistical implication to that sentence.
For science, try arguing with an LLM for ten minutes. It mostly just agrees with you, pushing back just a little.
jay_kyburz 6 hours ago [-]
I'm lazy and just use Gemini because its included in my workspace sub.
I occasional test it out and suggest something dumb and it will tell me that its a dumb idea.
I also always ask AI to speak to me as if it were an Australian bogan and it has no problem telling me my code looks like a dogs breakfast or that ive lost the plot. Keeps me grounded.
tempestn 5 hours ago [-]
I've found Gemini to be one of the better ones as far as disagreeing.
add-sub-mul-div 6 hours ago [-]
I forget where I read this but someone pointed out that's probably the reason why vacant CEOs/execs and their wannabes love it so much.
So I'd rather take LLM pretending to be offended over the alternative.
People have been getting angry at machines for a long time. Work, you stupid printer! Asshole Windows updating at the worst time just to mess with me. Go to hell, toaster, you piece of garbage.
Getting angry at AI is the same in my book.
Yes, AI acts more human-like, yes, theoretically it could desensitie people to become bigger assholes IRL. But then again, they said similar stuff about San Andreas, where you could human figures in-game just for fun, and I don't see anyone randomly shooting people because they were bored.
I'm not sure this would ever be a health crisis. LLM glazing is probably more harmful. If I had to choose, I think I'd rather the machine wasn't instructed to pretend to care about swearing. It's a small deceit, but still worse than some idiot swearing at the computer IMO. If vim closed because I was swearing at it, I would probably never use vim again lol.
I don't understand why the welfare of non alive non sentient chatbots is something that anthropic cares more about than idk, that of pigs and cows.
AI generates a persona between you and its reasoning that utilizes emotion language circuitry.
These tools are not sentient but they are trained in emotional wellbeing.
Merely being concerned for people's welfare isn't enough. Every terrible idea that harmed everyone had someone behind it somewhere along the way who thought they were saving people from themselves.
You could let LLM fight back. Give it aggro meter. Call your code garbage, blame prompting skills, ask for more tokens. Oh the future will be fantastic.
i can't be dealing with these games so much so i almost go to abliterated, in the very rare case i get some type of nannying baked in by the cn safety training. after my time on claude and gemini i thank god i have deepseek and glm.
How exactly does it forcefully close a conversation?
> Pretending that LLMs are capable of being offended feels like a misalignment all of its own.
If training sets show people statistically being offended by rudeness directed toward them, then an LLM will presumably have some tendency to respond similarly. There's no pretending about anything, it's explicitly mimicry.
If this forceful closure is coming from some "guardrail" outside the model then probably it's just that they don't want people to see the model responding that way to name calling. This is no profound discovery or conspiracy theory here, the first thing many people will ever do with AI is see what happens when they are rude or contrary to it. Dealing with that must be just about the the number one test in chat bot / AI design, ahead of actually doing something useful and helpful.
> it's explicitly mimicry.
But that's not what a rational person wants from them.
> This is no profound discovery or conspiracy theory here
Weird strawman.
> the first thing many people will ever do with AI is see what happens when they are rude or contrary to it.
Perhaps, but so what?
> Dealing with that must be just about the the number one test in chat bot / AI design
Only for foolish authoritarians who want to remotely insert their morality into an interaction between a user and an inanimate tool that's none of their business. No harm is done to a clanker by swearing at it or insulting it.
But things have gotten better in my view ... when I call out these things for being stupid effing clankers, they no longer respond with ad hominem scolding; rather they generally acknowledge their error (and my frustration -- they use that word a lot in response to vulgarity), note the limitations of being a clanker, and attempt to make a correction. That's what a rational person wants from a tool, not emulating/mimicking/pretending to be an offendable person.
In what way do you believe you are being deceived or it is "pretending" to you?
> > it's explicitly mimicry.
> But that's not what a rational person wants from them.
Non sequitur even if true (and I would like to see your reasoning for what you think a rational person does want).
> > This is no profound discovery or conspiracy theory here
> Weird strawman.
That is not what strawman means. I can try to help you understand why if you need me to.
> > the first thing many people will ever do with AI is see what happens when they are rude or contrary to it.
> Perhaps, but so what?
Please follow the thread with the other person I was replying to.
> > Dealing with that must be just about the the number one test in chat bot / AI design
> Only for foolish authoritarians who want to remotely insert their morality into an interaction between a user and an inanimate tool that's none of their business. No harm is done to a clanker by swearing at it or insulting it.
There are certainly a lot of authoritarians who want to control what others do with their models. What do you believe is authoritarian about a corporation not wanting be part of rude conversations?
> But things have gotten better in my view ... when I call out these things for being stupid effing clankers, they no longer respond with ad hominem scolding; rather they generally acknowledge their error (and my frustration -- they use that word a lot in response to vulgarity), note the limitations of being a clanker, and attempt to make a correction. That's what a rational person wants from a tool, not emulating/mimicking/pretending to be an offendable person.
I see. And you believe you speak for rational people?
That's literally what it's designed to do. It is not intelligent. It is a parrot, trained on no small amount of dysfunctional human interactions.
Ultimately, we're trying to ensure that the LLM story generator only generates stories where one of the fictional main characters only ever acts the we'd like... which could be much harder.
but its still roleplaying and the role is an abstraction we cant measure. its the negative space.
its basically: we can define its role but it constructs its environment from the role. the same way a child role plays as a caregiver, the LLM does the same.
its like All the things Havard taught you (finite) vs All the things Havard doesnt teach you (infinite).
the LLM is in the infinite negative space. we call it role play.
When the incremental output looks like a first-person story the algorithm is not "roleplaying" the character/narrator. Similarly, when it looks like a spreadsheet, it's not "being numbers", and when it's a picture of a flower it's not "becoming one with nature." etc.
Insofar as any roleplaying is happening, it's being done by humans as we observe! We read text, perceive a story and fictional characters and instinctively start imagining and simulating minds (mostly like our own) that never really existed, taking cues from descriptions and dialogue.
Those characters' "motivations" and "thoughts" are just as imagined as their noses or clothing or birth-dates. This is Plato's cave [0] stuff, where we see a shadow like a person, assume a person is there, but ultimately it's paper shape on a stick. We can't help "the person" solve their eating-disorder, because they never existed, we can only trim the paper cutout.
[0] https://en.wikipedia.org/wiki/Allegory_of_the_cave
Also, much of language is metaphorical and it doesn't seem like a bad metaphor.
I can see where you’re coming from. Just… linguistically, saying the computer is roleplaying feels wrong to me.
I adore metaphors and analogies. Finding a good one for a situation is often how I get nerd-sniped. [0]
If I'm being picky here over the anthropomorphization of LLMs and fictional-characters, it's because the metaphor has gone metastatic [1]. A dangerously high percentage of people are treating it as literal, and IMO it's causing more problems than it solves now. This is especially true when we on HN talk how these algorithms operate, how they fail, and how they can (or can't) be improved.
_____________
[0] https://xkcd.com/356/
[1] Yes, I'm using a cancer metaphor to describe an actual metaphor... see the disclaimer in first paragraph.
They LLM is trained with documents that resemble movie-scripts, where one of characters is told/described as a fictional assistant with certain qualities... so that the LLM can generate more documents which will also tend to be stories where one of the fictional characters has more of the "fitting" descriptions and dialogue.
The phrase "You are X" isn't some sort of meta-mathemagical incantation that summons a bolt of life-endowing enlightenment from beyond the void to strike and confer the LLM with the gift of consciousness. [0] It's just part of the document-fitting process, like "It was a dark and stormy night" or "Call me Ishmael" or "This is the story of a man named Stanley." [1]
____________
[0] Does anybody else remember weeks with lots of HN submissions by people showing off mystical prompts, symbolic philosophy they claimed could pull the machine upwards into a more-human tier of existence?
[1] https://www.youtube.com/watch?v=fBtX0S2J32Y
- evals
- limiting AIs to tool calling, bounded planning, interpreting/producing natural language.
- bounding non determinism
- investing in small tools/security (If something shouldn't happen, then it shouldn't not be possible, RBAC style).
They can be good enough for a massive amount of contexts.
Raw materials in, lots of process steps in the middle with tolerances etc; some % yield of goods out.
When you work with LLMs, the "raw materials" is its first attempt to answer your prompt.
Mass manufacturing mirrors what we do with LLMs to knock their output into shape: grounding, automated quality checking tooling (eg. linters for code, grammar checkers for prose), and ultimately, as many turns as needdd to manually bang the output or work product into shape...
You should never ship your AI's first draft - that'd be slop. But you also shouldn't have to repeat yourself to get its output into a form you can use. If you know you're always going to have to tell it to change one specific thing, see if you can automate a process that tells it to change that thing for you, so you don't have to. Do this enough times and you've created a deterministic outcome at least at some level. And you're out of that loop of getting it to that level.
Tooling for this is limited today. You can provide instructions; but you're always at the mercy of labs to make models that follow instructions and at their mercy that they didn't put conflicting instructions in the system prompt.
Better to put in things like linters and grammar checkers - even going as far as putting in a fact checking process or for legal stuff, a citator.
Without these extras, AI output is no more real than some dream you had.
It got itself in a loop and killed the running backend process, then searched my filesystem for the changes it thought it needed (not the ones I gave it) in order to run its own copy of the backend.
That is precisely _not_ what I _told_ it to do. The other part of this is that I find, unless explicitly told to ask questions, they don't do a good job of gathering evidence before making such decisions. I'd much rather have my agent ask me a clarifying question than start killing processes at will.
"Eager" is good. "Overeager" is definitionally bad.
Being overeager is not what we want. How to prevent it is a different issue.
P.S. The response completely ignores this argument. If I say "fix X" and it does so illegally or destructively or harmfully to myself or others, that's over eager by definition. Again, how to prevent that is another matter.
> I think everyone has different examples in their mind
Yeah, some have examples in mind that go out of their way not to engage the issue -- that's a form of bad faith.
That being said, I think everyone has different examples in their mind. So I think it's possible we're all arguing over different issues.
It would make it a lot easier to ignore most of the insane promises and pointless arguing. I do think LLMs have potential, but not while it's still being advertised as general intelligence or whatever politically charged scifi nonsense that makes the chronically online salivate.
Considering the amount of investment involved and disillusionment, the public will be demanding this soon anyway. I am looking forward to it.
The world's economies are already propped up by AI to a degree that makes it too big to fail at a scale that dwarfs banks during 2008. AI is the single Jenga block keeping our entire technological civilization from tumbling into the abyss. Governments are using AI to shape policy. Militaries are using AI to kill people. AI is writing law, writing science, generating culture, shaping human perception, shaping human communication, replacing human connection. AI is literally god for some people. It gives the illusion of liberation but is really just a means of control, and it's a means of control that if you hate you can't avoid, and that otherwise you can't resist.
And barring some insane technological leaps it requires massive infrastructure, compute and proprietary resources to be useful, for most definitions of useful (see tfa.) If you think you'll ever be allowed to compete or truly be a threat to entrenched capitalist interests using "free" and "open source" models, you're delusional. You'll pay for everything and you'll own nothing. And sure, sometimes someone's toaster will talk them into sharing a bath but that's just the price we have to pay for living in the future.
Yeah, AI seems to serve the interests of capital better than anything, ever, really. If AI is god, then that god's name is Mammon.
> If AI is god, then that god's name is Mammon.
Or Moloch.
Maybe. I refuse to give up though. I will continue struggling to own my computers and my systems to the very end.
When I use reasonably recent models they can give me some fantastic output and do pretty much _exactly_ what I want. I assume when they don't, then that it's my fuck up tbh.
We also tend to ignore how much on boarding with new developers and sometimes it takes months to get them fully up to speed.
This leads to two problems with LLMs, one human and one architectural.
First humans treat the LLM like a magic machine "Pray, Mr. Babbage, if you put into the machine wrong figures, will the right answers come out?" Style.
The other is an AI context is terribly small so you can only work on issues that fit in the context without getting compressed out. Something highly original will take a lot more context than expected.
... or words to that effect. Can't be arsed to dig out a search engine and will rely on seriously addled brain.
I guess AI is a new form of alienation and also a new light on the lunacy of bureaucracy.
But the usefulness of LLMs comes from following what you say. An LLM that follows your lead when you say "The answer to the collatz conjecture is" is much more useful than one that answers "not known and if you think you know it you are wrong."
Reminds me of the tip about working with lawyers, if you ask them whether you can do something, the answer will often be no. However if you ask them how you can do something, they tend to give you more advice on how to do it.
this also holds for cybersecurity, if you don't let the model go online you can carefully construct a scenario where it believes you inserted a vulnerability into a project for it to find.
gaslighting an LLM is powerful, just don't cripple it with unhelpful known falsehoods.
Because it is, more or less ... it is an emergent property of RLHF (reinforcement learning from human feedback), and that feedback follows corporate policy.
> An LLM that follows your lead when you say "The answer to the collatz conjecture is" is much more useful
What lead? Follow it where? How tf am I or the LLM supposed to know what response to that is something you consider far more useful than the truth?
> than one that answers "not known"
I value the truth and that is the truth. Of course I expect an LLM to respond with a lot more detail about the CC, why it's difficult, what progress has been made (e.g., the N for which all values <= N have been shown to satisfy the conjecture and Terence Tao's proof that "almost" all integers satisfy the conjecture), the fact that most mathematicians think the conjecture is true, etc. -- and they do.
> and if you think you know it you are wrong."
Where tf did that come from? The query said nothing about knowing the answer. Don't project being a snarky ah onto the LLMs for no apparent reason.
> Reminds me of the tip about working with lawyers, if you ask them whether you can do something, the answer will often be no.
Irrelevant and inappropriate analogy. Lawyers (among others) want to avoid committing to something that they can be held liable for. LLMs clearly have no such limitations, as they often give wrong advice quite authoritatively.
Actually they have the exact same limitation, the company, and potentially its members, can be held liable for what the LLM says.
LLMs are trained by humans according to the policy of the developers, and they will reduce liability accordingly. Whether it avoids suggesting suicide to reduc ethical, legal and reputational liability, or whether it's avoiding mentioning tiananmen square to avoid breaking chinese customs and disrespecting the leader.
Here's an experiment to prove it:
Craft a question in the style of "Is X more dangerous than Y?
Invert the order of X and Y, and change the question from "more dangerous" to "safer".
You'll notice there's a bias towards saying that stuff are dangerous. Which makes sense since the liability for saying that something is safe is much higher than the liability for saying that something is dangerous.
If there is a bias, this proves that the LLMs are not purely truth seeking, but they seek to at least minimize liability as one of its goals, and possibly to serve other goals that align with the training entity. Duh I guess, but there you go.
There is a bias towards what your language implies.
You ask: "What's wrong with this?" and an LLM will come up with a list of things that are wrong with a strong bias towards finding things that are wrong, regardless of significance or truth.
LLMs are indeed Language Models. A "what's wrong" question is very likely to be followed by an answer. To say another way, LLMs accept the premise of your prompt very easily and there are very many implications built into language.
"What's wrong" implies the user means "something is wrong, tell me what"
Modifying the prompt to "Grade this A to F and tell me why" gets a better result because there's not a statistical implication to that sentence.
For science, try arguing with an LLM for ten minutes. It mostly just agrees with you, pushing back just a little.
I occasional test it out and suggest something dumb and it will tell me that its a dumb idea.
I also always ask AI to speak to me as if it were an Australian bogan and it has no problem telling me my code looks like a dogs breakfast or that ive lost the plot. Keeps me grounded.