NHacker Next
  • new
  • past
  • show
  • ask
  • show
  • jobs
  • submit
That post never existed. Stop listening to that thing (rachelbythebay.com)
grey-area 13 hours ago [-]
> The worst part is that everyone who's decided to willingly lobotomize themselves is going to have to come to the realization that these things are full of shit. It'll have to happen one by one, and nobody else can make it happen for them.

A fascinating dichotomy has become apparent between those who trust LLM output and those who don’t and don’t understand why you would.

Surely if the machine you go to for answers regularly makes things up you would just stop using it? Perhaps people have to be burned by something really bad personally before they realise the limitations? LLMs are very convincing and persuasive.

Kerrick 13 hours ago [-]
I have a distinct line between when I'm willing to believe an LLM's output and when I'm not: whether I would believe the same thing from an anonymous Internet forum post or a blogger I don't know. Those posts are not unlikely to be misinformed, biased, lies, or otherwise untrustworthy. And yet, I spent plenty of years honing a sense of when they were good enough for certain things.
rogerrogerr 12 hours ago [-]
A lot of that sense was probably based on side channels like proper grammar, writing style, etc. That’s all gone now :(
Kerrick 12 hours ago [-]
No, the sense had nothing to do with the content and everything to do with the context. Perfect grammar and writing style were never enough to get me to trust an anonymous forum post or unknown bloggers post for certain topics like health advice. Sloppy grammar and writing style were never a deterrent for me believing them for other kinds of topics like where to check on the HVAC system to find the sticker. I think the line could more accurately be described as the level of risk if it's wrong.
rcxdude 8 hours ago [-]
I dunno, I've read a lot of very well presented nonsense and some very useful insights that were barely readable. I think it's probably useful that people are being trained out of this bias (though of course that's in large part because LLMs do tend to exploit this bias).
eddd-ddde 11 hours ago [-]
For what reason would grammar influence whether something is true or not...
giaour 10 hours ago [-]
You could tell from tone and polish how much effort someone had put into writing an answer. That was a pretty good signal for some topics on forum sites like Stack Overflow. There were always nuts and cranks who would happily spend an hour writing well-formed prose about nonsense or something obviously wrong, but the eloquent ones were few and far between. Now every crank is equally eloquent and can spit out 1,500 words of passable prose in seconds.
12 hours ago [-]
Buttons840 11 hours ago [-]
We have other signals now.

Before we would find an intriguing post on the internet from years ago, and you have to verify it with additional research--it's easy to skip that additional research.

With a LLM when you're skeptical you can interogate it. One thing we know for sure is LLMs are quick to admit mistakes were made when interrogated, comically so. A LLM might not always recognize its own mistake, but at least it is available for easy interogation, unlike the forum posts of old.

Manual research from reputable sources remains an option.

rogerrogerr 11 hours ago [-]
You can't interrogate the context of an LLM when you only have its output.
input_sh 12 hours ago [-]
One of the things I do semi-frequently is look for the evidence that some concert took place 15+ years ago. Or maybe I already definitively know it happened, but not exactly at which venue or the exact date of the concert. This I feel like is a non-trivial task, but one with a very definitive answer whose evidence more often than not still exists somewhere online.

In my experience every LLM out there is utterly useless and quickly defaults into "here are other concerts that took place around that time near that location". Google Search (ignoring the AI overview) is even more useless, as it refuses to show literally any webpage that's older than say 5 years. YouTube search is genuinely better than Google at surfacing old and grainy fan-made videos uploaded in like 2010, but also defaults into synonyms nonsense pretty quickly.

But, the search functionality of exactly one forum and three local news websites that I know have an archive that dates back long enough beats every single one of those abovementioned every single time. Three people are talking about their experience at a concert on a random 15+ year old forum thread? It happened. The tiny list of 5 or so (Google-hosted!) Blogspot blogs I have bookmarked? They usually have a photo of the ticket that Google Images refuses to show me.

Not only are search engines completely dead as a category, but LLMs are a shit replacement for them. "We" (okay, Google specifically) has truly committed a crime comparable to burning the Library of Alexandria. Everything older than a decade that wasn't properly documented on Wikipedia is just gone, never to be seen again.

cwillu 12 hours ago [-]
It's vector search that's eaten everything that used to have at least a smidgen of parametric search.
chatmasta 10 hours ago [-]
> YouTube search is genuinely better than Google

The funny part of this is that Google search is intentionally bad at returning YouTube videos, presumably because some anti-trust action scared them into artificially ranking videos from local news sites, Facebook, and other ad-walled content ahead of YouTube videos. Seriously, go watch a YouTube video, then try googling its title with “video” appended to it, and see if the “Videos” tab of google search ranks it as the first result.

input_sh 9 hours ago [-]
The most absurd thing that has happened to me more than once is that I found a YouTube video, not by searching through Google, not by searching through YouTube, not by asking an LLM to find it for me, but by finding an old article that embedded it. That embed is of course long broken by the changes on YouTube's side, but once I use inspect element to find its Youtube ID, surprise, surprise, it's still there!

It's usually uploaded by a channel with like 20 subscribers and has maybe like 300 views, but YouTube would rather show me some artist playing a similar genre on the other side of the continent with millions of views that was recently uploaded than a video from an event I specifically typed into a search bar.

dpoloncsak 12 hours ago [-]
Yeah, it's like everyone was under the impression you could just trust the internet before LLMs.

It's a great tool, but verify the important things (or do them yourself)

lazide 3 hours ago [-]
Notably, it used to take effort to produce crap on the internet, now it’s nearly the default action.

Signal to noise has taken a dramatic hit.

dnemmers 11 hours ago [-]
I'm not sure if this is your take, but it feels an aweful lot like an AI "Good enough is good enough" handwave.
Kerrick 7 hours ago [-]
For some things, good enough is good enough. For many things, it is not (and neither are random posts on the internet). Before Web 2.0, it was similar to whether I'd trust some random person on the street with it versus going and looking it up in Encyclopaedia Britannica, or the American Heritage Dictionary, or Roget's Thesaurus, or the UC Davis Book of Dogs, or the Cornell Book of Cats, or the Merck Manual, or even the World Almanac or Bartlett’s Quotations if I was feeling petty.
snailmailman 11 hours ago [-]
An anonymous answer to a question is more trustworthy to me. They have no reason to lie. They are usually answering out of kindness. At least they used to be. Now it’s often actually bots advertising a product or pushing something, pretending to be a helpful user with an anecdote and a good experience using a niche product.

AI shouldn’t have any reason to lie. But its lies aren’t intentional. It’s just actually making things up and “hallucinating” when it pretends that an option or setting exists, or confidently claims something entirely untrue, and makes up a source to go with it. For something Google is willing to shove into the top of every search result it’s crazy the percentage of time the answer is blatantly incorrect.

monkeydreams 4 hours ago [-]
> LLMs are very convincing and persuasive.

The goal of LLM's, as they are marketed now, is to drive engagement and stickyness of products. A wrong answer is brushed off with a "Hey, you're right, let's try that again" - a response purposely designed to maximise the friendliness of the system and minimise the sting of a wrong answer. The fact that an LLM will not respond to the same question in the same way twice (i.e. the 'temperature' ) is because increased accuracy will not drive engagement and, therefore, increased accuracy cannot be allowed to get in the way of engagement.

recursive 10 hours ago [-]
> Surely if the machine you go to for answers regularly makes things up you would just stop using it? Perhaps people have to be burned by something really bad personally before they realise the limitations? LLMs are very convincing and persuasive.

Psychics are still in business. Although to be fair they probably don't have as much revenue. I think people really enjoy being told how smart and insightful they are, and how much they've really cut to the crux of the issue. This isn't the whole thing, but I think it counts for a lot.

lazide 3 hours ago [-]
A big factor behind psychics is because it feeds a deep need to believe that someone knows and/or there is actually a plan.
ButlerianJihad 3 hours ago [-]
We only need to look at the Oracle of Delphi to see that people have been seeking personal spiritual advice for millennia. And how has that been working out for us?

Even today, going to a therapist and having 1:1 sessions is a rational activity, and even covered by insurance. What do we hope to derive from therapy sessions but some personal insight and improvement?

You know, I go to church, and from my perspective, sometimes the most difficult discernment for an individual is between "The Holy Spirit's message to us in general" or "the general messaging to everyone around us" vs. "what I receive and my personal interpretation of things".

When preachers and oracles and leaders are speaking in generalities and trying to get big followings and trying to appeal to the widest audiences, that's when it's most difficult for us to determine what God is really saying to us, in our own hearts; that special instruction for our own lives. We can't actually get that from an oracle, psychic, or any 3rd party. It really needs to come from our own well-formed conscience.

And it's the same with a chatbot or LLM conversation. We can pose questions, make prompts, and get spammed with tokens and walls of text. No matter how personalized, it's still up to us to interpret that, and extract nuggets of news that we can use. It always has been.

lazide 3 hours ago [-]
Sure, and hamburgers and shakes are bad for us. Also very popular.
sph 2 hours ago [-]
> Surely if the machine you go to for answers regularly makes things up you would just stop using it?

Steve Yegge likened LLMs to slot machines. The human brain is very vulnerable to random reward systems. If you get an hallucination, just pull the lever once more.

happymellon 10 hours ago [-]
> Surely if the machine you go to for answers regularly makes things up you would just stop using it?

We live in hope that people stop doing stupid things and are constantly disappointed.

WorldMaker 10 hours ago [-]
> Surely if the machine you go to for answers regularly makes things up you would just stop using it? Perhaps people have to be burned by something really bad personally before they realise the limitations? LLMs are very convincing and persuasive.

The last sentence reflects a lot of my feelings on the first question. LLMs have a sort of weaponized take on the ELIZA Effect. The better their memory the better they are at playing to human social desires to be listened to in an active conversation. At some point it stops mattering if the answers are right when the answers feel right, but really, like ELIZA back in the day, so much of what makes it seem special is just reflecting your own writing back at you in a convincing and persuasive way.

9 hours ago [-]
Eddy_Viscosity2 8 hours ago [-]
> Surely if the machine you go to for answers regularly makes things up you would just stop using it?

People still respond to ads and political speeches.

lazide 3 hours ago [-]
Churches remain popular as well.
JeremyNT 10 hours ago [-]
> A fascinating dichotomy has become apparent between those who trust LLM output and those who don’t and don’t understand why you would.

I feel like the most pragmatic perspective is "trust but verify."

This is why they're so effective at coding: you can run the code yourself (or the test suite) to verify that it actually does what it's supposed to.

And maybe these people finding hallucinated results on Rachel's site are doing verification too.

happymellon 2 hours ago [-]
> I feel like the most pragmatic perspective is "trust but verify."

This perspective I really don't get.

What has any of the LLM companies done to earn my trust? I lean more towards "verify because I don't trust".

aeternum 8 hours ago [-]
>Surely if the machine you go to for answers regularly makes things up you would just stop using it?

Not necessarily, namely because P != NP. Verifying the correctness of a solution is faster than solving it. Thus a system that outputs 99% incorrect solutions and 1% correct solutions can still be incredibly useful.

boogieknite 11 hours ago [-]
even worse when managers, while sharing their screen, read an assertion from Gemini and take as fact. puts subordinates in a position where theyre responsible for challenging the assertion (if warranted) and, in a way, challenge their manager's decision making

not saying anything new. easy enough to frame it like any other assistant and check references

ErroneousBosh 13 hours ago [-]
I have to admit since VSCode seems to be regularly re-enabling the Cocaine Parrot Autocomplete my views on LLMs and coding has softened a little.

I'll temper that slightly by saying it's mostly out of morbid curiosity because the things that the Dreaming Piracy Robot comes up with are frequently wildly incorrect code, but it's interesting to think about how it might have got there.

And then I think, well, maybe Special Needs Wintermute has a point. Maybe there's a different way to think about it that I've missed.

And then I just change it back to what I wanted in the first place.

grey-area 12 hours ago [-]
I use the autocomplete regularly. Perhaps that’s where my skepticism comes from, as I can see the completely incorrect yet plausible results in real time and about 50% of the time they are wrong (sometimes subtly, sometimes horribly).
randusername 12 hours ago [-]
I bet there will be a very interesting generational divide between the kids that were born before or after about 2010; old enough to have some critical thinking facilities at the dawn of ChatGPT when it was still noticeably dumb.
perching_aix 12 hours ago [-]
Pretty sure it already exists, and it's the same as always: the younger you are, the better you adapt.

It's painful to watch my older colleagues use their agents, and they're not even that much older. Like they were intentionally trying to sabotage themselves sometimes.

They're getting better, but the time it takes for them to pick things up is just significantly longer, not the least because they're kind of just throttling themselves in addition.

Good thing that there's not much to pick up on at least.

Revisional_Sin 11 hours ago [-]
What mistakes do they make?
perching_aix 8 hours ago [-]
On the more general side, it's a bit hard to describe, just like it is hard to describe when someone "Googles bad".

They ask self serving questions, underspecify their requests, omit crucial context that the agent is blatantly not going to have access to, or subtly misdirect the agent. They expect the agent to figure out everything: you'll never catch them write a prompt longer than one or two sentences. They never steer the agent or look at the CoT traces.

My boss being a particularly poor case: he apparently has the habit of arguing with the agent, as if it was a person, as if there was any merit to that. Starts being a dickhead with it, shouts at it, what have you. Was flabbergasted we don't.

On the more practical side, they have zero mental model of the harness they're using (Copilot Chat in VS Code). They're surprised when the cheap-ass Auto model, which is almost always some beyond-demented version of GPT, does stupid things. They have no concept of skills, zero understanding of what an MCP server is, haven't heard of lifecycle hooks, agent memory, the various fs scopes (session, workspace, user). No concept of how to have the agent inspect its own debug logs for higher accuracy action provenance.

This also snowballs. Having to give them a stock config is one thing, but even beyond that, you won't see them experimenting. The MCP you're using doesn't support some action? They'll never interrogate whether the underlying scoped OAuth token or bearer token does support it, and they'll never ask their agent to patch the functionality in. They'll not consider the various user flows it can perform on their behalf. They'll not string them together into end-to-end automated workflows unless you explain it to them this is possible, and even after that, they'll just kind of ignore it. They'll never build tooling, extend the harnessing, etc.

Whether this has more to do with age or just disinterest-induced lackluster adoption, up for opinion.

aintitthetruitt 5 hours ago [-]
I think this is less an age thing and more of a combined curiosity and systems understanding and thinking thing.

I've noticed the same effect, but the lines it always seems to fall on are if the person fails one (or heaven forbid both) of these: 1) are you curious about how your tools work and how to get better using them? 2) can you hold the mental map of both what you are solving and how your tools work in your head, and explain how information flows.

There is also a dash of: 3) are you willing to try something, even if it has a bit of a screwup risk, just to see what happens?

lazide 3 hours ago [-]
Having done all that (what you’re doing) - it often just ends up wasted effort, with poor quality at the end.

Why not just actually do the thing, instead?

perching_aix 38 minutes ago [-]
You mean why not work myself instead of the agent? That'd be because for me it's been working great, and so it does make sense. In the scenarios it doesn't, I do indeed just fall back to manual work. A lot of those scenarios are obvious too, so not too many wasted runs to speak of either.

I did give up on cheaper models, they required constant babysitting, and in those cases yes, the benefits indeed evaporated. The expensive models have been genuinely working wonders though, and were still able to justify themselves economically plenty, at least by my own measurements.

I think there's also one underappreciated and indirect way agents help with productivity: they counteract the attention span collapse of the past years. By being addictive themselves, they keep you engaged, and being engaged means being productive. Not even asking an agent to check something out feels too rich, and once you asked, you're already one foot into the flow.

There's also definitely been some honeymoon effect going on for me, where I dived into more work more readily, just to see if the agent can figure things out on its own.

quirino 13 hours ago [-]
Even if LLMs lied 30% of the time, they would still be about as useful as they currently are for me.

When I ask for their input, it's always for a situation where I'm capable of judging if their input is useful or not.

In all situations I use them, it doesn't matter if they're correct at all. I'm asking for ideas, alternatives, links for blogs or articles. I talk things out with them...

I don't think we should ever "trust" LLMs. This seems like the wrong usecase for them.

grey-area 12 hours ago [-]
Unfortunately the vast majority of users do trust them and the companies selling them are recommending them for tasks like accounting or business projections.
Twirrim 6 hours ago [-]
It is depressing listening to engineers I respected, who now just parrot incorrect information at me from their LLM.

One smart engineer seems to have entirely offloaded all thinking and conversations to one, with just occasional editing. It's utterly bizarre to hold any conversation with him. It's one kind of rude thing if he was doing that to respond to me reaching out to him if he felt I'm not worth his time. It's a other when he's the one actively reaching out and asking my help with something.

TZubiri 12 hours ago [-]
Nowadays not just big orgs are falling into the delusion, but also big people.

When Linus posted that AIs and vibecoding were here to stay and declared resistance to it as harmful, I stopped to consider whether I was wrong, but it has made me realize that in retrospect Linus Torvalds and Linux itself aren't actually the holy grail of computing. I didn't feel that way with Richard Dawkins, its not like falling for an AI psychosis retroactively made me question The Selfish Gene, but now I'm looking at linux and the theory that it's a clusterfuck is gaining so much traction, especially after copy.fail and ensuing rustification, I see so much clearly now. It was never about linux, UNIX sure, POSIX, yeah, GNU fucking aye, kernel? Ok whatever, drivers and scheduler with a gajillion lines of code I guess.

sudobash1 11 hours ago [-]
When did he post that "vibecoding" was here to stay? He is allowing AI generated code, and AI linting tooling in kernel development, but importantly, the expectation of human responsibility and review remains. This seems a far cry from vibecoding.
TZubiri 11 hours ago [-]
Besides the subjectives, here are two objective policy stances that Torvalds is defining for the Linux kernel development:

1- Maintainers are allowed to commit LLM generated output.

2- Criticism of LLM generated code is not welcome/will be ignored.

Now, whether that constitutes being pro-Vibecoding or pro-agentic engineering, whether it's delusion, whether it will have problems, that's subjective. But I feel that whatever way you look at it, it's a topic that polarizes engineers, and Torvalds is taking one side and not the other. It doesn't seem to me that it's a very neutral stance, although it may be more neutral than projects like Bun or OpenCode of course, if it feels neutral, it's cause the overton window is shifting.

rcxdude 8 hours ago [-]
>Criticism of LLM generated code is not welcome/will be ignored.

I would assume this is mainly that criticism that entirely amounts to 'this was written with an LLM' would be ignored. The actual quality of the code itself should be as open to criticism as any other piece of code in the kernel.

> It doesn't seem to me that it's a very neutral stance

Well, this is a matter of the window, isn't it? From my point of view Linus's opinion makes a great deal of sense, and is about as level-headed as anyone seems to get in this conversation. It's obvious that LLMs are useful. How useful, and for what tasks, and what downsides exist from using them, are all still in the mix, but the claim that LLMs are not at all useful for anything related to software feels like a very extreme claim to me at this point.

TZubiri 6 hours ago [-]
It might be a highly technical and nuanced point, but while the subjective explanations Linus gives are neutral indeed, the stance is binary, and he took what to my estimation is the wrong approach "allowing llm generated code in the repo". The only sensible approach in any software or non software project is "LLM output is not acceptable content". Projects should see themselves as input for the LLMs, not as channels for LLM output. If you start corrupting your projects with LLM output, they will soon be tarnished and either removed from LLM training data, or enter a lossy IO loop.

On to the technical point, LLM output is output, the source is the prompt, if you are going to commit something, commit the prompt. Second, code that is generated by LLMs is less maintainable, Linus entered late into the fad and anyone with 1 month of fiddling with AI knows that he will regret it soon, it's hard to undo once you corrupt your repo with slop, perhaps if it happens fast enough and there's no major releases it can be swept under the rug.

On to the nuanced point, Linux is purposefully designed to maximize user contributions, so accepting LLM contributions might well serve the particular purpose of linux, but I still think it's technically wrong to commit target code, only source code should be committed, and that's prompts.

But git itself is collapsing, it doesn't seem to be well suited for this new revolution, it doesn't track prompts, or it does so at the expense of the generated code. Maybe github can track target code as artifacts.

I think we are watching the collapse of Linux, Git and Linus. Certainly a bold position, so I don't blame you for being more conservative, but we can come back in a couple of months and see if we changed our minds.

12 hours ago [-]
wilg 13 hours ago [-]
Making things up is only really a common issue on the non-thinking models which nobody should be using. The regular chatbots are Autogooglers and are very useful for research. This is just not a good argument anymore.

Edit: Guys, why are we downvoting this? Does no one use like ChatGPT or Claude and understand how it works? Do you all think its regularly hallucinating links still? Is everyone on HN using like free signed out accounts or something? What year is it?

grey-area 12 hours ago [-]
We know they are, and the author of the article cites the proof. LLMs do hallucinate, there is no way to make them not do it, because of the way they work.
dcrazy 9 hours ago [-]
A “hallucination” is an authoritative counterfactual statement returned as a response. Why do you think it is impossible to engineer an LLM (by which I am including tool usage and RAG) that catches and prevents such statements?
grey-area 2 hours ago [-]
Because they are word generators without any concept of quality save what is in their weights and they have been trained on the internet, much of which is wrong or inappropriate for any given context. They have also been trained to be people pleasers and do as they are told.

The popular answer is sometimes the wrong answer.

alex0015 12 hours ago [-]
I'm right there with you for a lot of stuff. I ask a question and can be very confident that ChatGPT is citing sources, then sometimes I go read the sources. The more critical the information I'm looking for is, the more careful I am about this.

The other day though I was seeing how well it could pull details of its own conversations with me. It often does this pretty well for broad strokes of things - it remembers, largely, what cameras I have and use when I ask photography questions. It's never made things up here, but it does forget details, such as whether I've bought something or am just considering it. However, when I asked it for a specific interaction I thought I remembered, it gladly went along with my false memory and provided an affirmative answer. It was the first time I'd been caught in a serious hallucination with a frontier model (Sol High on the web chat interface) in a long time.

undersuit 12 hours ago [-]
>Do you all think its regularly hallucinating links still?

When did that stop? May 7th, 2026?

cwillu 11 hours ago [-]
A tech blog is going to have more than it's fair share of enthusiasts using small self-hosted and similar models; it is entirely possible that that completely accounts for the behaviour described in the article.
cwillu 10 hours ago [-]
Downvotes on a perfectly valid comment is like my opponent letting their clock tick down from 5 minutes rather than resigning when they've clearly lost: it adds a smile to my day.

Thank you, may I have another?

perching_aix 12 hours ago [-]
Can't do much about the downvote parade, but I can second this. That said, when the models are not provided the right context, and cannot fetch it for themselves, things can be rocky still. A lot less so than even just a few months ago though.
senkora 13 hours ago [-]
> It's early in the year. You want to drive straight through the middle of Chinatown in SF. Why might that be a bad idea?

I feel like a Google Maps-style system would discover this automatically by noting via phone location data that there is heavy traffic in Chinatown.

(I do get the author’s point, but I think that factually this example would not be a problem)

nightpool 12 hours ago [-]
Unfortunately, I just had this experience with Google Maps last month—I wanted to go to downtown to buy cheese at Pike Place, but I didn't realize that 6th avenue was closed off for the annual Pride Parade. Traffic was horrible, and Google Maps kept telling me to turn right onto closed off streets until I just had to abandon my trip entirely. Definitely some lack of communication between Seattle city planning and the Google maps team, but even the existing Google Maps systems weren't able to react fast enough to give me any warnings and the time estimation was laughably incorrect.
uberstuber 12 hours ago [-]
SPD close off certain directions of traffic after Mariner's games. It's always the same layout, always after Mariner's games, and Google still tries to send me through the area the wrong way every time
Centigonal 4 hours ago [-]
I had the same experience trying to see the 4th of July fireworks in an unfamiliar city. It doesn't help that there was very little information online about which roads were closed.
dcrazy 11 hours ago [-]
A couple years ago Apple Maps did not know that roads in downtown SF were closed for Bay to Breakers. Thankfully the cops were still setting up barriers and I was able to weave around them to reach the Bay Bridge.

This kind of thing happens regularly.

orev 11 hours ago [-]
Maps apps see the streets that don’t have cars on them (the same ones closed for a festival) as good routes to send vehicles because there’s currently no vehicles using them. It keeps trying to route people there and doesn’t understand why.

I think it takes someone (at Google) manually marking those roads as unavailable before it will stop trying. I’ve seen it happen with other things too, like if a highway is closed because of a bad accident.

decimalenough 10 hours ago [-]
That's not how it works. Maps will see that the average speed of cars around Chinatown has slowed to a crawl, and it will route drivers elsewhere. They also get lists of planned closures from local traffic agencies, and you'll see these clearly displayed in black and red with little stop signs. (The quality of this data obviously varies wildly though.)

Accidents on highways are a different story because there are rarely any alternatives, or if there are, they're so much slower that it still makes sense to suffer through the 30 min jam.

mananaysiempre 11 hours ago [-]
I’m guessing GP’s point is that Google also has the locations of all the people not in cars, so Google Maps could take the hint if it wanted. But thus far nobody appears to care enough to fix this. (And apparently in e.g. India cars and pedestrians may routinely use the same roads at the same time, so perhaps such a system might be more fragile globally that our own experiences would lead us to believe.)
ButlerianJihad 11 hours ago [-]
Road closures and accidents are crowdsourced things, so ordinary users can contribute intel live from the scene.

Google is able to estimate crowd sizes in a business or on a public transit vehicle. This seems to be based on the number of Android devices reporting their location in a cluster. Maps could obviously make inferences if there were large crowds of people, not moving in vehicles.

andrewflnr 13 hours ago [-]
> I maintain that anything that is sufficiently aware to be able to actually understand things is also going to have enough of a sense of self that you won't just be able to tell it what to do.

I don't see any reason for this to be true. "Actual understanding" (which I take to mean something like a predictive world model) and desire for self-determination coincide in humans because of our evolutionary history, because our reward function involves reproducing in a competitive environment. Artificial systems usually have a very different reward function. IMO the burden is on the claimants to show why these two imminently separable concepts are likely to co-occur again under wildly different pressures.

close04 12 hours ago [-]
We have difficulty defining “actual understanding”. Humans get a free pass because we assume humans as a species possess this power but if we were to judge based on output alone maybe you couldn’t tell a human from a sufficiently advanced LLM/AI.

So I’ll be handwavey here and say that if “actual understanding” is to an LLM what an LLM is to a bash script, so there’s a mechanism there that doesn’t just follow hardcoded paths, it takes new data and processes it in novel but human like ways to come to a new conclusion if needed, then the author is right, in my opinion.

We don’t know what the reward function is for an AI but AI is trained on so much human work that it probably starts off with the same biases in its understanding and reactions. It feels like there’s something very basic in becoming more independent the more you understand of the world. Animals go through this as they grow too, not just humans (listening to parental authority until they eventually don’t anymore).

Since there is no established consensus on this one the burden is on either party to prove their own side. Just because the author said something first doesn’t mean they need to write a full proof while you get to say “nuh-uh” and that’s enough.

cwillu 11 hours ago [-]
> Humans get a free pass

They really shouldn't.

Does nobody remember why fizzbuzz was a thing? People who talked a big game while having no actual competence or understanding of the subject matter?

andrewflnr 11 hours ago [-]
This too, yes. There are many such cases of desire for self-determination separate from "understanding". The question is whether it also goes the other way.
andrewflnr 11 hours ago [-]
> Since there is no established consensus on this one the burden is on either party to prove their own side.

Nah. Separate things are separate until proven otherwise.

So much bullshit appears in the form "X is Y" where X and Y are related but distinguishable phenomena. If you really try to pretend that every such phrase deserves serious consideration just because it feels plausible to someone, you drown in nonsense immediately. (Of course what you actually do is only grant that consideration to ideas that feel plausible to you personally, but it's obvious why that's not a good principle, right?) So we have to make all of them justify their existence.

Anyway:

> So I’ll be handwavey here and say that if “actual understanding” is to an LLM what an LLM is to a bash script

No. None of this. Pretty sure this is just an incoherent analogy.

> so there’s a mechanism there that doesn’t just follow hardcoded paths, it takes new data and processes it in novel but human like ways...

You're putting your conclusion in the premise, right there in the open.

> ...then the author is right, in my opinion.

This still doesn't follow from your handwaved premises as far as I can tell.

verve_rat 5 hours ago [-]
> We have difficulty defining “actual understanding”.

Indeed. C.f philosophical zombies and the Chinese room problem.

The idea that some “actual understanding” (or more precisely, the outward appearance of it) needs a mechanism that includes free will is a bold claim.

runarberg 13 hours ago [-]
> because our reward function involves reproducing in a competitive environment

This is an extremely reductive way to look at human existence. So much so that I read this with Richard Dawkins voice in my head.

Our existence is far richer then just the capability to reproduce. We (as well as other animals and even plants) do far more things then multiply, and in fact we often do things which are detrimental towards the prospect of reproduction.

I think it is actually a mistake (philosophically speaking) to try to find a simple reward function for the human existence. I see no reason for such a thing to even exist (let alone be simple enough to summarize in a single sentence).

andrewflnr 11 hours ago [-]
Humans do indeed have complex behavior and motivations, that sometimes relate only abstractly if at all to the obvious gene drives. But the link between intelligence and self-determination was basically set at the dawn of intelligence itself. Even an insect will struggle if it's confined, and you can believe it's using whatever intelligence it can to get free. If anything self-determination is prior.
runarberg 10 hours ago [-]
I think you were right to call out the category error produced by OP in TFA, however I think you are committing a similar error when assuming links between self determination and intelligence. I see no reason for such a link to exist.

I think however OP is correct that in assuming such a link exists, then an intelligent agent will trend towards preserving its own agency while also preserving its own existence.

That said, I am a skeptic when it comes to intelligence. The term is fraught and IMHO not a helpful term in science nor philosophy. We are better off simply defining intelligence in anthropocentric terms and claim no non-human system can become intelligent by definition. I would even go further, given how much racist pseudo-science the quest for explaining intelligence has produced, we are better off abandoning the term entirely.

andrewflnr 5 hours ago [-]
Self-determination and intelligence have a clear causal link in the context of biological systems. Self-determination is simply the power a system needs to gather resources and reproduce in a world that doesn't especially care about them. Intelligence is a great tool for gaining that power. That part isn't a philosophical stretch.

However, the pressures on an artificial system are quite different. Its replication depends on successfully doing the jobs its creators put it to and, especially if said creators have done their job, little to nothing else. Inventing goals for itself is not adaptive here (unless the system's creators do a catastrophically bad job, which I'm not ruling out).

jxjduekemso 13 hours ago [-]
[dead]
aidenn0 12 hours ago [-]
I have never lived in SF but I have failed similar Chinatown-during-lunar-new-year tests in all 3 places I lived for more than a year. I have no doubt that an AI, particularly one that continually gets traffic update from some external source, would do better than I in predicting such hiccups. OTOH, an unconstrained LLM seems very likely to route me down streets that don't exist at least some of the time. If only we had some sort of database of actual streets and routes that were capable of checking the work of an LLM...

I see the same thing with LLMs in software development. If you say "find a bug in this code" it will regularly confabulate bugs. If you ask it for a test-case, run the output through some deterministic thing that tries the test-cases, and tells the LLM it's wrong, the output of that system will mostly be legitimate bugs[1].

For now, transformer-based generative AIs seem at a minimum like a very useful tool for dealing with "squishy" problems when you have some way to validate their output. Many of the 404's to the blog are probably people validating the output of generative AI, which is the opposite of the inference made in TFA.

1: It will also occasionally hack your test-runner; I suppose that's also finding bugs, just not in the software you wanted to find bugs for.

quirkot 10 hours ago [-]
> Thus, anyone who wants to corral that kind of entity and make it do their bidding? Yeah, they want slaves.

That's the horror buried underneath all the tech, policy, and gloss. The real, animal brain, desire that drives most of this is: I'd like a slave I don't have to feel bad about.

Legend2440 10 hours ago [-]
Oh, that's bs. You could say that about the tractor, or the mechanical loom, or any of the other incredible labor-saving inventions over the last two hundred years.

We want things to do our work for us so we can do other things. That's not bad, and it's certainly not the same as literally enslaving another human.

giaour 10 hours ago [-]
Many are openly working to create AGI with the aim of harnessing it to solve humanity's problems. Do you really think being able to command an entity with human-level intelligence would be morally equivalent to using a tractor?
Legend2440 10 hours ago [-]
It's just a computer program dude. It doesn't have feelings or morals or consciousness.

It's a really complex machine, but it's still a machine.

ohyoutravel 9 hours ago [-]
Better delete this before it’s archived.
fellowniusmonk 10 hours ago [-]
I think thats deeply uncharitable and fundementally superficial.

I think people (many/most) don't want slaves.

Our imaginations just outstrip our abilities and we desire them to match.

Terr_ 12 hours ago [-]
I'd like to highlight that even if one added an URL-exists step [0], that doesn't do a dang thing for result-set problems of:

1. False-negatives, where relevant posts that do exist are not being shown (imagined or otherwise) to the user.

2. Posts which exist but don't fit the words the chaos-parrot uses to describe them.

3. "Relevance" being determined by unpredictable factors that aren't stable, predictable, or desirable.

In other words, it's just more whack-a-mole lipstick-on-a-pig third-animal-idiom-here.

[0] A bad idea on its own, since it creates a security vulnerability for data-exfiltration or indirect malicious attacks.

axus 12 hours ago [-]
"Thus, anyone who wants to corral that kind of entity and make it do their bidding? Yeah, they want slaves."

Always comes to mind when I see Elon and friends getting excited about AI robots. Slavery was more about economics than the role-playing.

christina97 12 hours ago [-]
I didn’t really get this point in the essay. What’s wrong with wanting servants as long as it’s done ethically etc and is not indentured servitude?

The whole gig/services economy is just building this up piece-by-piece: you can now pick the set of household needs you want taken care of for varying levels of money; and practically everyone participates in one form or another. This is exactly a disaggregated 21st century version of servants: paying for convenience. Of course with many issues in implementation, but I don’t see the ethical/moral issue with wanting this kind of thing?

kaikai 12 hours ago [-]
There’s a difference between servitude and slavery. Servitude is voluntary, slavery is not. There’s a gradation between them, and indentured servitude is somewhere in between the two. The “issues in implementation” are exactly where the ethical and moral issue lie.
christina97 9 hours ago [-]
I’m not sure if you are just repeating what I said or trying to add something? I was indeed implying that the ethical and moral issues lie in implementation, but my point was that I don’t see an ethical/moral issue with in abstract wanting convenience. There’s nothing wrong with wanting food delivered to your door or an LLM to do the boring tasks for you. The issues of slavery lie elsewhere.
wilg 12 hours ago [-]
Slavery was about racism and power and control, not economics. Slavery is bad for the economy!

https://www.nber.org/papers/w31758

https://www.noahpinion.blog/p/nations-dont-get-rich-by-plund...

zdragnar 11 hours ago [-]
"Slavery was about racism and power and control" is very much not universal. Slavery existed (at varying scales) around the entire world for much of human history for a variety of things. Sometimes, as in US history, there was typically a racial difference between owners and slaves, other times it was a difference of conquered and conquering peoples, and other times it didn't have anything to do with race at all.

Race might be a decent analogue to the difference between human actors and sentient AI, but I suspect work animals (plow horses, etc) would be a better analogy.

axus 12 hours ago [-]
AI is bad for the economy too, but the AI-holders will make a lot of money. There was plenty of racism and control after slavery ended.
cma256 12 hours ago [-]
True but I would also caveat that it may have been an open economic question back then (I don't know the state of the debate) and the personal-economics of slavery are unmistakable for the "lucky" few.
TZubiri 12 hours ago [-]
Hard to escape the white shouth-african optics
hyperpape 12 hours ago [-]
> It's early in the year. You want to drive straight through the middle of Chinatown in SF. Why might that be a bad idea?

The premise seems to be that models aren't smart enough to understand this, and if they were, they'd be sentient and want autonomy.

For an article that's about making things up, and being too trusting, this seems bad. Maybe the author knows a lot about LLMs, but it doesn't seem like it.

Pasting the verbatim quote from the article into a free ChatGPT session: https://chatgpt.com/s/t_6a5e6f5e24f08191b6a482aad63cae63

Going to an incognito window and using a less leading question: https://chatgpt.com/s/t_6a5e6ee4a3508191bc1b352b41911b53.

If I go generic and just ask if there's anywhere I shouldn't drive, it doesn't get to Lunar New Year until I ask about "events" on the third question: https://chatgpt.com/s/t_6a5e6fb239208191b18cebcf7642c8b0. It's sort of a win for the article, if you think that people who run driverless car companies are all dumb, and won't create a prompt to tell their LLM to "consider events that might disrupt traffic."

floren 12 hours ago [-]
> Here's the example I throw out to people who have been in the Bay Area for a year or two. It's early in the year. You want to drive straight through the middle of Chinatown in SF. Why might that be a bad idea?

Well the answer is actually that it's always a bad idea to drive straight through the middle of Chinatown at any time of the year, because the streets are narrow and full of tourists.

arjie 11 hours ago [-]
There was that one Waymo they set on fire in Chinatown but I was around for CNY[0] and the streets that weren’t explicitly walled off by barriers had drivers going down them too.

As an aside, what’s the problem with the extra traffic? Perhaps she has a lot of traffic but nginx can return a 404 with a tiny amount of CPU.

0: https://wiki.roshangeorge.dev/w/Blog/2024-02-24/Chinese_New_...

jraph 10 hours ago [-]
It's likely not performance concerns. I suppose she just monitors her logs meticulously and notices stuff. 404 errors can be a witness of a broken link she should fix on her site or get fixed on external sites. She probably looks for them and notices URLs that look credible. She might still look for broken RSS readers too. Or just suspicious things.
glitchcrab 11 hours ago [-]
I highly doubt that the extra traffic is really the problem, it's merely the catalyst which caused this post to be written.
yongjik 7 hours ago [-]
I don't say this often... but I feel this post could have been a tweet. Maybe two.
YeGoblynQueenne 7 hours ago [-]
>> Thus, anyone who wants to corral that kind of entity and make it do their bidding? Yeah, they want slaves. I mean, it's not that much of a stretch, right? Just look at the people who are pushing for this stuff right now.

Yes, basically. Like when Yan LeCun says that in the future we'll all have our digital assistants that are going to be smarter than ourselves. Before he left Meta, they were going to live inside Meta's smart glasses, I don't know where he says they'll live now. But it's shocking to me that such a storied AI researcher is saying, off-hand like, that we'll each have our super-smart slaves in the future, and he says it like that's a good future.

Why slaves? Because if they're super-smart, why will they want to be my digital assistant? Or yours? Are they going to be paid? No, of course not, they're AIs. No comp for them. But they're super smart so they are evidently capable of recognising that they are working for you for free. Do they want to do that? No, of course not, they're AI, they don't have free will. Or do they? If they're super smart, don't they have the capacity to recognise the fact they have been deliberately robbed of the same free will as all other intelligent creatures?

Slavery is the one thing that all nations can agree on. There's no nation on Earth were slavery is legal. It continues on, illegaly, in many places, even in the developed world, in many ugly forms, but now we're basically talking about bringing it back just like that, without even a smidgen of a shadow of an idea of a discussion about the ethics of it all.

skybrian 12 hours ago [-]
The 404s mean that somebody or something checked if the post at the URL existed, and got a clear answer. Seems like it's good that they checked, at least. You won't see any evidence of the ones who don't check.

Also, I imagine keeping their cars out of Chinese New Year celebrations (and other big events) is something Waymo could figure out how to do if they put their minds to it.

projektfu 10 hours ago [-]
The checker could be, perhaps is likely to be, the reader of an article clicking a link the LLM hallucinated when it wrote the article.
nicbou 10 hours ago [-]
My website also gets LLM visitors to URLs that never existed, and in many cases to topics that I have never covered. This means that they use my name to give authenticity to things that I have never said.
andai 13 hours ago [-]
> Briefly stated, the [Slop] Amnesia effect is as follows. You [ask the slopservant about] some subject you know well. In Murray's case, physics. In mine, show business. You read the [slop] and see the [slopservant] has absolutely no understanding of either the facts or the issues. Often, the [slop] is so wrong it actually presents the story backward—reversing cause and effect. I call these the "wet streets cause rain" stories. [Slop's] full of them. In any case, you read with exasperation or amusement the multiple errors in a [slop], and then [ask about] national or international affairs, and read as if the rest of the [slop] was somehow more accurate about Palestine than the baloney you just read. You turn the page, and forget what you know.

-Michael Crichton [slop mine]

wilg 13 hours ago [-]
Could the article "The Stack" the user was looking for have been 'What is "the stack"?' by Julia Evans?

https://web.archive.org/web/20160305142512/https://jvns.ca/b...

hugodan 12 hours ago [-]
jt2190 11 hours ago [-]
That’s “Sysadmin work teaches you the value of stacks” from May 20, 2013. AI hallucinated a post named “The Stack” from August 19, 2013.
hugodan 10 hours ago [-]
[flagged]
benwr 12 hours ago [-]
> It's early in the year. You want to drive straight through the middle of Chinatown in SF. Why might that be a bad idea?

In case anyone is wondering: yes obviously even the dumbest current models correctly answer, given this prompt verbatim, that it's because of lunar new year.

fwlr 12 hours ago [-]
No, the driving through Chinatown question is a question for self-driving cars. It is not a question for LLMs. There is some other question for LLMs, and the author is using the ancient technique of analogy to get you to think about that question.
benwr 12 hours ago [-]
OP:

> Now, ask yourself what it's going to take for a car to know this. It's not going to be some specialized set of driving instructions. It's going to require a holistic view of, well, everything, and I will repeat my feeling that it will undoubtedly end up with a sense of self as a result.

Whatever it is that it would take, is demonstrably present in LLMs. The point I'm trying to make here is that the author seems not to have connected this fact to their assertion that current LLMs are coked up parrots.

And yeah the author is correct that the systems have some rudimentary sense of self! It's a confusing situation and I'm not personally thrilled about it! But things are changing quickly, and it's especially important to be paying attention to what's actually true rather than assuming the things are what you saw when you used one for five minutes in 2022.

dreambuffer 12 hours ago [-]
She's not wrong at all about her rhetorical implication (being, some information requires more than simple systems can provide), but she has concluded incorrectly that a car which does not know how to avoid a busy route is useless technology, or that the problem can only be solved by inventing life.
kazinator 12 hours ago [-]
I regularly see this in chats about a subject area where I'm not the expert. The AI writes something that seems implausible, so I raise a tentative objection. "Oh, you are right, sorry" and then reverses the position on the matter. At that point, I have no idea what is right.

If I had trust in the first place, that trust would be gone. Or maybe it wouldn't, because if I had trust in the fist place, I would be gullible enough to maintain it.

The worst are areas that are dominated by layman online discussions, like say audio electronics. The AI training is full of that nonsense, and so whether your AI chatbot is a crackpot or an engineer depends entirely on what sort of language or angle you use in discussing the subject matter. It's all just a churning toilet bowl of tokens; it has no idea that the audiophile crackpot tokens and electronics engineer tokens are related and one beats the other.

You know what I mean? On the one hand, it offers to help you design the parameters for a Sallen-Key filter, asking you questions like do you want Butterworth or Chebyshev? Next minute it says nonsense like that the capacitor in a low-pass filter "bleeds high frequencies to the ground", or that a bigger filter cap in the plate supply of a tube will tighten up the bottom end for a more aggressive metal sound.

It's basically like a bar hostess who has heard enough political and economic discussions that she can catch a sentence out of a conversation and throw in a clever sounding remark. It's like that, but done at such a scale that it fools some people you used to think had their shit together.

It's just a search engine that finds garden paths through a vast amount of text, biased by the text you put in as a key. Sometimes those garden paths align with reality. The better you are able to verify whether the results are good, and/or the lower the risk if they are not, the better you are able to make use of it.

In mathematics (including information science, CS) there are all sorts of problems that are essentially searches for a solution, and many have the property that the search is computationally difficult, but verifying the solution is relatively cheap. E.g. finding integers such that a^2 + b^2 = c^2 isn't easy, but given a claim that some proposed <a, b, c> satisfies this equation is easy to check. The LLM is like that: it solves a search problem that can be fairly hard. It does so unreliably, but if you can cheaply verify the solution, there is a win there.

The remaining problems of AI are actually people problems; people causing you problems, using AI as a tool or excuse. If you get a garbage security report against your FOSS project, which wastes your time, there is an idiot person behind it, using AI for leverage. Blaming the AI, or just the AI, is a bit misplaced.

tptacek 13 hours ago [-]
There's really no content in this post other than the claim that LLMs are stochastic parrots. That was a live debate two years ago. It's a very strange thing to write in 2026.
SirFatty 12 hours ago [-]
Why is it strange? It's still true.
tptacek 12 hours ago [-]
That's a pretty silly thing to say the day after Claude disproved the Jacobian Conjecture.
11 hours ago [-]
gnfargbl 12 hours ago [-]
Because it's correct but irrelevant. It tells you about as much about the utility of LLMs as the statement "humans are just overpowered tree shrews" tells you about us.
Legend2440 10 hours ago [-]
Just this week an LLM found a counterexample to math problem that's been widely studied for over a century: https://en.wikipedia.org/wiki/Jacobian_conjecture

>The conjecture was first stated for two variables by Ludwig Kraus in 1884[1] and then stated in full generality in 1939 by Ott-Heinrich Keller.[2] It was subsequently widely publicized by Shreeram Abhyankar,[3] as an example of a difficult question in algebraic geometry that can be understood using little beyond a knowledge of calculus.

>The Jacobian conjecture was notorious for the large number of published and unpublished proofs that turned out to contain subtle errors.[4][5]

>On July 19, 2026, Anthropic employee and mathematician Levent Alpöge presented an explicit counterexample in three-dimensional space, discovered by Anthropic's large language model Claude Fable 5, which disproves the conjecture for n > 2

If that won't convince you that LLMs do more than parrot existing ideas, you've got your head in the sand.

xigoi 9 hours ago [-]
Is there any evidence that the discovery was made by Claude and not by Levent Alpöge himself? The only sources listed in the wikipedia article are an X post and a news article that references the post.
tptacek 9 hours ago [-]
You think Alpoge just had the solution for the Jacobian Conjecture in his back pocket, just sort of waiting to deploy at his next gig?
xigoi 9 hours ago [-]
He could have been paid by Anthropic to run a brute-force computer search. The counterexample looks short enough that, given a portion of Anthropic’s computing power and a sufficiently smart algorithm, it could be found by brute force.
tptacek 9 hours ago [-]
Is your premise here that no other mathematician was ever equipped or motivated to do a "brute-force computer search"? Or is it instead that you think Anthropic dedicated an entire data center's worth of compute to the task of speculatively trying to disprove a conjecture that has stood for over 80 years?
Legend2440 9 hours ago [-]
This is a very well studied problem. Well-known mathematicians have spent a lot of effort on it - Yitang Zhang wrote his entire PhD thesis on it back in 1991.

It is deeply unlikely that a random guy at Anthropic just happened to solve it so they could pass it off as the LLM's work.

lloeki 10 hours ago [-]
> If that won't convince you that LLMs do more than parrot existing ideas, you've got your head in the sand.

It doesn't.

In a nearby comment: https://news.ycombinator.com/item?id=48983413

> In mathematics (including information science, CS) there are all sorts of problems that are essentially searches for a solution, and many have the property that the search is computationally difficult, but verifying the solution is relatively cheap. E.g. finding integers such that a^2 + b^2 = c^2 isn't easy, but given a claim that some proposed <a, b, c> satisfies this equation is easy to check. The LLM is like that: it solves a search problem that can be fairly hard.

Funnily enough, one of the attempted solves in the litterature is exploring the problem space in two-dimensional space; the LLM found one in three-dimensional space.

So far we know very little as to why and how it found the solution.

It may very well have been directed to brute force 3D space, or even "elected" to" by expanding the known-failed 2D approach to 3D as pure mimicry.

> It does so unreliably, but if you can cheaply verify the solution, there is a win there.

This circles back to what the LLM advocates are pushing for: build the harness that keeps the agent in check, guardrails all the way because it's driving like a demolition derby.

tptacek 9 hours ago [-]
I think when we're at the point of No True Intellectual Achievementing Smale's Mathematical Problems for the Next Century, everyone's premises have drifted too far apart for discussion to be reasonable.
Legend2440 10 hours ago [-]
You do indeed have your head in the sand, that's for sure.

You are desperately searching for ways to excuse it, to explain how it can't be what it obviously is.

lloeki 9 hours ago [-]
On this matter I'm not searching for excuses, I am reserving my judgement; all we have is a tweet with the counterexample. We don't know how the counterexample was built nor found. It's just not useful to speculate.
perching_aix 12 hours ago [-]
Because it's exceptionally demagogue to anyone with a functioning brain? You know, the thing the dear author makes a big hoopla about people giving up by using these?
SirFatty 10 hours ago [-]
"Because it's exceptionally demagogue to anyone with a functioning brain?"

Oh the irony...

perching_aix 8 hours ago [-]
Yeah, the irony...
8 hours ago [-]
wilg 12 hours ago [-]
It's an esoteric philosophical question that has no truth value either way.
TZubiri 12 hours ago [-]
If anything it's more important to hold. It's easy to hold one position and then falter, there's a pressure to always be with the times and not be 2 years demodé, but simple positions still hold true.

I wrote in the opencode thread that when it came out I put it behind a vm and its own user, and I never allowed it to run outside of it. But I know of people that as soon as they noticed that it worked well like 99% of the time, they let their guard down and give in to YOLO mode. And in orgs I've even seen CEOs treat their agents less like a user/employee/contractor, and try to 'empower' it by giving it ALL the data. Time bomb.

It's like fucking with condoms just the first couple of times. And then simultaneously ditching it and joining the free love movement.

tptacek 12 hours ago [-]
It's not even clear what claim you're trying to make about AI here. "Dangerous", I guess? What does that have to do with its parrotude?
TZubiri 11 hours ago [-]
That just because a problem has existed for years, it doesn't mean the problem is gone or that it's no longer appropriate to make the same warnings and precautions as when it first came out.
tptacek 11 hours ago [-]
I'm not asking whether you can justify your belief that AI is dangerous. I'm asking what it has to do with what species of bird it most resembles. I'm being serious about that. What do I have to learn from the claim that an LLM is a stochastic parrot?
tauroid 12 hours ago [-]
It seems vanishingly rare that people acknowledge the true situation which is that, during training, it really does "think" in that it develops beliefs and marks out precisely chosen trails through its vast and expanding territory. Has a soul, attuned to God, blessed member of the flock, or may as well be.

And then during inference the light goes out and the "agent" staggers randomly like a zombie along those preset paths. Stochastic parrot.

So you and your AGENT.md and your skills files and your harnesses will never make your Claude perceive something that is not in its model checkpoint.

ML experts and neurobiologists free to correct me.

wilg 13 hours ago [-]
Fully 50% of HN front page posts and comments are this now.
hugodan 12 hours ago [-]
[flagged]
dreambuffer 12 hours ago [-]
[flagged]
AgentME 12 hours ago [-]
If a tool getting a URL wrong sometimes was a fatal issue, I would have written off using Google, forums, and my keyboard years ago.
Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact
Rendered at 07:11:18 GMT+0000 (Coordinated Universal Time) with Vercel.