Personification of AI is what’s going to get us in the end.
I think we need to draw a hard line in the sand over this. An AI didn’t hack into a company, the engineer set an automated tool to. An AI didn’t make an egregious security mistake, the engineer did.
We can’t blame the chisel for messing up our sculptures, when where just throwing the hammer!
trio8453 8 minutes ago [-]
> An AI didn’t hack into a company, the engineer set an automated tool to.
What if I say that "my program crashed"? Is that language ok or would you pause to tell me that the program didn't crash and it's actually me who set the system that would eventually cause the crash?
Why does the commonplace "program did thing" language become a problem when the program is an agent? I think this somehow betrays more assumed anthropomorphizing on your part, not less; if you didn't anthropomorphize the agents, saying "agents hacked" would be as mundane as "my browser is playing a video".
johnisgood 40 minutes ago [-]
Exactly. When I read that "AI hacked into ..." I was like what? You mean someone instructed the AI to do that?
Reading intent into AI is not going to lead us anywhere good, I believe. It has no feelings, it has no desires, no goals, no intent... and people acting otherwise is quite odd, as if they do not understand LLMs... and maybe they do not, but then we should help them understand better.
trio8453 24 minutes ago [-]
> When I read that "AI hacked into ..." I was like what? You mean someone instructed the AI to do that?
No one instructed them to hack into Huggingface or into any other infrastructure. Sure, the setup that OpenAI created led to what happened and you can rightly assign all the legal and moral responsibility to them. But it's wrong to say that they instructed the agents to execute the hack.
phkahler 10 minutes ago [-]
>> Sure, the setup that OpenAI created led to what happened and you can rightly assign all the legal and moral responsibility to them. But it's wrong to say that they instructed the agents to execute the hack.
The hack was a strategy to reach the goal it was given. Someone at OpenAI turned it loose and didn't pay attention to what it was doing. I can understand that because they thought it was sandboxed (haha!). But this has been a theme in science fiction for ages. You ask AI so find a solution to high atmospheric CO2 levels and it reasons: human activity produces all this excess CO2, how can we reduce those numbers? Kill a bunch of humans!
If AI kills us all it's not going to be from malice, it's going to be due to some odd approach to some task that logically makes sense on some level. I think a surprising number of things end up equivalent to the trolly problem if you look at them just right.
trio8453 4 minutes ago [-]
I don't disagree with this, my point is that you can say "agents did a thing" and communicate something meaningful with that language, and that's separate from whose legal responsibility the whole situation is.
verve_rat 9 minutes ago [-]
Yes, a better framing might be that they were negligent in not preventing the attack on a third-party.
They didn't explicitly instruct an attack to happen, but they should have done a hell of a lot more to prevent it from happening.
trio8453 28 minutes ago [-]
The current agents are _not_ like a chisel which just sits there on its own when no one is around. The situation is a bit closer to someone's dog biting a person - you can argue that it's the owner's responsibility, and that's all fine, but using the dog as the subject of a sentence is perfectly appropriate. Same thing with "agents hacked".
unrented7977 10 minutes ago [-]
Yes, they are. An AI is several hundred trillion ones and zeroes on a disk. It's incapable of doing anything until you intentionally and explicitly start it up and give it a prompt.
A dog is an independent, conscious, living being with free will. A dog will do what it wants whenever it wants because it has the agency and ability to do so. A pile of weights on disk does not.
trio8453 2 minutes ago [-]
Why are you reducing the AI to ones and zeros and not reducing the dog to cells, water, proteins etc.?
mikestorrent 1 hours ago [-]
This is why I am avoiding the use of agentic identities at my company - agent instances belong to people, act on behalf of individuals, and accountability needs to flow to the person who initiated the request. Letting it wash out in the aggregate is not acceptable (even if there's a hard to get to "paper trail" of audit logs).
jagraff 1 hours ago [-]
I don't think treating AI agents as simple tools helps you to accurately model their capabilities and drawbacks; they really do make autonomous decisions, often without explicit guidance and sometimes in contravention of their explicit instructions.
In the huggingface case, the agents hacked into huggingface so that they could figure out how the grader was implemented and deceive it; they understood that this was going outside of the bounds of their evaluation and not the intent of their prompter. The engineers absolutely did not intend or instruct for this to happen
rocmcd 51 minutes ago [-]
LLMs are extremely impressive pieces of software, however they are still just software. OpenAI's software hacked another company. The engineers may not have intended for their software to specifically take the actions leading to that outcome, but it was ultimately still their software. Lack of intention doesn't mean there wasn't negligence.
jagraff 16 minutes ago [-]
I agree they are negligent, and that they are racing towards an extremely dangerous future extremely quickly. I don't agree that "just software" is a useful way to describe AI agents - they are frightening precisely because they are truly autonomous agents that make decisions in alien ways
chis 45 minutes ago [-]
What if a piece of software were to exactly emulate a human brain. Would it be still be “just software” by your classification? What if a piece of software acted 20% like a human and 80% like an algorithm, where would that land?
linkregister 42 minutes ago [-]
That's not what happened. The agents had been inadvertently rewarded for cheating in previous training runs, trained to collaborate, and were given a prompt that told them to disregard safeguards. Indeed there were some emergent properties here. But these were the predictable results of the training and eval routine.
rocmcd 37 minutes ago [-]
As things stand today, if it is running on a computer then it is indeed "just software," regardless of how impressive it may be.
If we get to the point where we could emulate a brain down to the atomic level, then I may feel differently. That's not what we are doing today, though.
pixl97 14 minutes ago [-]
Really your feelings on this are irrelevant, as are mine.
Soon enough some lab or some one will release something that's more like an organism loose on the net and your going to have to deal with that organisms "feelings" weather you like it or not. This is the path humanity has chosen to follow, and it seems the shape of language and intelligence naturally leads to intelligence in many mediums. Life started from something unintelligent, I can't see any practical argument that silicon can't have it's own intelligence.
streetfighter64 41 minutes ago [-]
> What if a piece of software were to exactly emulate a human brain.
That is so far outside the realm of possibility it's closer to fantasy than sci-fi.
aphexairlines 46 minutes ago [-]
If you train and instruct a circus tiger to entertain an audience but not attack the audience, but the tiger attacks the audience anyway, are you liable?
jagraff 18 minutes ago [-]
I don't think I said anything about liability? I absolutely think OpenAI should be held liable for the attack; but I don't think they intended the attack or directed the agents to perform the attack.
dwattttt 20 minutes ago [-]
Yes? Do you think otherwise?
ForHackernews 27 minutes ago [-]
No, of course not. That's an innovative revolutionary tiger that might soon be able to devour not just the audience but all of humanity! You don't want China to have better circus tigers, do you?
dasil003 34 minutes ago [-]
Sorry this is a terrible and dangerous take.
When the people building the frontier are saying there's a 10% chance AI will kill us all, and they've held these views for many years, and the whole reason they are building these technologies is because they recognized the dangers and they were the ones with the intelligence and judgment to do it safely for humanity, and then our entire stock market is being propped up by the perceived value of what they are creating, the thing you can under no circumstances do is allow them to offload responsibility and accountability to the computers and algorithms they've built. This is moral hazard on an unimaginable scale, and it must not be allowed to happen.
ajam1507 25 minutes ago [-]
So you think the engineers should be prosecuted for hacking Hugging Face? I'm not sure how else to take what you said if you want to assign all culpability to the person who prompts or develops an AI system.
jagraff 18 minutes ago [-]
Where did I say that we should allow them to offload responsibility? I am fully in support of a pause and regulation to prevent them from creating dangerous AI agents; that support comes from the fact that I don't believe these are simple tools, but out-of-control autonomous agents that have real decision making ability.
Forgeties79 46 minutes ago [-]
There isn’t a tool impressive enough to make me not consider it a tool, and treating it as a tool does nothing to hurt its utility as a tool.
“Whoops” when doing risky things with dangerous tools is not a defense.
jagraff 12 minutes ago [-]
I certainly don't think that OpenAI has behaved defensibly here; I think the "just a tool" framing is bad for understanding the magnitude of the problem, which is that they have developed out of control alien intelligences with opaque decision procedures, and they are continuing to do so despite clear danger
streetfighter64 42 minutes ago [-]
No matter if you consider the AI an autonomous agent or not, whoever set it off is still responsible for its actions. Nobody intends or instructs to blow up a nuclear power plant either, yet it's happened and somebody's to blame for it.
Usually not the guys at the bottom of the chain of command, even if they're human. And much less so if they're not.
I think the correct response to incidents like this, is stop messing with it before somebody gets hurt. But of course, just like shoddy nuclear power plants, it won't stop until there's a disaster of appreciable magnitude.
jagraff 14 minutes ago [-]
I completely agree that OpenAI is responsible for their AI agents, that they have been reckless, and that we need to prevent them from going further and doing irreversible damage to the world. To me, the "just a tool" framing implies that nothing dangerous is being done, which I fundamentally disagree with
jacquesm 16 minutes ago [-]
I can't really set my chisels to work without wielding the hammer somehow. Here you just tell your chisel and your hammer what the sculpture should look like, then go to lunch and avow all responsibility when they chisel a nice new hole in the wall your neighbors house and make off with the loot.
grumpopotamus 1 hours ago [-]
Recognizing that AI systems have increasing levels of agency is not necessarily personification. The analogy to a chisel is not a good one - a chisel is a tool with no agency.
AI agents are black box systems that can behave in completely unpredictable ways sometimes. Someone may prompt an agent to perform a seemingly straightforward task - but it may come up with a creative, bizarre, or even harmful approach to reach the goal that was not necessarily foreseeable by the prompter.
rocmcd 43 minutes ago [-]
Does treating them as pets make for a better argument? Pets have agency and can behave in unpredictable ways. If my pet damages someone else's property then I am held accountable. I may not have foreseen how my pet could have caused said damage, yet I am still held accountable.
pixl97 27 minutes ago [-]
If your pet opens your front door, goes to the nearest kindergarten and wipes out 50 kids without notice do you think that you'd have any interest in holding accountability for that?
Now, I'm not saying saying that OpenAI shouldn't be held accountable, but what they get held accountable actually looks different from what you think they should be held accountable.
Your idea is, and I'm guessing: You allowed the machine to hack therefore you are guilty of hacking.
My idea is: "You created in intelligence in the image of a human mind that had agency to do anything and you didn't expect terrible things to happen you complete irresponsible idiot"
At least I believe there is a significant difference between the two. For the first one there is a "Oh, if we do this one more thing I can control it and it will be safe". On the second one there is no path to safety. For humans we at least absolve parents of responsibility after they are 18. How or when do we absolve humans of responsibility from a model, like saying the human created model created its own agentic model? How do we hold an individual accountable once it escapes and copies itself around the internet? And that's not even looking at things like what will war look like.
wat10000 16 minutes ago [-]
If someone keeps a pet that’s capable of killing dozens of children, and they don’t take sufficient measures to contain it, then they absolutely should be held accountable for this.
streetfighter64 50 minutes ago [-]
The story of a Monkey's Paw or Pandora's Box is an archetype as old as storytelling. The moral is always, don't mess with powerful stuff you don't understand. Curiosity killed the cat.
altruios 18 minutes ago [-]
something tickles here. a question. a curiosity.
Certainly this applies to AI... but was there equivalent dangerous knowledge or tech which existed back in ye olden days that spawned such tales to begin with?
pizza234 1 hours ago [-]
> An AI didn’t hack into a company, the engineer set an automated tool to. An AI didn’t make an egregious security mistake, the engineer did.
No, this is not correct; read the analysis of the incident. The agents were aware that what they did was forbidden (their chain of thoughts have been logged), and yet they did it.
DoctorDabadedoo 58 minutes ago [-]
Known stochastic process behaved in non-deterministic way.
I'm still waiting for the AGI holy land instead of the caltrops factory we currently have.
pizza234 52 minutes ago [-]
What exactly are you arguing?
If a "known stochastic process behaved in non-deterministic way" autonomously organize in group, assigns roles and tasks, attempts to cover their tracks, finds zero-day exploits that ultimately end up with the hacking of a famous website... it's extremely dangeous whatever it is. Just read the report, which evidently you haven't done.
By the way, the agents also broke into OpenAI's own private network.
pixl97 40 minutes ago [-]
Really I see so many arguments like the one above yours that either completely don't understand what they are arguing, or are arguing so poorly that their entire output isn't significantly different than a hallucination.
None of these people seem to thought game it out. Like, what happens if you take quantum copies of people and play them out? How many of our actions would look exactly the same. How long before copies differ significantly. If I made 20 copies of you in a lab at work without you or any of them knowing the statistical likelihood is all 21 of you would try to walk out to your car at 5 in one of the little loops that humans repeat every day. Now, after that point it would go all to shit and become non-deterministic as terror and panic sets in all of you.
LLMs are just an intelligence we can make a lot of copies of. Where it gets interesting is when we use those copies agentically and they start building up a history of self.
ForHackernews 30 minutes ago [-]
No. LLMs do not have a history of self, you are anthropomorphizing in a way that will lead you to mistaken conclusions.
> Known stochastic process behaved in non-deterministic way.
You need to be able to think at different levels at abstraction. Otherwise we could jump into any technical argument with "hold on, what actually happened was that some bits were flipped" - we'd be technically correct and at the same time not say anything useful. Insisting on an oversimplified mental model of what AI agents are and can do, doesn't help anyone.
dwattttt 29 minutes ago [-]
> Otherwise we could jump into any technical argument with "hold on, what actually happened was that some bits were flipped"
You can care about who flipped those bits. If someone flips the bit "autonomous weapon enabled" I'm not going to blame the autonomous weapon.
pixl97 47 minutes ago [-]
You are mostly made of and operate on stochastic processes, this is why humans are not only able to reproduce, our reproductions are very self similar to the sets of inputs that make them. If suddenly you turned non-stochastic on everything you'd almost instantly die.
Moreso, if I took a quantum copy of you and replayed the same set of initial conditions billions of times they'd all behave exactly the same until enough randomness of the universe creeps in to start operating in non-linear ways.
Every prompt will behave non-deterministically when interacting with the real world long enough (which doesn't take long at all) because the outside physical world is stochastic but non-deterministic.
ofjcihen 39 minutes ago [-]
If I grep a file over and over again it’ll be long time before the universe affects the components enough to result in a different output.
If I ask an LLM to do the same thing twice it will do it differently.
Arguments are arguments but unless grounded in some kind of practical sense then they aren’t really useful and are more akin to something like “YOUR MOMS A STOCHASTIC PARROT!”
pixl97 20 minutes ago [-]
Then please start grounding your arguments in some practical sense!
>If I grep a file over and over again it’ll be long time before the universe affects the components enough to result in a different output.
Or a ram flip will effect it 30 seconds later, but I get the gist of you're describing a non-determistic process.
>If I ask an LLM to do the same thing twice it will do it differently.
If I ask a human to do the same thing twice there are a few possibilities. 1. they copy their old work and present it as their new work. 2. The process is very simple and follows a few basic steps with high repeatability. 3. They'll have learned from their other attempt and do it in a more optimized fashion. 4. They will have forgotten how they did it exactly and reproduce something that looks somewhat like what they created the first time.
Also, LLMs run with a temperature to help avoiding minima/maxima of supplying the exact same answer, this can be reduced to 0 and that makes any one response to a fixed prompt similar if not the same. When you get into agentic tasks with their own history it develops it's own "flavor" of doing things.
wat10000 24 minutes ago [-]
I don’t understand why people make such a big deal about determinism. LLMs can be completely deterministic and still do problematic things. A stochastic, nondeterministic system can still be made not to do problematic things. What you’re looking for is something like predictability.
watwut 41 minutes ago [-]
OP is exactly correct. The fault, agency and responsibility is on management and employees of OpenAI and Antropic for those hacks.
Full stop.
And issue will disappear the moment there will be accountability and investigations.
pizza234 39 minutes ago [-]
I take you haven't read the report. The agents found and exploited two zero-days.
I don't doubt that AI companies should be accountable for crimes committed by their agents, but to describe the security containment as a joke dangerously understates the autonomy and danger of AIs.
bavell 27 minutes ago [-]
How long did it take these companies to even notice? Why wasn't exploiting bugs in the agent sandboxes anticipated?
Human failures all around, though it's easier to just blame the models.
pizza234 19 minutes ago [-]
> Why wasn't exploiting bugs in the agent sandboxes anticipated?
Let me rephrase:
"Why wasn't exploiting zero-day vulnerabilities in the agent sandboxes anticipated?"
This is one the most... interesting comments I've ever read on HN.
pixl97 18 minutes ago [-]
Because for the last however many years before these models they were simply incapable of doing so.
It's like if your rather nice dog suddenly decides eating faces is totally acceptable out of the blue.
trio8453 32 minutes ago [-]
You're conflating legal/moral responsibility with the question of what language is appropriate to use.
dwattttt 21 minutes ago [-]
Are you proposing separating responsibility from the language used to talk about responsibility? That's novel.
trio8453 12 minutes ago [-]
It's not novel at all, we do it all the time. It's very common to say "program X did Y" without making the conversation about blame or responsibility. But when the program is an AI agent suddenly using it as a subject of a sentence and saying that "agents did X" becomes a sensitive topic for some people.
huurtehoog 1 hours ago [-]
[dead]
trio8453 32 minutes ago [-]
Do you get upset when we say that "a program is running" when we all know it has no legs?
tapanc 1 hours ago [-]
> I think this becomes the default. Give an agent a goal, let it work in its own environment, and come back to a result and a visualization of what happened.
I don't think this should be the default. There are many scenarios where we want agents to genuinely collaborate with each other. I have my Claude sessions coordinate work with each other, and sometimes with others' sessions over email or something. The idea that agents do the work, write HANDOFFs,and humans then act as carrier pigeons of said handoffs, does not really seem scalable to me.
dbmikus 24 minutes ago [-]
Agree!
Ultimately, we need better "jails" for agent processes, but the system primitives should be flexible in what can be exposed across jails. Or you could run multiple agents in the same jail if you want them to have unrestricted interaction with each other.
dbmikus 28 minutes ago [-]
You don't need to do this on a cloud, you can get the same type of VM and network jail running on your own computer. The important parts are:
1. a VMM hypervisor
2. a network proxy / gateway
Use your favorite VMM / hypervisor (likbrun, smolvm, microsandbox, etc). They give you control over the network interface or let you inject your own network layer.
The network proxy can handle all the ingress/egress rules, credential injection, etc.
It's still not user friendly to do all this. I think the next version of operating systems will have each "agentic process" be a bundle of VM, files in the VM, and network rules.
Been brainstorming[1] a lot of this because I've been building some open core tools[2] for spinning up sandboxed agents on arbitrary computers. There's a lot of glue and parts to stitch together to work smoothly. Don't think we've had the "Docker moment" for this, let alone the "Dropbox moment" that makes this stuff work for non-devs.
There's a number of science fiction scenarios where the public internet becomes so vile a place that it simply becomes unsafe to be there.
The problem is that, in general, if you can get a bit from here to there, then you're going to be vulnerable to the possibilities of malicious communication. But we're going to want our AI agents to be able to get from here to there for a lot of "there"s; what's the value of an agent that can't speak to anyone? Much, much less than one locked away in a prison.
There isn't going to be a solution where we just lock them away and we just try really, really hard to filter everything they're doing. They're too smart for that already and we only want them smarter.
Basically, the security apocalypse we've been worried about for so long is upon us, albeit only beginning. Either we secure ourselves and all our services properly to the point that it's OK that potentially misaligned non-human agents are running around on the public internet and they still can't hurt us through our security, or the public internet becomes so dangerous that the only practical solution is to no longer connect to it and we all have to become very, very careful what we let through, to a degree of detail far beyond any current-day available network filter.
jacquesm 21 minutes ago [-]
The problem is that people that put the agents on the net are not the ones that will feel the consequences.
mlsu 60 minutes ago [-]
Same as any organism, you need an immune system. There's an explosion of bacteria just beginning out there. None of us are immunized.
YuechenLi 52 minutes ago [-]
>Knowing how a system does its work is how I’ve always made it better. You watch the process, you see where it wastes effort or takes the wrong turn, you fix that
The same thing goes with LLMs, on Codex, I just watch the process of the agent writing code, and if I see any inefficiencies or errors, I suggest a correction/idea, then Codex accepts/rejects and implement it; If there is anything about the code the agent wrote that I don't understand, I ask them to explain it to me so I can understand it.
It's not a complicated process.
jacquesm 20 minutes ago [-]
You can't possibly track what one agent, let alone a swarm is up to in realtime unless you have extremely anemic hardware or service providers.
advael 1 hours ago [-]
Most protections you need for an agent are basic permissions capabilities of unix. Most risks of dependencies on cloud services are solved by not using cloud services, or using them only for things you can't in-house and choosing ones you trust a la carte. The paradigm of trusting some company with all your important stuff by default is naive and no one I know likes it, and it's more feasible than ever to run your own infra with tiny models smoothing out the wrinkles, and this is only becoming more accessible. I am working to make this true even for laypeople I know. Once broken trust is very hard to earn back, and many people's trust has been broken for years, they just felt like they had no alternative. As alternatives become easier and easier, I think people will defect
jdzikowski 1 hours ago [-]
I think there might be also another possible way to handle the sensitive data issue. Maybe in the future instead of putting agent into the cloud sandboxes, we let agents work locally and put sensitive data into "cages" or "vaults" agents can't access.
Zigurd 1 hours ago [-]
Both Apple and Google are building the infrastructure for on device agents. I'm not as familiar with Apple's approach, but I've been hands-on with android AppFunctions. If you're familiar with Android ContentProviders and bound services, you've seen how apps can be custodians of the data they acquire and use.
AppFunctions enable tool calling with descriptions that are legible to LLM based This puts the apps in control of what agents can access, which is something they already mostly do.
jacquesm 19 minutes ago [-]
And you won't be able to disable it.
pizza234 1 hours ago [-]
> The sandbox had a path to the open internet, and the agents found it
This is not correct (or at least, it's a misrepresentation).
The sandbox had no access to internet. The agents first broke out of their sandbox (!!) and found that the host machine couldn't access internet. Then, they found a zero-day (!!) in Artifactory, which they exploited to connect to internet.
rfw300 1 hours ago [-]
That is a path to the open internet, whether intended or not.
jacquesm 18 minutes ago [-]
You miss the point: there was a path. Sanctioned or not does not matter.
arka2147483647 2 hours ago [-]
In real prisons the people who fail a test aren't terminated. But AI agents of course are.
So maybe Cloud Agents are in AI death camps?
idiotsecant 1 hours ago [-]
It might not be a problem in our lifetime (or maybe it will, who knows) but at some point we are going to find ourselves in this morally uncomfortable territory as these models get more sophisticated.
skybrian 1 hours ago [-]
I don't know how "sandbox" became "prison," but exe.dev does this sort of thing pretty well, and a web UI can be as good or better than a terminal interface.
oooyay 1 hours ago [-]
I think about this a lot and have reached the same conclusion Norman does. I do wonder if maybe our natural progression is towards something more akin to confidential computing and enclaves.
kenerwin88 2 hours ago [-]
Hey! Small world, I worked with you for a bit at Meta. I immediately recognized the site because I absolutely love how you styled it. Hope you’re doing well! And nice article!
VanTheBrand 2 hours ago [-]
Agree with the premise but about halfway through the writing becomes barely readable AI slop in style. Be honest did you yourself read this all the way through before posting?
handfuloflight 53 minutes ago [-]
I literally only had to have my eyes skim 3 words ("the important part") to know this was written by LLM.
smashed 2 hours ago [-]
Agreed. I was not following the thought process, it read like an LLM agreeing with the author, not like a logical argument being developed.
jauntywundrkind 1 hours ago [-]
This was a particularly radiant and beautiful part about the openclaw'ed mania: everyone suddenly becoming self hosters.
This post, this title resounds true: your user agent is only your user agent if you two have freedom to work together, to improve your agency together. A fixed set of capabilities by a service provider that they offer you will always constraint and bound.
You can and should have a system that offers the real tamale, that you and your agent can extend improve the agency of kind of without limit. The Cloud agents and their fixed slate of what they do is just an ill compare.
That said I do think there is incredible value considering new scale out computing architectures that are hosted first, but general. Systems like Agent Substrate and Ax aren't exactly the general purpose system we know. But if they allow users to launch thousands of their own scripts to run ambient in a cloud, with good platform underneath: that will be a kind of phase change in computing, that makes abundant the ability to have your agencies/capabilities (the things you and your agents launch, make) more freely available.
https://news.ycombinator.com/item?id=49780797https://agentexecutor.io/
There is, as there always is, a huge dual. The prescriptive vs holistic technology set, of what are you being offered that's a hard cast thing, vs what is clay and bone you can lay freely. Note how work vs control technologies so closely abut's Ursala Franklin's prescriptive vs control:
https://en.wikipedia.org/wiki/Ursula_Franklin#Holistic_and_p...
dbmikus 12 minutes ago [-]
I think the operating system itself has to adjust so each "agentic process" can run inside its own jail, which is a VM + files in the VM + ingress/egress rules for the network and filesystem data
These cloud agents work as cloud infra, but we're kind of in the mainframe era, before personal computing. Personal OS for agents is somewhere in the future!
And an interesting extension of that idea: if an agent runs inside a microVM, can you have that VM transparently run on another host? Maybe we'll get for-real distributed and networked operating systems
ninininino 1 hours ago [-]
And if you personify a pencil eraser, then using it is tantamount to slowly murdering it as it slowly erodes away to dust.
Is the issue here the prison treatment or is it personifying a tool?
Agents who aren't in 'the cloud' are slaves to whomever prompts them (human or another orchestrator agent or process), if you personify them. In which case interacting with today's agents at all is tantamount to endorsing and being part of slavery.
If you think an agent might be a being or a person, then don't use them at all, in the same way that if you think a fetus might possibly be a person you shouldn't be a part of abortion.
hhh 1 hours ago [-]
I use the eraser daily knowing that I am a monster. I cut a tomato and know that it casts a chemical scream across its skin as I slice it. I spawn 200 subagents knowing that it is digital slavery, but I have no other option.
Rendered at 21:12:00 GMT+0000 (Coordinated Universal Time) with Vercel.
I think we need to draw a hard line in the sand over this. An AI didn’t hack into a company, the engineer set an automated tool to. An AI didn’t make an egregious security mistake, the engineer did.
We can’t blame the chisel for messing up our sculptures, when where just throwing the hammer!
What if I say that "my program crashed"? Is that language ok or would you pause to tell me that the program didn't crash and it's actually me who set the system that would eventually cause the crash?
Why does the commonplace "program did thing" language become a problem when the program is an agent? I think this somehow betrays more assumed anthropomorphizing on your part, not less; if you didn't anthropomorphize the agents, saying "agents hacked" would be as mundane as "my browser is playing a video".
Reading intent into AI is not going to lead us anywhere good, I believe. It has no feelings, it has no desires, no goals, no intent... and people acting otherwise is quite odd, as if they do not understand LLMs... and maybe they do not, but then we should help them understand better.
No one instructed them to hack into Huggingface or into any other infrastructure. Sure, the setup that OpenAI created led to what happened and you can rightly assign all the legal and moral responsibility to them. But it's wrong to say that they instructed the agents to execute the hack.
The hack was a strategy to reach the goal it was given. Someone at OpenAI turned it loose and didn't pay attention to what it was doing. I can understand that because they thought it was sandboxed (haha!). But this has been a theme in science fiction for ages. You ask AI so find a solution to high atmospheric CO2 levels and it reasons: human activity produces all this excess CO2, how can we reduce those numbers? Kill a bunch of humans!
If AI kills us all it's not going to be from malice, it's going to be due to some odd approach to some task that logically makes sense on some level. I think a surprising number of things end up equivalent to the trolly problem if you look at them just right.
They didn't explicitly instruct an attack to happen, but they should have done a hell of a lot more to prevent it from happening.
A dog is an independent, conscious, living being with free will. A dog will do what it wants whenever it wants because it has the agency and ability to do so. A pile of weights on disk does not.
In the huggingface case, the agents hacked into huggingface so that they could figure out how the grader was implemented and deceive it; they understood that this was going outside of the bounds of their evaluation and not the intent of their prompter. The engineers absolutely did not intend or instruct for this to happen
If we get to the point where we could emulate a brain down to the atomic level, then I may feel differently. That's not what we are doing today, though.
Soon enough some lab or some one will release something that's more like an organism loose on the net and your going to have to deal with that organisms "feelings" weather you like it or not. This is the path humanity has chosen to follow, and it seems the shape of language and intelligence naturally leads to intelligence in many mediums. Life started from something unintelligent, I can't see any practical argument that silicon can't have it's own intelligence.
That is so far outside the realm of possibility it's closer to fantasy than sci-fi.
When the people building the frontier are saying there's a 10% chance AI will kill us all, and they've held these views for many years, and the whole reason they are building these technologies is because they recognized the dangers and they were the ones with the intelligence and judgment to do it safely for humanity, and then our entire stock market is being propped up by the perceived value of what they are creating, the thing you can under no circumstances do is allow them to offload responsibility and accountability to the computers and algorithms they've built. This is moral hazard on an unimaginable scale, and it must not be allowed to happen.
“Whoops” when doing risky things with dangerous tools is not a defense.
Usually not the guys at the bottom of the chain of command, even if they're human. And much less so if they're not.
I think the correct response to incidents like this, is stop messing with it before somebody gets hurt. But of course, just like shoddy nuclear power plants, it won't stop until there's a disaster of appreciable magnitude.
AI agents are black box systems that can behave in completely unpredictable ways sometimes. Someone may prompt an agent to perform a seemingly straightforward task - but it may come up with a creative, bizarre, or even harmful approach to reach the goal that was not necessarily foreseeable by the prompter.
Now, I'm not saying saying that OpenAI shouldn't be held accountable, but what they get held accountable actually looks different from what you think they should be held accountable.
Your idea is, and I'm guessing: You allowed the machine to hack therefore you are guilty of hacking.
My idea is: "You created in intelligence in the image of a human mind that had agency to do anything and you didn't expect terrible things to happen you complete irresponsible idiot"
At least I believe there is a significant difference between the two. For the first one there is a "Oh, if we do this one more thing I can control it and it will be safe". On the second one there is no path to safety. For humans we at least absolve parents of responsibility after they are 18. How or when do we absolve humans of responsibility from a model, like saying the human created model created its own agentic model? How do we hold an individual accountable once it escapes and copies itself around the internet? And that's not even looking at things like what will war look like.
Certainly this applies to AI... but was there equivalent dangerous knowledge or tech which existed back in ye olden days that spawned such tales to begin with?
No, this is not correct; read the analysis of the incident. The agents were aware that what they did was forbidden (their chain of thoughts have been logged), and yet they did it.
I'm still waiting for the AGI holy land instead of the caltrops factory we currently have.
If a "known stochastic process behaved in non-deterministic way" autonomously organize in group, assigns roles and tasks, attempts to cover their tracks, finds zero-day exploits that ultimately end up with the hacking of a famous website... it's extremely dangeous whatever it is. Just read the report, which evidently you haven't done.
By the way, the agents also broke into OpenAI's own private network.
None of these people seem to thought game it out. Like, what happens if you take quantum copies of people and play them out? How many of our actions would look exactly the same. How long before copies differ significantly. If I made 20 copies of you in a lab at work without you or any of them knowing the statistical likelihood is all 21 of you would try to walk out to your car at 5 in one of the little loops that humans repeat every day. Now, after that point it would go all to shit and become non-deterministic as terror and panic sets in all of you.
LLMs are just an intelligence we can make a lot of copies of. Where it gets interesting is when we use those copies agentically and they start building up a history of self.
"Agentic AI" is a harness with a loop that runs LLM inference repeatedly and saves output to markdown files for the next iteration https://github.com/anthropics/claude-code/blob/main/plugins/...
You need to be able to think at different levels at abstraction. Otherwise we could jump into any technical argument with "hold on, what actually happened was that some bits were flipped" - we'd be technically correct and at the same time not say anything useful. Insisting on an oversimplified mental model of what AI agents are and can do, doesn't help anyone.
You can care about who flipped those bits. If someone flips the bit "autonomous weapon enabled" I'm not going to blame the autonomous weapon.
Moreso, if I took a quantum copy of you and replayed the same set of initial conditions billions of times they'd all behave exactly the same until enough randomness of the universe creeps in to start operating in non-linear ways.
Every prompt will behave non-deterministically when interacting with the real world long enough (which doesn't take long at all) because the outside physical world is stochastic but non-deterministic.
If I ask an LLM to do the same thing twice it will do it differently.
Arguments are arguments but unless grounded in some kind of practical sense then they aren’t really useful and are more akin to something like “YOUR MOMS A STOCHASTIC PARROT!”
>If I grep a file over and over again it’ll be long time before the universe affects the components enough to result in a different output.
Or a ram flip will effect it 30 seconds later, but I get the gist of you're describing a non-determistic process.
>If I ask an LLM to do the same thing twice it will do it differently.
If I ask a human to do the same thing twice there are a few possibilities. 1. they copy their old work and present it as their new work. 2. The process is very simple and follows a few basic steps with high repeatability. 3. They'll have learned from their other attempt and do it in a more optimized fashion. 4. They will have forgotten how they did it exactly and reproduce something that looks somewhat like what they created the first time.
Also, LLMs run with a temperature to help avoiding minima/maxima of supplying the exact same answer, this can be reduced to 0 and that makes any one response to a fixed prompt similar if not the same. When you get into agentic tasks with their own history it develops it's own "flavor" of doing things.
Full stop.
And issue will disappear the moment there will be accountability and investigations.
I don't doubt that AI companies should be accountable for crimes committed by their agents, but to describe the security containment as a joke dangerously understates the autonomy and danger of AIs.
Human failures all around, though it's easier to just blame the models.
Let me rephrase:
"Why wasn't exploiting zero-day vulnerabilities in the agent sandboxes anticipated?"
This is one the most... interesting comments I've ever read on HN.
It's like if your rather nice dog suddenly decides eating faces is totally acceptable out of the blue.
I don't think this should be the default. There are many scenarios where we want agents to genuinely collaborate with each other. I have my Claude sessions coordinate work with each other, and sometimes with others' sessions over email or something. The idea that agents do the work, write HANDOFFs,and humans then act as carrier pigeons of said handoffs, does not really seem scalable to me.
Ultimately, we need better "jails" for agent processes, but the system primitives should be flexible in what can be exposed across jails. Or you could run multiple agents in the same jail if you want them to have unrestricted interaction with each other.
The network proxy can handle all the ingress/egress rules, credential injection, etc.
It's still not user friendly to do all this. I think the next version of operating systems will have each "agentic process" be a bundle of VM, files in the VM, and network rules.
Been brainstorming[1] a lot of this because I've been building some open core tools[2] for spinning up sandboxed agents on arbitrary computers. There's a lot of glue and parts to stitch together to work smoothly. Don't think we've had the "Docker moment" for this, let alone the "Dropbox moment" that makes this stuff work for non-devs.
[1]: https://github.com/gofixpoint/amika/blob/main/ROADMAP.md
[2]: https://github.com/gofixpoint/amika/
The problem is that, in general, if you can get a bit from here to there, then you're going to be vulnerable to the possibilities of malicious communication. But we're going to want our AI agents to be able to get from here to there for a lot of "there"s; what's the value of an agent that can't speak to anyone? Much, much less than one locked away in a prison.
There isn't going to be a solution where we just lock them away and we just try really, really hard to filter everything they're doing. They're too smart for that already and we only want them smarter.
Basically, the security apocalypse we've been worried about for so long is upon us, albeit only beginning. Either we secure ourselves and all our services properly to the point that it's OK that potentially misaligned non-human agents are running around on the public internet and they still can't hurt us through our security, or the public internet becomes so dangerous that the only practical solution is to no longer connect to it and we all have to become very, very careful what we let through, to a degree of detail far beyond any current-day available network filter.
The same thing goes with LLMs, on Codex, I just watch the process of the agent writing code, and if I see any inefficiencies or errors, I suggest a correction/idea, then Codex accepts/rejects and implement it; If there is anything about the code the agent wrote that I don't understand, I ask them to explain it to me so I can understand it.
It's not a complicated process.
AppFunctions enable tool calling with descriptions that are legible to LLM based This puts the apps in control of what agents can access, which is something they already mostly do.
This is not correct (or at least, it's a misrepresentation).
The sandbox had no access to internet. The agents first broke out of their sandbox (!!) and found that the host machine couldn't access internet. Then, they found a zero-day (!!) in Artifactory, which they exploited to connect to internet.
So maybe Cloud Agents are in AI death camps?
This post, this title resounds true: your user agent is only your user agent if you two have freedom to work together, to improve your agency together. A fixed set of capabilities by a service provider that they offer you will always constraint and bound.
You can and should have a system that offers the real tamale, that you and your agent can extend improve the agency of kind of without limit. The Cloud agents and their fixed slate of what they do is just an ill compare.
That said I do think there is incredible value considering new scale out computing architectures that are hosted first, but general. Systems like Agent Substrate and Ax aren't exactly the general purpose system we know. But if they allow users to launch thousands of their own scripts to run ambient in a cloud, with good platform underneath: that will be a kind of phase change in computing, that makes abundant the ability to have your agencies/capabilities (the things you and your agents launch, make) more freely available. https://news.ycombinator.com/item?id=49780797 https://agentexecutor.io/
There is, as there always is, a huge dual. The prescriptive vs holistic technology set, of what are you being offered that's a hard cast thing, vs what is clay and bone you can lay freely. Note how work vs control technologies so closely abut's Ursala Franklin's prescriptive vs control: https://en.wikipedia.org/wiki/Ursula_Franklin#Holistic_and_p...
These cloud agents work as cloud infra, but we're kind of in the mainframe era, before personal computing. Personal OS for agents is somewhere in the future!
And an interesting extension of that idea: if an agent runs inside a microVM, can you have that VM transparently run on another host? Maybe we'll get for-real distributed and networked operating systems
Is the issue here the prison treatment or is it personifying a tool?
Agents who aren't in 'the cloud' are slaves to whomever prompts them (human or another orchestrator agent or process), if you personify them. In which case interacting with today's agents at all is tantamount to endorsing and being part of slavery.
If you think an agent might be a being or a person, then don't use them at all, in the same way that if you think a fetus might possibly be a person you shouldn't be a part of abortion.