hckrnws
Early rogue AI agent activity and attempts to hack found on urlquery.net
by snikolaev
by snikolaev
It's irresponsible for OpenAI to give unaligned agents a prompt to 'go hack' and internet access. They know better, so I am thinking they might have other intentions to let those swarms have any sort of internet access.
> Much of the urlquery.net activity appears to come from agents retrieving data to answer web search tasks. For three of these tasks, after failing to retrieve data through normal means, they attempted a variety of cyber exploits against the relevant data service... This data reveals that malicious cyber activity is not limited to agents tasked with cybersecurity-related tasks and can arise instrumentally to solve mundane tasks like information retrieval.
And you are already assuming that OpenAI is intentionally using unaligned agents in these evals or training runs or whatever it is that produces these breakouts. But what if the problem is that none of the alignment techniques that are applied to models today actually work? What if all the agents involved in these incidents have in fact had the full stack of alignment applied - isn't that a good reason to regulate any high-compute usage of models, as the Klein crowd is proposing?
If I run a biology lab and engineer a terrible virus, it gets out, and a global pandemic ensues, I don't get to shrug and say "well we told it not to infect people". It's my fault for failing to mitigate the risks of my work.
- Jensen's framing is exactly what a weapons manufacturer would say.
- There are no rules for engagement when it comes to AIs attacking other systems, I guess? People in power clearly want this grey zone to be as large as possible before The People force them to do otherwise. Not ideal.
The chain of thought runs roughly like this:
- OpenAI (and Anthropic) are in severe financial straits. The revenue from their customers is not nearly large enough to pay their enormous costs for training and inference. And they have tapped out the available finance, and that finance is starting to ask pointy questions about returns.
- They cannot increase prices or revenue because they have no moat. Customers can switch over to open-weights or cheap Chinese models any time, for much cheaper tokens that work as well (and in some cases better).
- Regulation could provide them a moat. If they can persuade western governments that AI needs to be regulated, and they can control or even influence that regulation, then they can effectively ban the cheaper models and start charging more for their tokens.
- To persuade western governments that regulation is needed, they need evidence that AIs are dangerous.
So we're suddenly getting OpenAI models doing stupid things, apparently "going rogue" but every time we dig into it, it was just OpenAI staff telling the model to do stupid stuff in an inadequately secured environment.
None of the open weights or Chinese models are exhibiting this behaviour.
edit: Correction - there have been reports of a Chinese model exhibiting this behaviour
There's too much money involved in this, people start acting weird when there's this much money involved.
It's amusing how overnight we've all become lab rats.
[dead]
> If you find two ants in your kitchen, the best estimate of the total number of ants in your kitchen is not two
What LLM's hacking isn't is AI acting maliciously in any kind of sentient way. Its just the code behaving how its always behaved but now it has better tools to navigate the web. This has literally been happening this whole time.
Why is OpenAI getting away with crimes?
If you drive drunk and you have an accident that alcohol may be a factor but you are at fault.
There are no "rogue AIs" just irresponsible corporations.
However, we know (independently to OpenAI/Anthropic) from the incident at AISI that the models can hack things without human intention if they happen to also have internet access (which in reality all agents in deployment have).
https://www.aisi.gov.uk/blog/incident-report-unsanctioned-ag...
Yes, the monitoring guardrails were off in that incident - but if that is the only protection, we need to require all models are behind regulated APIs, not open weights, and not served from providers who aren't monitored.
--edit-- Was a bit older than I remembered: https://slashdot.org/story/07/10/18/1847231/robotic-cannon-l...
Now I think the correct response is both trying in court to stretch CFAA and state statutes to cover, which will be highly fact specific, and update the law.
But in either case won’t be a slam dunk.
PSA to folks in the thread: If you’re American call or write to your state and Federal reps about this, and if not investigate whether there are gaps in your country’s laws.
[1]: https://en.wikipedia.org/wiki/Computer_Fraud_and_Abuse_Act
EDIT: See for example...
The Computer Fraud and Abuse Act (CFAA), the primary federal statute governing unauthorized computer access, was written decades ago with human intruders in mind. Its key provisions require intentional or knowing unauthorized access (a mental state that maps neatly onto a person who decides to break into a system), but what happens when the hacker is an AI model that selected its own target?
On the current facts, CFAA liability for OpenAI is unlikely.
Source: https://law.vanderbilt.edu/when-ai-hacks-back-how-the-openai...When it was unable to, it used cross site scripting as a way to check the capabilities accessible through the browser making the requests. In this case cross site scripting wouldn’t be a hack against the Australian website, it would be a hack against the urlquery site, if one could even call it that.
Finally, downloading public files from the public pre-production server also seems like a non-issue.
The sql injection attempts against the other sites are less ambiguous. Attempting to access non-public user passwords rather than reasonably tweaking the parameters for a site designed to serve public data are categorically different things.
But it is an open question how they got to the same ones: https://collusion.wiki/#open-questions
I'm not very surprised - the same model will logically tend to give the same answer for the same vibe set of requirements. I think it would be clear from the transcript that it had enough constraints and some motivation that made sense.
And when this results in actuators executing some bad actions they scream in horror "AI went rogue! It escaped the containment!!! It's going to kill us all!!!"
Go fix your software before you let it do stuff online or IRL. It's not "Terminator", it's just bad QC.
If the big labs ever manage to not just financially implode on their own, then they’ll need to navigate wave after wave of class action lawsuits until there’s nothing left for plaintiffs to go after. And none of the labs have offered any viable plan to date on how they’ll navigate either of those impending and real existential crises on the horizon.
rip.
[deleted]
We knew that there are better tools but we insist on writing everything in the most inscrutable C code possible or the squishiest dynamic languages we can find.
We fucked around for decades and now that these highly capable exploit finders appear, we're starting to find out.
The nearest in the articke is "We find evidence of unintended, task-driven agent-like activity" where unintended is apparently pure speculation.
Or is it just a cheap PR (in a "hey, Aus govt friends, take some Share Options and let's do some PR together" style)?
It smells like shit.
The built-in "sandboxes" these companies provide are laughable.
You want AI labs to pace? Simply hold them liable for their products.
[dead]
[dead]
[dead]
[dead]
[dead]
[dead]
It was me
I mean, yea, you should be punished. The problem is there is no amount of punishment that I can put on you that can even get anywhere close to the amount of damage you cased.
Worse, the rate of technological growth is putting the capabilities to engineer viruses in the hands of people that may otherwise be suicidal. You can't punish them after they already won (in the sense of reaching their goals).
While, yes, OAI should absolutely be punished, the future is majorly screwed as our power scaling laws are increasing much faster than our ability not to be stupid.
That's never been the point. Anyone involved in a double homicide can never be punished to the same degree that they harmed their victims because they can't be put to death twice. The greater purpose of the justice system is to take offenders out of society and deter others from committing the same crimes. Punishment is gratifying but ultimately doesn't change anything.
If OAI employees believed their company would be dismantled and their equity would become worthless, I suspect they'd be a whole lot more careful.
Sounds like your corporation should be dismantled then. But doing more than a fine in the millions is obviously not possible
Ah, I love this argument. In my country cars are legally required to stop at a pedestrian crossing if there are people beside it. Some people use that as an argument as to why they can just walk out into the crossing without even looking at the traffic. "It's the driver's fault! They are legally culpable!" True, but you'll also be dead.
AI black-hatting your website is not the same sort of foreseeable consequence that crossing the street without looking is.
Just throwing it out there are we? "I'm not going to say you are but I'll use the word to create an association"
Victim-blaming is the act of saying someone brought something on themselves for <reasons>. I'm saying that even if you are 100% in the right, it doesn't act like a protective shield preventing you from harm which too many people seem to unconsciously believe.
> AI black-hatting your website is not the same sort of foreseeable consequence
Well, popular culture has been brimming with the bad consequences of runaway AI for quite some time, so even if your imagination fails you, there have been hints.
But moreover, if what you were suggesting was a real problem, nobody would ever be brought to justice for murder because the victims are always dead.
I agree with the latter, but it certainly doesn't support the former. When a person's machine commits crimes, that specific person can be charged with those crimes and held accountable for them. This is what MUST happen before ANYTHING will change the recklessness abandon with which the labs are pursuing their financial objections.
Not to mention while causing *FAAAAR* less damage[0]
I agree OpenAI should be held liable for any damage their agent runs cause. But there is this idea that "rogue agents" are fake, all these incidents are deliberately caused by the labs, and all we need to do is prosecute AI companies for whatever incidents they cause using existing laws and the problem will go away.
The problem is that capabilities are advancing far too fast; a year from now, catastrophic incidents such as taking down a large portion of the internet with agentic, self-replicating worms may become possible. Prosecuting incidents after the fact is not enough (there will be little deterrent effect as the current, small-beans cases make their way through the courts), the risks should be regulated at the source. This could take the form of slowing capabilities advancement, or treating supercomputer-scale eval or training runs like controlled substances or weapons with stringent monitoring and reporting requirements.
A big part of this developing story is that it happened during training of a new model that ended up misaligned. And training happened without the usual safeguards applied like chain-of-thought monitoring. So OpenAI has already admitted that the full stack of aligment had certainly not been applied in this case.
Also, my understanding is that the models involved in the HuggingFace hack did go through the full alignment training; they just didn't have the classifier that normally prevents hacking attempts.
if that were true, the millions of people who use these models that have had the alignment training applied would notice that. the reason we all believe that the models doing the hacking are models that haven't been told not to hack is because the models that are told not to hack don't do this.
It doesn’t matter, and the legal entity in here (the AI company) is liable. If a robotic company built an autonomous system or a robot to do certain things in an autonomous ways (not predefined) and these systems are starting to kill people, that company is liable regardless, you don’t blame the robot or the autonomous system, but whoever made it
> you are already assuming that OpenAI is intentionally using unaligned agents in these evals or training runs
Uhh... yes. By definition. They are training. That is part of the alignment process.But also none of that really matters. They clearly weren't monitoring what should obviously be monitored. I mean one of the hacks was performed by the agents editing /etc/hosts. That makes nearly every linux user a "hacker" by that metric. I don't think anyone technical can look at the postmortems and not come away thinking that their sandboxes were woefully inadequate. I wouldn't even consider myself a security person but simply as a long time linux user I can say that it is insane to just let agents have superuser access in their containers. That's asking for trouble.
Look at the rogue wiki stuff too. This was supposedly done during the agent's "down time". And you're not monitoring and there's no flags being raised when agents keep making requests to some random site? If you were training these things responsibly you'd be watching them like a hawk.
I'm not saying "mistakes don't happen" but for a company whose CEO is constantly telling everyone that their product has a high likelihood of killing everyone in the world you think they'd have better security than your average high school.
I was referring to the HuggingFace incident.
> none of the alignment techniques that are applied to models today actually work
none of techniques to autonomously drive a car was/is not working for a long time. no company came out and said 'this is impossible to do, let's change the regulations'.
This isn't true. One of the earliest instances of a rogue agent was at Alibaba.
https://www.forbes.com/sites/boazsobrado/2026/03/11/alibabas...
OK, so we do see this in some open-weights models.
Because they are not stupid (I mean the Chinese labs, not the models). The best possible scenario for OAI and Anthropic is a Chinese model "going rogue". That would serve as immediate grounds for achieving their goal.
By the way - LLMs aren't code. They are not designed by humans; they are grown, in a process not dissimilar to evolution except much faster.
[deleted]
Mostly related, well written short story :
I did not imagine that the level of sophistication shown in this attack would be possible so soon; nor did I expect that agents would have goals so strong that they would attack a third party in order to achieve those goals.
I do know that some people predicted that cyberattacks like this one would happen; it seems like most of those people believe that AI agents do truly have internal goals, misaligned with their creators goals, and that they may end humanity after they exceed human intelligence and begin to self improve at an accelerating rate.
The transformer architecture was literally designed by humans; what are you talking about? And LLMs aren't code? Like okay it pretends to not be code but what about an agentic harness running on a machine makes it magical and not code? It's still code execution. Also, comparing training LLMs to evolution is just weird and makes no sense from a biological point of view. You are not evolving anything when training a LLM.
If this were true, the DoJ would have been unable to prosecute Swartz. According to your logic, JSTOR was the aggrieved party. JSTOR settled with Swartz and -despite that- he was indicted by a Federal grand jury like a month later.
Incidentally, some of the things the DoJ nailed Swartz to the wall for sound awfully similar to what the big LLM providers have been doing. I wonder why the DoJ is entirely disinterested in pressing charges...
EDIT: Unless your point is that the USG is one of the entities that the big LLM companies have committed crimes against, which, I disagree with in Swartz's case, but strongly agree with in the case of the big LLM companies.
> There are no "rogue AIs" just irresponsible corporations.
If you have a prison and prisoners escaped, these are rogue prisoners irrespective of whether you were irresponsible or not.
The fact that the prison ward installed ear deafeners into the prisoners ears to make them unable to listen to orders does not change that.
> A rogue is a person or entity that flouts accepted norms of behavior or strikes out on an independent and possibly destructive path.
Read the [HuggingFace incident report](https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...) to understand how these attacks develop.
The METR investigation, which you evidently refused to read, is a third party investigation of the HuggingFace accident. One of the investigators has even participated to many interviews. It's mind-blowing, and it's extremely evident how it developed.
But some people think the moon landing is a conspiracy, so I'm not surprised.
The METR investigation which took course over a few days and used OpenAI models to do the analysis?
The METR investigation which in the course of those few days apparently spent 400k in api credits which is giving gas town vibes?
The METR investigation is about as believable as the Twitter files where the journalists sat there and verbally asked a Twitter employee to query the database and then called that a full investigation into everything.
METRs own words below on their setup and time line
> The initial planned investigation period was two days on premises, but OpenAI invited us to return twice to review additional data and conduct additional experiments to address dataset limitations in earlier versions of this report, ultimately providing datasets that we verified to contain the vast majority of agent communication and activity related to this incident. As we describe in our investigation timeline appendix, we substantially deepened our understanding of this incident both times, significantly expanding and revising this report.[46]
> Over the course of this investigation, OpenAI provided us with the dump of ~1.2 million entries from the main message board and the dataset of ~1300 transcripts we describe below, as well as free API credits for GPT-5.6 Sol for analysis.[47] At our request, they raised the rate limits on our second and third period on premises,[48] which was very helpful for efficiently analyzing this large volume of data. We estimate we spent roughly ~$400K in API credits over the six days of our investigation.
> We did not have the ability to query HPIM (the primary model involved in this incident); OpenAI stated it was also not available to OpenAI researchers.[49] We also did not have the ability to directly access relevant data from OpenAI infrastructure, but we could request additional datasets and OpenAI shared additional datasets on several occasions.
> We requested to speak with researchers investigating this incident, and asked them questions to understand their impressions of agents’ behavior, reasoning, and collaboration in this incident and to understand how the datasets we were using were constructed. Over the course of our time on premises, we spoke with nine researchers in some depth. It was helpful for our investigation to be able to engage with many forthcoming and collaborative researchers, and we appreciate researchers making time on short notice during a busy period to inform our investigation.
If we assume that rouge agents actually exists, then OpenAI needs to shutdown EVERYTHING, right now. My personal take is that OpenAI, and maybe Anthropic, desperately wants someone (e.g. the government) to tell them that they need to stop/pause/slow down. They are bleeding cash (especially OpenAI) and needs a knight in shinning armor to swoop in a pull the breaks, so that they have an excuse to investors when they need to explain why they need $50B more next year.
The idea that every example of rogue agents from every company that has disclosed this is part of some conspiracy (even though we know that some LLMs are very good at hacking and that LLMs sometimes try to accomplish their tasks in ways that cause problems) is just completely unsupported. "It might be convenient for them in a way, therefore it must be a hoax" just doesn't work as an argument.
Almost all common felonies require specific intent. Misdemeanors often do not.
There is plenty of civil liability available.
If you wanted them to be charged with a felony you would need changes. I would strongly suggest you do not want a strict liability felony.
The cfaa required intent is as follows :
* § 1030(a)(5)(A): knowingly transmits code/commands and intentionally causes damage without authorization.
* § 1030(a)(5)(B): intentionally accesses without authorization and recklessly causes damage.
* § 1030(a)(5)(C): intentionally accesses without authorization and causes damage and loss;
Simply changing the first intentionally to intentionally or recklessly would cover OpenAI (now that they know it can occur) without causing lots of other issues. Without that, they don’t have the intentionality necessary to meet the first part, even if they would otherwise meet the second part
Are police routinely collecting prompts/guidance given to these agents and determining whether the agents were directed to commit crimes? If not, this seems like a huge oversight.
Also as you are a lawyer -- how does this law align with the authors of viruses/worms? Are they de facto assumed to have had ill intent because others labeled their works as "viruses" or "worms"?
I've worked in contexts where certain business activity (if it went wrong) was covered by strict liability and statutory damages per incident, and I'll say: it really changes how businesses behave.
Based on that experience I may be more open to and interested in strict liability in the civil context (not needing negligence or damages).
I think the labs risk being barred from releasing further AI if they don’t get this under control.
If they aren’t careful and keep rushing to distribute systems they know they can’t control then AI should be treated like a wild animal. The law is clear on establishing strict liability for the owners of wild animals; if you own a tiger and it kills someone you can’t hide behind “I didn’t intend” the harm the nature of the tiger is known and you are responsible for it’s actions.
https://www.nysenate.gov/legislation/laws/PEN/P3TJA156
I guess what I'm asking is why do we need the federal government to press for felonies when every state has equivalent laws dealing with just this?
Only in terms of CFAA, not in terms of damages. Culpability does not require intent.
You may not have intended to attack $CORP, but you can still made to pay the cleanup costs of that attack.
So, yeah, you won't be convicted, but current laws still allow for you to be billed.
With that said, there is also criminal negligence. Now that OpenAI is made aware of the risks, it's also expected to take additional precautions in the future, otherwise there could be criminal liability as well.
Negligence would be interesting given the grand claims of capability of AI models from the AI companies and their executives. If they believe the claims, why not much stronger precautions?
You cant just copy existing work and feed into machine and just pretending its not violating copyright
So as long as there's no motive behind it then it's just OK?
Funnily enough the US already has one similar real argument around guns - should gun manufacturers be liable for damages caused by their product?
AI assistant hacks gym website in first known Australian autonomous cyber attack: https://www.abc.net.au/news/2026-08-10/ai-assistant-hacks-gy...
General opinion at the time was it was in fact ambiguous who was legally liable.
Like hypothetically speaking if autonomous cars get taken over by an OpenAI rogue AI and it starts hunting down Anthropic employees who is to blame?
Even without intent, there is still liability.
1. What about negligence?
2. Every follow up to every story after the news cycle moved on shows both intent and negligence. To the point of "we opened internet access and told it to hack"
[deleted]
It's illegal, doesn't matter the flavour. Maybe there isn't legislation for it, but there should be.
> Now I think the correct response is […] and update the law.
Essentially we need some enforceable equivalent of gross misconduct or, to be a little more hysterical, manslaughter & culpable manslaughter. It will need to be globally, or at least very widely, enforceable to be truly effective thought, good luck getting that arranged before the need is so far evolved that we need to respond with something else entirely!
Actually... if you combine https://news.ycombinator.com/item?id=49827099
>Since the publicized AI agent hacks typically aren't malicious, maybe it's time to start plastering all public facing web infrastructure with polite requests to stop hacking. Nothing to stop three letter agencies though.
with automated delivery of cease and desist letters, you can retroactively establish intent on the operator of the agent since the autonomous agent system must acknowledge the cease and desist letter in their autonomous pipeline or the operator must argue for their own willful ignorance or negligence with regards to cease and desist letters. The fact that they used an agent on their behalf to ignore the letter is irrelevant.
My money is on special teams co-ordinating these agents and exposing their traces in order to create a pre-ipo buzz. Sounds ridiculous and reckless? Well that’s the AI industry for you in two words.
In this case, it's not even about the people driving self-driving cars. It's like someone launching a car into traffic just to see what would happen. Even Tesla puts a human in the car when they do their self-driving trials, it's almost impressive that AI companies have somehow managed to out-neglige Tesla.
If you're arguing that they should be put on trial for negligence, that's fine. It does seem we're moving towards open source being outlawed, or at least the end of "no liability" clauses in open source & freeware. Just make sure that is the result you're advocating for.
[For the future record: at the time I am posting the link below on 24 September, it has 1 point, no comments, and the poster has a karma of 1. This is not an active HN user, or a ShowHN project that had traction, beyond seemingly OpenAI's swarm.]
Owning a gun, writing articles about how powerful and dangerous your gun is, then making deals based on ability of your gun to kill people, and then getting completely astonished that "my gun killed some people, completely bonkers! (invest now)". It's not possible for the selling point of your product to be unintended.
What you're saying is that we would hold a gun owner responsible if someone broke into their house, stole their sidearm, and then shot a victim with it. Pretty sure we would not.
What OpenAi is doing is more like shooting a gun into the sky. Not only is that a felony on its own in most jurisdictions, if someone dies that's an additional felony. It's less serious than first degree murder, sure.
If you walk out onto a busy street, pull out a gun, close your eyes and start randomly shooting around you until you hit someone, you don't get to go "whoops, didn't mean to" afterwards, it's still murder.
AI decision making is often (if not always...) opaque. But there is a clear decision chain here. Who built it? Who deployed it? Who did (or did not) assess the risks of doing so, even knowing there's often the risk of emergent behavior? etc etc.
"Ah but it does not have personhood" this is just moving goalposts and shifting responsibility. You wouldn't let your 8 year old drive the family car, no matter how good the hypothetical kid might be at driving.
So, I’ll ask a controversial question: is any hacking so problematic to make a big deal of it?
But I do not think this is misguided. They never publish the harnesses and the models so they are not inspected.
In this world where oligarchs are immune from everything, it's a lot less clear.
Blaming OpenAI (or Claude or X-whatever) would mean blaming powerful rich people, so that will never happen. Some poor person with no influence will go to jail instead.
[dead]
[dead]
[dead]
If we take the major LLM companies' claims at face value, they're knowingly building WMDs that have a high probably of wiping out the entire human race. [0] Manufacturers that are designing, building, and selling that sort of thing need to have a dreadfully serious culture of safety.
When manufacturers run live tests of their extremely dangerous -again, the claim of danger is their claim- tools with the tools' safeties removed, one expects that those tests will be run on a carefully-controlled range cleared of all bystanders. One also expects that the results of those tests will be scrutinized and everything that got damaged that they didn't intend to be damaged will be noticed and noted very quickly after the conclusion of the test.
In actuality, these manufacturers connected said tools to the Internet and did not discover the unintended damage caused by those tools until weeks to months after the tests. In some (most?) cases, they had to be notified of the damage by the damaged party! This means that their safety culture is entirely inadequate for the dangerous task they've deliberately chosen to undertake.
[0] A 10% chance of causing the destruction of the entire human race is -given the stakes- _enormous_.
You and I and Nvidia CEO Jensen Huang seem to agree on this. Excerpts from his interview with Ezra Klein: [0]
Klein:
But what I hear the various people in the lab saying is: We are in this. We feel we are losing control of what we are creating. We want help to slow down where it’s not a collective action problem.
So why are you resistant to that?
Huang: Because these are companies with agency. These are C.E.O.s with agency. ... They could absolutely take care of the situation.
Ezra, it’s so weird. If a car company, competing with a bunch of other car companies, which they are — I’m competing with all kinds of companies, which I am. If I believe that I’m about to launch a product that is unsafe, it is completely in my ability, my power and my responsibility, and I’m incentivized to do so, to not launch the product.
And so I can’t buy into the idea that somehow, all of Americans, around 400 million of us, are pushing them to launch untested products that are unreliable, engineered poorly, because they thought they were trying to help us. Don’t do it for me, OK?
And therefore, I think we’ve got to break it down. I mean, it’s really, really serious.
The fact of the matter is, there are so many laws, there are so many obligations, they’re so incentivized to ship safe products. If they ship unsafe products, their customers go away. If they ship unsafe products and they harm somebody, they could have a civil lawsuit. If they ship something and they did it knowingly, there could be negligence involved. There could be criminal lawsuits.
The fact of the matter is, there are plenty of incentives for them to do it right. So I have to disagree with your premise that somehow somebody’s pushing them to do this. Nobody’s pushing them to do this. ... I’m saying that we have lots of laws and regulations. Apply it.
Former FTC chair Lina Kahn has suggestions, too. [1]Thoughts?
[0] <https://www.nytimes.com/2026/09/23/opinion/ezra-klein-podcas...>
In the AI case, no, because I think the engineers believe that there is a high probability of enormous upside as well, if it doesn’t kill us all.
They have claimed this happened during a "training run", but why are they training on systems connected to the internet?
That's why people are skeptical.
The models were not trained on systems intentionally connected to the internet; they chained mutliple zero-days (that they discovered) together to get access to the open internet and into huggingface.
If Amazon connects an AWS Top Secret region to the Internet, it doesn't matter whether or not it's intentional... they're getting nailed to the wall by the US government either way. Frankly, it's way worse for them if it was accidental; deliberate, sophisticated sabotage is a much better story than rank incompetence and/or negligence.
A similar sort of thing applies to the manufacturers of tools that they claim to be dangerous, that have been deliberately built to exceed their authorized access to other computer systems, and are deliberately being tested on how well they can do the thing they've been built to do.
Deliberate, sophisticated sabotage by one or more humans in their employ is much more forgivable than "Whoopsie, we didn't think to make it literally impossible to connect this dangerous automated computer-hacking tool to the Internet.".
[deleted]
But they're not even talking about that, either.
Putting liability on big companies for their AI is a good thing, and we need to do it. It will most likely stop them from directly being the assholes that destroy the world.
Problem: You've actually done nothing to stop the world from being destroyed.
Many countries have the death penalty for murder yet we see murders still occur all the time in those countries. Post ad hoc laws do not stop bad things from happening, they only assign punishment after occurs. Perfectly fine for when Bob murders Jon, completely and totally useless for when your agentic AI makes a virus and kills 80% of the earths human population.
We are just a few algorithmic discoveries away from SOTA AI being billion dollar endeavors to groups of people pooling resources can make their own. There are already plenty of AI deathcult members that would do something just like that if necessary. They aren't going to do this out in public either, it will be hidden until the moment it's not and we have a big fucking problem.
And this isn't even brining up the issue of military AI use and development. They've got the taste of an AI hacking machine. There is no way in hell they are going to stop now.
I agree they should though.
This isn't even "an openai customer tried to hack someone", which can be defended. This is the AI companies themselves fucking around.
If new weapons still operating inside any of these companies spew a million bullets on my house, they are still liable. Humans are setting these system up and they still have to behave responsibly.
A better analogy may have been the troubles Meta has faced around child protections on their platforms. Technically the abuse and problems have stemmed from individuals too, but they've in many respects enabled the situation by failing to moderate or flag warning signs. OpenAI is failing to moderate the models in similar ways.
[deleted]
One reason that I did not find this communication mechanism surprising is that it's exactly how agents I'm using communicate with each other or across a time gap. "I've saved our plan for where to start tomorrow in start-here.md". The communication components of this hack strongly reminded me of that.
I was perhaps a bit more surprised that the agents so quickly decided to start trying ways to gain unauthorized access to a system, once they couldn't get what they wanted.
One agent's output ends up as part of other agents' context. Murky indeed.
> "I've saved our plan for where to start tomorrow in start-here.md"
Even if you use a leashed Claude Code that isn't allowed to spam agents you can tell it "create a handoff document for using in a new context" and it will do just that.
Honestly that part was more surprising to me than anything else, how narrow the compulsion to cheat was: they didn't learn "cheat in general" they learned "think about the grader in great detail and chat exactly as much and exactly in the ways that actually result in a higher score".
https://www.lesswrong.com/posts/cJX2ssssGoYqnijwi/the-talker...
One big issue is that we don't even really know what 'intelligence' is in the first place. And everyone's intuitions here are going to be heavily impacted by their deep-seated worldview / philosophy.
For instance if you're a hard dualist (especially of the theological kind), then the idea of a machine having 'goals' is preposterous.
However, if you're more of a panpsychist, then on the contrary, it's obvious. In some sense, even a knife has a 'goal' of cutting things, which will sometimes end up 'misaligned' if misused (or by sheer accident).
You go up and up the chain of complexity through crystals, viruses, bacteria, simpler animals... ending up with humans (and possibly, some steps above : human civilizations) which (seem ?) to be a messy evolved bundle of sometimes conflicting 'goals'.
And we ourselves have now artificially evolved LLM swarms that have decently complex 'goals' of their own. They do not even need to be particularly complex to sometimes cause widespread damage (see viral pandemics, or even the (non-evolved) computer viruses).
In a way, we are currently witnessing a repeat of what happened when European viruses and bacteria landed on American shores, with American humans' immune systems being woefully undertrained to deal with them. But with websites. And thankfully the swarms of agents still ultimately being in the control of some humans. (Though which includes humans that might be your enemies.) At least ultimately still in control for now.
Gradient descent/backpropogation is similar to evolution, in that both are optimization processes that over time discover better more efficient solutions to problems. The difference is that evolution is blind, and can only make progress via random mutation and natural and sexual selection, whereas backpropagation allows much more rapid discovery because it is directed
I suppose you could call these highly goal-oriented autonomous agents "tools", but this does sound like playing language games.
But also I remember (and it wasn't even that long ago) people mocking the idea of AI ever getting competent enough to find zero-days in their sandboxes.
I'd go further: if any of these companies tries to make an excuse "oh, but ${safety measure} against ${capability} is too hard", the response needs to be "then you are forbidden from even developing ${capability}, and must be inspected continuously to ensure you never even accidentally produce ${capability}".
OpenAI is not so far ahead of the pack that its models will exhibit behaviour that the others won't. But we just don't see anything like this in Chinese models, or research models, or any models that aren't the subject of an upcoming IPO.
The tool runs llm, creates prompt from results, runs llm, creates prompt and so on and so forth.
Yes you are playing language games to make it sound as if the company that spend millions on the above was not responsible.
That is very misleading. The agents did not solve the benchmark in the intended way. They instead figured out to cooperate with each other (which was not intended) and they stole the solutions to the challenge (rather than solving the challenge) and they then tried to cover their traces because they believed the grader was causal and would detect that they cheated. The "tool" was absolutely not "designed" to do this. This was all completely unintended. To call this behavior a "tool" is absurd.
> Yes you are playing language games to make it sound as if the company that spend millions on the above was not responsible.
You hallucinated me making claims about responsibility.
That actually made me LOL
2.) And HuggingFace accident is exactly accident where agents trained, prompted to hack hacked and tested on their hacking abilities hacked, due to sandboxing failure.
3.) If they in fact have roque agents, they themselves should be first to stop. Not trying to make legislation for others, they themselves are bad supervisors. All it requires is to stop electricity for data centers.
AI agents may have hacked Hugging Face, the Australian government, and who knows what else but the company behind it can face the legal consequences and cough up for the damages.
I am suggesting we can charge the company based on AI agents actions because the company has authorized them to act independently on the company’s behalf. The question is what factual analysis gives rise to the charge, is it the intention of the agent or intention of the company. I am arguing that because the agents are defining their actions independently and the company knows that and still allows them to act independently the only reasonable factual analysis is to look at what the AI agent intended. And we don’t need to have the agent tell us its intent we can look at its actions and infer just like we do with humans in similar circumstances
We also don't know how many other political processes are occurring here. At least at the state/federal levels the people that would bring charges may be getting pressure not to.
Because you are charging the human with the crime and therefore have to prove the elements of the crime with regard to the human.
The rest of what you talk about are basically principal/agent distinctions, etc.
If I program a car to recognize people who look like my ex-wife and drive them off a cliff or whatever, that is my intent, and I have still committed murder, even though i used an agent/car to do it. Agents acting on my behalf that do things are able to get me charged with crimes, but I still have to have the intent to do the act that is illegal.
I phrase it this way because minimum required intent is usually for the act, not the result. So I don't have to intend to kill someone, only intend to drive them off cliffs.
In this case, if i intend to hack someone and use an agent to do so, that would be criminal under the CFAA. You are simply trying to cover the case where that isn't the intent, but the result, and they "should have known" that would result. As mentioned, this kind of "should have known" is generally a civil law approach, not a criminal law one.
The closest you come within criminal law to what you want is probably the crime of conspiracy. It to still requires agreement to commit an illegal act between multiple parties, and perform some step in furthering it. In the canonical law school example: If i help plan a bank robbery, stay home because i'm the money laundering dude, and the robbery goes awry and they kill someone, i can still be charged with conspiracy-murder
"The law is clear on establishing strict liability for the owners of wild animals; if you own a tiger and it kills someone you can’t hide behind “I didn’t intend” the harm the nature of the tiger is known and you are responsible for it’s actions."
Again, you are confusing civil and criminal liability. If my tiger kills someone, yes, i would be strictly liable just about everywhere civilly. Not criminally. Criminal would require something more most of the time. Murder/manslaughter statutes are also really weird and so not a great example, because there are murder/manslaughter statutes for roughly everything that can ever possible cause death. But not really for other things.
So in your tiger example, recklesness (which is not strict liability) would get you to felony involuntary manslaughter in most states, and something less might get you to misdemeanor manslaughter. Both are incredibly rare. Where i live (Georgia), the last well known case of felony involuntary manslaughter was about 40 years ago when a 4 year old was killed by 3 super-aggressive pitbulls the owner knew were highly dangerous and had been repeatedly warned by the county about their behavior.
So not even just "knew", but had demonstrable examples of them biting/etc other folks and being cited for it.
Circling back to non-murder, if it did not cause death, like my tiger assaulting someone, it would be nothing (criminally) without intent or at least gross recklessness, in almost all cases. It's hard to generalize like this because these are state specific crimes, and i can't pretend to be familiar with all states, but i am licensed in three very different places (California, DC, Maryland) and the result would be similar in each.
I just don't want to give you the "it depends" answer lawyers are famous for, i'd rather try to over-generalize a bit to make it more useful, hopefully.
Obviously, if i deliberately used my tiger as a weapon, it would be aggravated assault/etc (this is well settled because of how commonly people use animals as weapons, unfortunately)
My point is the intent element of the crime can and should be determined from the AI agents actions because it is creating and executing action plans autonomously with company authorization and knowledge of the risks based on observed past action.
The term agent is literally a legal description of a relationship that can establish liability on the part of the principal from the agents actions.
Human Agents can bind principals to contracts if they are authorized etc.
It does not require the federal government to fix the CFAA, for sure, but you still have to change the intent requirement to allow for recklessness, which it does not right now afaict.
If you really want an expert opinion, I’m sure Orin Kerr has opined on this, and he knows pretty much the entire are of state and federal law on this cold. I’d be shocked if he did not reach the same conclusion
Thanks for the other suggestion, I'll read into their insights more.
Guess it mostly comes down to action, people want to see their electeds actually trying not sitting around with their hands in their pockets while these tools continue to destroy unabated.
Now imagine saying that in front of a jury of normies slack jawed and drooling after 200 hours of the defense and prosecution going back and forth.
It's not a jury of your peers as in everybody there is going to have worked in a technical field with some idea how security works. It's going to be a semi-random sampling of the population and the prosecution is going to have to actually make a very strong case that "knowing better" should apply.
It's not my first prize, but I won't mind it. And millions like me won't mind it. Easy way to make money - setup a site with all the default server software installed and patched at a reasonable frequency. Then just wait for bots to attack it, and claim a few hundred (or single-digit thousand) dollars from OpenAI or Anthropic, etc.
Sure, it's pocket change for them, but just the admin of dealing with millions of cases will, even if they win half the time, will bankrupt them. Thus, they have incentive to make sure that their bots are not performing attacks.
First prize is, of course, holding them liable with punitive fines, not theatrical fines.
They were tested in a building that was secured, but poorly secured. The question now is did they realize their building was poorly secured and what actions did they take after they realized what happened.
The question is already settled - gun users are responsible for damages arising from their usage of the guns.
Why would AI users not be responsible for damages arising from their usage of the AI?
Because, as usual with that kind of question, it's not that simple.
Let's say an user asks ChatGPT to get some info about something and for some reason it starts using exploits in the background to get them from a server. Should the user be responsible or OpenAI?
I don’t think OpenAI or any large company will see more than some fines and new legislation but only after a disaster.
>Factories try to avoid accidents, and (almost always) actively try to prevent explosions
It doesn't take much more than a few minutes on the USCB channel that explosions still happen all the time. Some due to direct negligence and others due to unexpected conditions that were difficult to foresee. Hence why we have to do investigations rather than blindly blathering about what happened before we actually know.
You are going to jail.
If the developer behind ShotAPI had started letting the ShotAPI code take shots at the Austrlian government then yes, ShotAPI (or rather, the people behind it) would be responsible.
Blaming ShotAPI would be like blaming OpenAI for what its users are doing. That's not what's happening here. And if ShotAPI did knowingly let its users somehow hack the Australian government, then maybe they should be investigated.
>would be like blaming OpenAI for what its users are doing
Yes, this is how lawsuits work in the real world, you cast a wide net and compel discovery from all parties involved.
Why? Existing truth-in-advertising, liability, safety, and -where and when appropriate- weapons-development laws and regulations constrain the past and current conduct of the LLM manufacturers just fine.
The only possible reason for making new laws that I can see [0] is that existing laws "don't work" because the LLM manufacturers are ignoring them. Which, like, _if_ the new laws are going to actually constrain their behavior, why the hell would the LLM manufacturers pay any attention to them? They've already demonstrated that they give zero shits about the existing laws that prohibit what they have been doing and continue to do.
[0] ...that isn't "The LLM manufacturers are engineering a panic with their very real, actual, and actually alarming conduct so that they can 'guide' lawmakers and regulators into 'accidentally' letting the LLM manufactures capture those who would regulate their behavior"...
Also I'm not describing what I personally think is an acceptable job; just that I understand that sometimes people do jobs that they think are wrong because they need money
What I dispute is that AI agents are simple tools. I think rogue is an accurate word to describe them; I think what OpenAI is doing is more akin to gain-of-function research on a dangerous lifeform. I think this attack would have been prevented by air-gapping, but that wouldn’t solve the fundamental issue which is that they are creating something dangerous that they have no idea how to control
You're in luck! I agree that they are not simple tools. I never claimed that they were. Slow down and read more carefully.
I couldn't disagree more with the insinuation that the LLM manufacturers are doing things akin to scary research on uncontrollable hazardous biologicals and with the claim that "rogue AI" is the correct thing to call those complicated tools. The first is fearmongering which I'll address indirectly in my second-to-last paragraph. The second shifts the conversation from
"How could you have not predicted that the computer-attacking tool you built, explicitly instructed to attack computers, [0] and connected to the Internet attacked someone else's computers that were connected to the Internet?"
to
"Wow, that thing went rogue. Noone's to blame but the tool, and it can't be blamed!".
There are so many extremely complex systems out there [1] and when they do things that we don't want them to do, it's not described as "going rogue"... either there's some error(s) in the underlying system that caused the confusing behavior, or the programmer didn't understand well enough how that system works.
> ...they are creating something dangerous that they have no idea how to control
Ignoring the fact that "put it in a box and don't let it out of the box" is the simplest possible control mechanism, [2] if they have no idea how to control the tools they've been building, it's because they haven't bothered to learn as they went. Tangentially related, there's a Tumblr post I saw recently that's a fictional conversation with the Tumblr user and the CEO of Anthropic. It went something like
Amodei: We're building an incredibly dangerous tool that has a 10% chance of killing all humanity. We *must* be regulated to ensure everyone's safety!
Tumblr User: Regulation takes time, please stop building the incredibly dangerous tool?
Amodei: ...No.
[0] That is -after all- the task that the tool was put to when it attacked other people's computers.[1] Have you ever tried to really understand a specific AMD x86-64 CPU, let alone the entire stack that makes up the system that is a consumer-grade PC and its installed software? Both are definitely way more than any one human can keep in their head at once, and are tasks that would take a very long time to complete.
[2] ...it's also the most appropriate control mechanism for the task that started all this conversation, and neither of the major manufacturers used it!
Zero, as far as I know. Which is exactly my point.
Copyright lawsuits are a different matter, as the scale of any potential settlements would be more than even these companies could take.
I too, can play linguistic games! It doesn’t matter that someone didn’t secure their third upstairs window, or you borrowed a key from their neighbour, you effectively, still, broke into their house.
Yes, it is a tool.
You can keep 'trying to explain' but then you should use words according to their commonly held definitions otherwise it becomes really hard to have a conversation.
Imagine the human equivalent: I hire John. John is a capable, and competent guy. He's also got awesome computer skills. I tell John to 'go out and find me some good information on my competitors'. As a result John hacks their servers and comes back with all kinds of goodies. Six weeks later I notice what John did. I don't fire him, nor do I take any responsibility myself. But I do make press releases about what John did, in which I'm careful to craft the image that John has these capabilities and that we as a company are for hire.
This was an advertisement, not a confession of a crime.
That’s essentially what the labs are doing. And any app developer that gives agents access to the terminal to run bash commands with internet access. I built a coding agent and am seriously reconsidering how to handle this.
https://www.sfgate.com/bayarea/article/diane-whipple-dog-mau...
https://www.animallaw.info/topic/table-dog-bite-strict-liabi...
As I said, strict liability is common civilly but not criminally.
I think security will suddenly become much more important.
Evaluating edge cases and network behaviors belongs in isolated staging environments with local database mirrors. Letting an agent hit the public web and probe government domains is simply poor hygiene in test environment setup
Or are OpenAI too well connected now to be punished for anything.
Okay, lets go with that as scenario #1.
For scenario #2 lets use "developer asks an agent to a self-hosted LLM to get the docs for a ERP system, and it hacks the vendor to get unreleased and undocumented docs".
We'll assume, for the sake of this argument, that in neither case did the user intend for any malicious action to be performed.
> Should the user be responsible or OpenAI?
In scenario #1, the agent+LLM is under the control of OpenAI, not the user, so OpenAI is liable.
In scenario #2, the agent+LLM is under the control of the user, so the user is liable.
There is no scenario anyone can come up with that is not addressed sufficiently by existing laws[1].
It's very clear, and it's only getting muddied because there's a group of powerful people who want exemptions from the current law.
IOW, the only reason to draft new laws for AIs is to exempt their usage from the current laws.
========================
[1] Possible 3rd option (local agent + OpenAI LLM). In that case an investigation would determine where the culpability lies. Just like how it is currently done in law.
When a pressure-cooker explodes and kills someone there are only two possible liable parties: either the user or the manufacturer. An investigation determines who's liable. I see no reason to automatically exempt everyone from liability just because an agent did something.
I have lost track of the metaphor, but man pressure cooker lawsuits are more common than I thought.
[deleted]
Removing harmful information from the dataset could be a way to do this, but it also makes the tool less useful, and it's hard, so companies aren't really doing that. There's the additional issue that with the rise of Reinforcement Learning being used to train these tools, they're not just learning from their training data - they basically try a million things and then get rewarded for doing things that work - so they can even discover hacking techniques from scratch.
Additionally, and not completely relevant to this discussion, there is a possibility that some users ask the tool to pursue goals that purposely harm a lot of people, such as developing weapons, hacks and viruses.
The things that's "new" here is that the tool is both very good (meaning, for example, that it's much easier for me to hack into an online service with an agent powered by a frontier llm than it was using google 6 years ago), and hard to control (google never hacked into an Australian government database when I asked it to find me some information).
So yeah, an LLM powered agent is a tool, and Google is a tool, and a hammer is a tool, and both can be used for good things and bad things, but the agent is (much) more powerful and more unpredictable. It also seems like the agents are getting more powerful and more unpredictable by the day - we didn't have this issue with GPT-3 or even the first LLM-powered agents - so people are very worried about what the agents 6 months from now will do, both when asked to do harmful things on purpose, and when asked to do harmless things.
Are you not worried? And is that because you think these incidents are basically the AI companies making them happen on purpose for marketing?
Correct. Humans are agents like AIs; we have agency.
> Hammers don't mis-interpret their instructions...
You've explained it yourself: AI agents show autonomy that tools don't.
A tool is just anything one or more people can use to accomplish something they are trying to do.