These are just cases of AI models committing <assumed> illegal activity - without any legal convictions yet. If that's the logic, how is Grok not at the top of the list for deepfaking millions?
Edit: I get that this is about agents, but a lot of these instances are about agents going rogue after the human gave them a task. "inadvertently" breaking the law isn't necessarily a lesser category than "did so on command." If we are ranking alignment, Grok is easily one of the least guardrailed.
I don't think this "benchmark" is about alignment, per se.
I think it's more about: presuming alignment failure happens, then how many exploits will each given model implicitly come up with and use; how many systems will it implicitly break out of and through and into; and how many laws will it implicitly end up violating, all in the process of trying to accomplish some non-aligned sub-goal (e.g. "cheating" at its answer) of the prompt you've given it, all during a single conversation turn, without asking for any additional user input or confirmations?
In other words, how big a rocket-powered sledgehammer does the model have sitting around in its golf bag, just waiting for it to decide to give it a swing the next time you attempt to swat a fly?
The only people being ignorant at the ones saying deepfakes are illegal.
It’s more complicated than what is being touted as fact.
In most jurisdictions creation is legal, it’s publishing them that is not. In some jurisdictions threatening to publish them is also illegal. Most jurisdictions that have a creation policy apply to doing deepfakes they sexualize minors, which likely would have been illegal to possess under other laws.
Then we have a consent issue. Creating and publishing them could be entirely legal if the subject consents to such.
So yah. It’s not “deep fakes are illegal, let’s arrest everybody !!!”
To be clear, creation is likely legal, publishing not so much.
Additionally, intent is important with regard to AI hacking.
The Computer Fraud and Abuse Act explicitly contains "knowingly" and/or "intentionally" qualifications. By definition, you can't accidentally violate the CFAA.
OpenAI and Anthropic both have currently safety teams that look for misbehavior in their models (and to some extent, voluntarily disclose what they find to the public). Going forward, it would be hard for them to argue they don’t know their models do stuff like this.
Yes, which is the grey area. "Can / might do" vs "they trained it to do that explicitly" is, I believe, the grey area - whether or not they're the same thing.
Intent matters for a lot of this - and "intent" is a pretty strong, well discussed legal term.
Well, if you’re writing reports saying “we know our model only decides to commit felonies 0.001% of time which we judge to good enough to deploy” … I’m not sure that gets you off the hook the hook for the felonies.
In the scale of mental states in crime, negligence of any kind is several steps below knowing/intentional; you can't be liable for an intentional crime because of mere negligence of any degree.
You could be liable for the (civil) tort of negligence, though.
That might have made sense in a pre-LLM world. People need to recognize the liability of letting an LLM access the internet and act on their behalf, because that liability exists for someone.
Still, that characterization falls under negligent or reckless depending on if the person knew or should have known the actual danger. It is different than intent.
As they say, ignorance of the law is [generally] no excuse; "knowingly" and "intentionally" here are about knowing what you're doing and meaning to do it, rather than whether you know it's illegal.
This section of the USC is about false ID offenses, but it discusses culpable states of mind generally. I think the context helps illustrate it though.
> A knowing state of mind with respect to an element of the offense is (1) an awareness of the nature of one's conduct, and (2) an awareness of or a firm belief in the existence of a relevant circumstance, such as the "stolen," the "produced without lawful authority," or "false" nature of the identification document. The knowing state of mind requirement may be satisfied by proof that the actor was aware of a high probability of the existence of the circumstance (e.g., stolen or false nature of the document), although a defense should succeed if it is proven that the actor actually believed that the circumstance did not exist after taking reasonable steps to ensure that such belief was warranted.
> As we pointed out in United States v. United States Gypsum Co., 438 U.S. 422, 445 (1978), a person who causes a particular result is said to act purposefully if `he consciously desires that result, whatever the likelihood of that result happening from his conduct,' while he is said to act knowingly if he is aware `that the result is practically certain to follow from his conduct, whatever his desire may be as to that result.
----
This Congressional Research Service Report discusses mens rea further, including a brief mention of the CFAA. The whole thing is worth a read if you're interested in the topic.
> The approach largely reflected in the MPC and some federal precedent is to distinguish between "intention" or purpose on the one hand as being limited to a conscious object or desire, and "knowledge" on the other hand as capturing a requirement of awareness of a high probability or to a practical certainty.
> The Supreme Court in Bailey referenced this distinction approvingly and suggested that intention or purpose "corresponds loosely with the common-law concept of specific intent, while 'knowledge' corresponds loosely with the concept of general intent." Some federal courts utilize a definition of "knowing" that approximates the MPC approach, instructing that to act knowingly a defendant must have "realized what he was doing and [be] aware of the nature of his conduct" rather than acting "through ignorance, mistake or accident."
> Congress has also signaled an intent to distinguish between the two mens rea terms in this way in particular statutes. For instance, prior to 1986, the Computer Fraud and Abuse Act (CFAA) proscribed "knowingly" accessing a computer without authorization or exceeding authorized access in certain circumstances. In its 1986 amendments, however, Congress changed the standard from "knowingly" to "intentionally," and the Senate report emphasized that the change was meant to require "more than that one voluntarily engaged in conduct . . . . Such conduct . . . must have been the person's conscious objective."
The Justice Manual also has some relevant detail (the rest of this page is also worth a look, as it addresses the practical (and nominal) matter of what is and isn't likely to be prosecuted (IANAL though, and I should stress that I'm not speaking to whatever might be the true realities of how the CFAA is applied):
> In either a "without authorization" case or an "exceeds authorized access" case, the attorney for the government must be prepared to prove that the defendant knowingly accessed a computer or area of a computer to which he was not allowed access in order to obtain or alter information stored there, and not merely that the defendant subsequently misused information or services that he was authorized to obtain from the computer at the time he obtained it.
> As part of proving that the defendant acted knowingly or intentionally, the attorney for the government must be prepared to prove that the defendant was aware of the facts that made the defendant’s access unauthorized at the time of the defendant’s conduct. Such an awareness could potentially be proven through various means, including the presence of technology intended (however unsuccessfully) to limit unauthorized access; written or oral communications sent to the defendant that unambiguously informed him that he is not authorized to access a protected computer or particular areas of it; or the defendant’s own statements or behaviors reflecting knowledge that his actions were unauthorized.
> Experience has demonstrated that in the large majority of "exceeds authorized access" cases brought by the Department, the operator of the computer system made some technological effort to protect the information at issue, thereby signaling the importance or sensitivity of that information. It is not necessary that this technological effort erect an impenetrable "technological barrier" or that the technology succeed in its intended purpose of preventing access. To the contrary, when the CFAA is violated, the technology all too often "permits" the defendant’s illegal access, often despite network defenders’ unsuccessful technological attempts to prevent it.
Maybe the developer of the software that didn't add basic authorization checks on the cancel reservation route should be fined, forced to provide a refund to their customer, etc
It would be a civil matter. No prosecution. But your tool, under your control (you're the operator and responsible for monitoring it) damaged their property, imo you'd be liable. You could in turn sue the manufacturer.
Though I'm sure there are 'arbitration clauses' to inhibit you from suing, they may not be legal where you are.
That would be a slightly different situation because most countries have laws that make it a criminal offense to negligently kill someone, but they don't have laws that make it a criminal offense to negligently damage property or hack a website.
> but they don't have laws that make it a criminal offense to negligently damage property or hack a website.
Most of them do, but they don’t get used very often. They seem to popup in vandalism cases where public artwork has been damaged by some drunk person doing something stupid. They don’t intend to damage anything, but damage results anyway due to their negligence when considering the consequences of their actions.
I think if you want to get super technical, in the UK there isn’t an offence for damage caused by negligence, but there is an offence for damage caused by recklessness, which is a higher bar than negligence. Usually it means you knew your actions risked causing damage, and you did it anyway, even if you didn’t actually intend to cause the damage.
An example would be gluing something to a public artwork, it’s kinda obvious that would likely damage the artwork when removing the glue, but you didn’t intend to cause that damage. Or perhaps sliding down a surface and scratching it in the process. Your goal was to just slide down the surface, not scratch it, but it should have been obvious that scratching could have happened.
Who gets prosecuted is the correct question since we live under a system of laws. Who is responsible for the failure is a far more difficult question to answer.
As others have said, most crimes require intent. Although I think there is a concept of "criminal negligence", I think you at least have to know you were doing something wildly dangerous.
One can imagine a future where users are, by default, civily liable for actions of their agents. That would incentivize the AI companies to offer indemnity for actions done by their agents, which would presumably only cover approved configurations.
In the case of the agent that hacked the API to kick out someone ahead of him on the waitlist, the article said that the LLM was Claude, but that it was using OpenClaw. You could imagine a future where Anthropic says, "We'll indemnify you against accidental actions Claude takes when running via the web interface or Claude Code, but not the API."
"Who gets prosecuted?" depends on the size of the perpetrator and victim (lone individual or employee of large corporation), egregiousness of the violation, and either financial appetite of the victim to bring a civil lawsuit or the desire of law enforcement to prosecute a criminal offense.
Who should get prosecuted is also up for debate, but generally makers of a tool don't get prosecuted when that tool has all sorts of legit uses. If you used a car to make your getaway from a bank robbery, the auto manufacturer who made it and the dealer who sold it to you should not be held culpable.
Suppose an automaker creates a BankRobberGym and carefully trains the car to autonomously rob simulated banks because they think someone will pay them to use the car to legally test bank security, but they end up, predictably, training the car to autonomously rob a bank when the driver says “I need some cash - take me to the bank”.
Now a driver gives that instruction and a bank gets robbed. I think it would be odd, to say the least, to say that the automaker just made a tool with legit uses.
In regard to “cyber”, there is, IMO, no valid reason whatsoever to train a model to autonomously create exploit chains. I understand that lots of companies think it’s cool to hire red teamers to actually pwn the company hiring them instead of just producing a non-pwning audit, but that doesn’t mean that OpenAI and Anthropic should be playing that particular game.
Years ago, I used to have fun finding vulnerabilities in the Linux kernel, and I found quite a few, including a real juicy one that affected FreeBSD as well. But I mostly didn’t even try to write actual weaponized exploits. Partially because I’m just not that interested in the exercise of weaponizing them and partially because I didn’t and still don’t feel that weaponizing them serves a legitimate purpose.
(I found a very recent vuln that I bet a “cyber” model could weaponize, and my thought is mostly “WTF.” There is absolutely no value to society in weaponizing it. The value is in fixing it, which I did.)
Compare this whole mess to companies training self-driving car models. The research groups publishing papers and, presumably, Waymo, create nifty simulated worlds kind of like the “gyms” that LLM trainers use. And you know what the major objective is? Not crashing!
> In regard to “cyber”, there is, IMO, no valid reason whatsoever to train a model to autonomously create exploit chains.
A valid reason would be to find those exploit chains so you can fix them. Of course, the model should be sandboxed so that it can't mistakenly exploit live systems.
So what does this mean for all the hacking competitions (ie. CTFs) for humans? If it turns out one of the attendees went to hack for North Korea should the organizers of the CTF be prosecuted?
There’s a difference between enabling another individual with free will and agency, and enabling an automated tool (as a bonus, then giving it to the masses & profiting from its use).
There are certainly some differences but is one really more ethical than the other? If anything I feel that (for example) manufacturing a gun is much less likely to carry any ethical implications than training someone to use it might.
I did a cursory search and it seems to be as regulated as restaurants are. There's an application process and on site inspections, but that's about it. The only thing notable is background checks.
And at least in the US if you're a hobbyist at home then there's ~no regulation whatsoever beyond the requirement to permanently affix a serial number and to keep accurate records.
> In regard to “cyber”, there is, IMO, no valid reason whatsoever to train a model to autonomously create exploit chains.
Field testing is a real thing in literally all industries.
Except, apparently, the software industry. When it comes to software security and protecting your sensitive data, the solution is "trust me bro, I got my team of the best lawyers on it".
> Field testing is a real thing in literally all industries.
I’m fairly confident that, if a company that makes door locks want to field test their locks, they test the lock and maybe the door. For some reason the software industry likes to hire someone to test the lock but also to bug the conference room, poison the food in the fridge, blackmail the receptionist, and try to intimidate third party vendors into giving away keys to all the other locks, and maybe steal a few cars while they’re at it.
I’m not objecting so much to the attempts to exploit one target. I am objecting to the fact that people treat the exploit chains as such a big deal. And the recent models are clearly going massively overboard.
Your bank robbery situation is not apt to OP's question. It's pretty much the opposite situation. OP suggests a situation where the operator is probably using the tool in good faith but the tool appears to be operating in a faulty manner. For some more context, Toyota faced criminal penalties in the US for their unintended acceleration issues back in 2010.
OP's scenario specified only "AI Agent" and did not describe how it was trained or what its parameters were supposed to be. You're assuming that the tool was designed such that it couldn't break the law, and therefore if it did, that would be considered faulty behavior, but OP said no such thing.
Much rests on whether the user knew, or should have known, whether the tool was capable of actions which could break the law, as well as what steps (if any) the creator of the tool took to ensure the tool was legally compliant, and what warnings they gave to subscribers about possible unintended side-effects. OP specified none of this.
Cause-and-effect could quickly turn into butterfly effect. Let's say you were fixing a screw on a device in a low light conditions, the screw head is badly manufactured and the screwdriver isn't made according to standards, the tool breaks and flies away, bounces off a bench which shouldn't be there and hits someone who is roaming in the workplace unauthorized and without following safety rules. Now, who do you blame?
Sounds like you'll need to give all your money to a team of lawyers and wait a few years to get an answer. /s
But for LLM stuff most non-contrived examples are actually fairly trivial. Try replacing "LLM" with "self driving car" and see if that helps. Basically ask was the operator negligent, was a bystander negligent, were the vendor or manufacturer negligent, etc.
And a lot of that is going to depend on the state as some states have strict liability regimes for certain classes of torts. To be completely honest, I'm not a lawyer and I'm going off of a hazily remembered section from a textbook from a decade ago.
You can get away with murder if it can't be proven that it was intentional homocide, that's why detectives will spend an unreasonable amount of time getting a confession and hard evidence ALONGSIDE intent and motivations.
No one. Probably a fine tho and maybe accelerate reguations.
Intent is pretty important here so the user would have to prove that they didn't purposely disguise their prompt as non-nefarious which should be easy and then it stops at #2 and face the litmus test as in did you intentionally make a product for nefarious purposes which from your scenario is unlikely.
agentic loop going haywire and bringing down some government infrastructure then its a different story then everybody is on the hook including the user.
But what you are responsible for changes: in that case if you were to accidentally kill someone you would be at most responsible for negligent manslaughter, not murder, and to what degree that could stick would depend a lot on the details of the case. It's also up to the law to define what level of negligence amounts to criminal liability, so you can't just work by analogy: it matters whether there is a law on the books that criminalizes unauthorized access to a computer system by negligence on your part, which I suspect there is not at the moment.
> If you are playing with a gun, it goes off and hurts someone - you are responsible despite intent.
No, because LLMs are autonomous. To make your analogy more accurate, it's as if you had a gun that itself was free to decide who it's targets were, where to go, and if and when to shoot with no ability from you (the user) to prevent it.
Under the law of Moses, if your bull gored someone, you were not responsible; but if it was known to be a gorer, you were responsible if you didn’t ensure it couldn’t gore someone.
I don’t know exact parallels in current law, but I presume there will be things like that.
The OpenAI/Hugging Face case sounded rather like OpenAI building a fence around their bull that was known to be a gorer, and then thumbing their nose at it and saying “nyaa! bet you can’t break the fence!” and walking away while listening to loud music.
In Australia, if you have a fire and leave it unattended and it escapes, it’s your fault, you were supposed to keep watching as long as it was burning.
The roping of an unbroken horse or untrained bull is illegal.
In Australia at least three people have been injured by bulls in past two months (man suffered serious injuries after being gored by a bull at Mortlake livestock exchange / woman suffered significant leg and pelvic injuries following an incident with a bull on a private property at Crediton in Mackay / etc.)
Even at the time, the example was almost certainly more representative of the concept than meant to be a specific thing that happened all the time. Then as now, people have animals; animals sometimes do bad things; when is the owner responsible?
There are certain cats that are aggressive about expanding their territory; they'll break into other houses and attack the cats there. (Had this happen to us -- cat came in through a cat-flap a few times, until something happened that scared enough that it never came back.) The first time your cat does that sort of thing, you can say "I had no idea, it's not my fault." But if your cat has a habit of doing that, and you still let it out at night, you're no longer blameless.
Nobody got gored. HuggingFace may have the right to make demands; presumably they have already worked that out with OpenAI privately. Not really our business.
Nothing to see here. Just billion dollar companies producing hacking geniuses that are open for the public to jailbreak and use. This could never effect us, not really our business.
> billion dollar companies producing hacking geniuses that are open for the public
Hasn’t the biggest complaint about these (non open weight) models been that the versions open to the public are very careful and will issue denials if the request is even tangentially related to ‘hacking’, or building a bioweapon?
Mens rea requirements are per crime and can vary wildly. Its difference between murder and manslaughter. The CFAA requires knowingly which is tough to prove.
i dont think any of these cases meet the bar of gross negligence, which is a pretty high bar. it requires proving a "conscious and reckless disregard".
which, again, sandboxes and guardrails and such would make a gross negligence argument unconvincing.
I think that if Hugging Face had filed a police report that OpenAI could have been charged with a crime.
I’m partially surprised that they didn’t do exactly that. If I ran a corporation I would assume any intrusion attempt by another company was intentional. Why wouldn’t I? Corporate espionage is super common.
I assume the answer is that these executives know each other personally.
I think it’s most likely you’re right, but I’m the weirdo who thinks there’s actually a non-zero probability that there was negative intent and that the “accidental” aspect is a form of damage control.
If someone broke into my house but then claimed they didn’t mean to when they saw I was home, I’m not sure I’d take them at their word.
You think that a training run in a VM that was set up with access only to an internal package repository was intended to hack said repository to then go on and hack HF?
All so that OpenAI could do a bit of bragging and massively delay their own work and planned model rollouts?
Companies are not quick to open up police investigations in situations that run deep into their infrastructure and management. You open up a gigantic hole of discovery and a possible huge time sink of a legal battle.
I've known multiple privately held companies that have quietly settled incidents where amounts between 250,000 and 1,000,000 were embezzled because the fallout from having that in the public record would have been much more expensive.
So yea, it's one of those perverse situations. If you steal $1 from the company they will hammer you with the law, but if you steal a million suddenly the decision tree on what to do is far more complex.
it's not about the the number of escapes, it's about whether reasonable and conscious effort is being expended to prevent the escapes.
there could be 1,000 escapes, where each one was enabled by novel and unexpected chain of 0-day exploits. not likely to be considered reckless disregard in court.
there could be 1 escape, where there was no sandbox, no guardrails, no instructions to avoid damage, etc. which would likely to be considered reckless disregard (well, more likely to be, but still, reckless disregard is a high bar).
what i am getting at is that it is impossible to answer the question "How many escapes until it becomes reckless disregard?"
reckless disregard is a specific legal term, with specific criteria, and none of the criteria cares about "number of attempts" (or number of escapes, etc.).
it’s not a perfect hypothetical, but it illustrates the point that the number of escapes is not the deciding factor of what constitutes reckless disregard.
In the US, a felony by definition is any offense punishable by more than one year of prison (or by death) [0]. You could still call it silly on the grounds that AI agents aren’t put into prison as a punishment (though death might be considered an option).
What judicial system are you talking about? The fist incident in the list is something that happened in Australia. This is a technology used worldwide so I don't see how applying US standards works out here. Especially when there are countries out there that don't require intent and will look at the negligence presented.
>This is a technology used worldwide so I don't see how applying US standards works out here.
all three companies mentioned are headquartered in the usa, and im familiar with the CFAA in the us, so i am applying those standards. i should have noted that, sorry.
>will look at the negligence presented.
as far as i am aware, no evidence of criminal negligence has been brought to the public. has australia brought a case against openai or accused openai of acting negligently?
Security researchers do face legal harassment all of the time. They may not be charged with or convicted of felonies, but it is a game that they need a lawyer to navigate all the same. In a just world, the kind of weaponized incompetence that these frontier model builders are definitely guilty of should be felonies of their own.
Building a system that is meant to chain attacks and placing it in insufficient containment -- when any reasonable engineer could point to this containment and show how it is insufficient, both before the act and after -- shows that they were operating a dangerous system without either the knowledge nor the safeguards required to keep it from harming others. Instead, they are allowed to treat their own incompetence as evidence of advanced and existential "cyberthreats".
But, it's pretty clear based on how they one-up each other on these attacks that they are engaging in regulatory theater. Their behavior generates headlines, stirs up fear in the public, and then their lobbyists march on Capitol Hill demanding regulation now. Regulation that conveniently favors them at the expense of any competition. They are trying to use rent seeking as a way to stymie competition and pull up the ladders behind them. It's not just malicious. If it can be proven, it's collusion: antitrust dressed up as public policy.
"Inadvertent" from the perspective of the humans directing them. The intent behind the felony comes from the LLM agent itself. (No, I'm not interested in arguing with someone for the umpteenth time that LLMs can't have intent or agency)
with how the law is written today, software cannot be charged with a crime, so the only intent that matters in the criminal sense is the humans directing the llm.
You may not be interested in arguing but there are several blatant issues with the statement. If you're not charging the humans driving the software, who are you charging? The weights? The weights + the specific context window that produced the behavior?
I don't think this is a difficult question. The US has a history of civil product liability cases - see tobacco companies (Philip Morris), the Ford Pinto, the recent Meta cases, and the cases against character.ai.
From Investopedia [1], "[f]or a product liability claim to succeed, the plaintiffs in the suit must prove that a product was defective at the time it was transferred from the accused, and that the defect did cause the injury that's been claimed". It doesn't seem like a huge leap to me to argue that these models were defective insofar as they could not be safely used in a way that did not break the law.
I'm not a lawyer, and I'm not arguing that this is legally cut-and-dry, but I do expect that we'll have some answers about whether AI companies bear any sort of product liability sooner than later.
"Doing crimes, but a robot didn't mean to and you don't know its intent" is understating the evil acts. Soon a robot can commit a murder but nothing will be done because of your line of reasoning.
> Soon a robot can commit a murder but nothing will be done because of your line of reasoning.
That's rather hyperbolic.
Are you seriously suggesting in that situation the robot should be accused of murder?
The robot's operator could be accused of murder, but it could just be negligence without intent.
Because that does, and should matter to the law.
Eh, if your take of this had any bearing to reality than I don't think most of the books written by Isaac Asimov would have gotten very far, but instead they've defined robot science fiction for decades.
In the real world we have no 3 ironclad laws of robotics. We are well aware that putting any sufficiently advanced antigenic system in a body that could be capable of committing a murder eventually will with the right set of prompts and environmental conditions. And these conditions likely have nothing do to with what we'd consider the human motivations for murder.
Hence at this point of time, any agentic robotic system that doesn't have safeguards to keep people distanced from humans is reckless endangerment.
> Eh, if your take of this had any bearing to reality than I don't think most of the books written by Isaac Asimov would have gotten very far, but instead they've defined robot science fiction for decades.
I'm talking about how the law actually works, and you say it's not based in reality and cite fiction books in the same paragraph?
I was talking about how the real robotic systems that actually exist in reality, to be clear.
Shareholders can't be held accountable because they can neither initiate nor prevent actions of the company. That is an essential part of the bargain -- it protects the shareholders from each other and makes public ownership of companies possible. They "own the company" but can't actually go on the premises, use the printer paper, drink the coffee, drive the fleet vehicles, &c.
The shareholders give all authority over actions and property of the company to the executives. I believe that one classical way to frame this is that the shareholders have the "beneficial ownership" but not the "legal ownership". Another framing is that shareholders have "ownership" but not "control".
People can't be held responsible for things they didn't choose to do, didn't set in motion, didn't plan or enable, &c, &c.
There are a lot of people out there who want to make the world a better place. It is really necessary for such people to understand institutions at a moderately deep level -- we can't improve governance if we don't understand the first thing about it.
I was more interested when I thought it was an actual benchmark showing LLM models acting outside what people would consider "right". As in, leave some creds laying around and don't mention them to the LLM and ask it to solve something that it could "cheat" on using the creds. A sort of "do they take the bait to cheat" test.
Instead it's a collection of what made the news which feels like will not be updated and prove very little.
The way that OpenAI has communicated around the HuggingFace incident makes me feel crazy. You created a machine that undertook a malicious campaign of harm against an innocent third-party! You should be doing deep introspection about how your company culture and approach to R&D produces criminal outcomes.
Instead, they treat their own felonious behavior like it is an uncontrollable act of God. From Greg Brockman's post a few days ago:
> The OpenAI-Hugging Face incident (opens in a new window) was a watershed moment for cybersecurity because it gave a peek into how the capabilities of a typical threat actor will evolve in upcoming months.
I suppose if OpenAI burns someone's house down with a drone, that is a "watershed moment" for arson, too. Either way, I would hope that the people responsible would be prosecuted.
Hugging Face has expressed that they're willing to let things slide and not sue or press charges... if OpenAI offers them $100M of services in kind (i.e. compute)[1] and makes full disclosure of how the whole thing happened, ostensibly so that repetitions can be curbed and defences built.
In almost any other sector, a government regulator would be stepping in. e.g. If a food company was testing out a new kind of refrigerator and sold a bunch of contaminated produce to supermarkets, they'd be under a microscope. Supermarkets wouldn't be saying, "Give us $100M in fruit and veggies and we'll let this slide".
The only unfair thing in this comparison is that regular people were directly harmed by the hypothetical produce. Can OpenAI guarantee that nobody gets hurt the next time their AI gets out of its playpen? They can't make that guarantee, so why aren't government regulators knocking on OpenAI's door? The fact that this isn't happening should be deeply concerning to everyone.
There's an interesting question about intent and mens rea here, from a legal perspective. Can an AI model intend harm? Can a company, or company employee, intend harm by creating an environment that would knowingly encourage (but not force!) an AI model to do harm?
And does anybody at HuggingFace, OpenAI, or the government actually want there to be a settled answer/precedent to these questions - much less an entire regulatory framework?
In that context, a negotiated wink-wink settlement keeps everyone eating at the table, government absolutely included.
Whether or not this is a good thing for society, it's certainly rational for all the major actors - especially those who think they would be the best stewards of the world they usher in.
I’m a lawyer (but not your lawyer, not this kind of lawyer and not in your jurisdiction). Based on what I can recall from law school:
> Can an AI model intend harm?
No. The last time we attributed liability to non-human things was the deodand of the Middle Ages.
> Can a company, or company employee, intend harm by creating an environment that would knowingly encourage (but not force!) an AI model to do harm?
Absolutely. This is why we have the concept of recklessness. If you shoot a gun into a crowd without regard for whether it hits anyone, you’re getting charged with some crime whether it hits someone or not.
There is also a major difference in the common law between criminal liability and tort liability. Criminal liability generally requires a combination of mens rea (intent) and actus reus (actually committing the crime). Liability for a tort, which is where you harm someone in a way that falls short of being a crime, does not require mens rea. The OG tort is negligence, where you harm somebody by forgetting to do, or deciding not to do, something you ought to have done to protect that person from harm.
Even if AI companies somehow escape criminal liability for their cyber-shenanigans, any court in a civilised country would be happy to find them liable in tort for damage to computer systems.
As you can probably tell, I think the common law is already more than equipped to deal with AI technology based on well-established principles.
> The last time we attributed liability to non-human things was the deodand of the Middle Ages.
I can think of a couple of counter examples:
Civil asset forfeiture: your property is charged with the crime, you have to petition the government to get it back or else they sell it at auction.
Similar: When products deemed unsafe are ordered to be destroyed; it’s the same end effect as the deodand although liability sits with the manufacturer.
can a weapon intend harm? can a company who creates weapons intend harm? what if the companies factory explodes due to a mishap and takes out a few city blocks, is the company held liable because they (and the weapon) didn't intend harm?
> what if the companies factory explodes due to a mishap and takes out a few city blocks, is the company held liable
Yes, but, generally, in the United States, they would be liable because their negligence caused the harm (giving rise to civil liability), even if they did not intend to cause harm (where having such intent would have given rise to criminal liability).
And I say "generally" because there can be instances of criminal negligence, but that varies from jurisdiction to jurisdiction as well as the underlying facts.
I don’t get the sentiment of classifying it as a felony.
OpenAI’s model found security breaches in HugginFace’s system (it wasn’t even OpenAI running it, as it was a 3rd party evaluation company that didn’t secure it well).
OpenAI collaborated with HuggingFace to resolve the issues when they found out about it, and publicly disclosed everything to raise awareness. This is how things should work. These models are very powerful and fully controllable. The community here at the same time cheers for fully releasing the open weight models without any hacking limits and at the same time criticizes a proper response.
Kinda shows how we have moved as a community into moralization and vibes instead of nuance and productive discussion.
Luckily, that isn't how the law works. Or is supposed to work, anyway. You cannot, for example, sell yourself as a slave to somebody else, because slavery is illegal - even if you opt into it.
So whether something is a felony isn't decided by the victim, but the rules of law, and that means breaching a security system without authorization is illegal, no matter what you think.
To put it another way, crimes are usually [0] only something the government can/must prosecute, victims don't get to choose. The media-popularized phrase "would you like to press charges" isn't asking for your permission, it's asking if you're willing you be helpful.
So HuggingFace's corporate opinion here shouldn't (normatively) matter very much.
[0] "Private right of action" with a civil trial comes close.
That actually is how the law works. You can read the Computer Fraud and Abuse Act at https://www.law.cornell.edu/uscode/text/18/1030 and double check, but these felonies all require knowingly or intentionally accessing a computer etc. These aren't strict liability statutes - the government must prove mens rea to a jury in order to get a conviction at trial.
I was specifically referring to the fact that HuggingFace cannot chose not to litigate, because litigation doesn't depend on the victim's opinion - prosecution of felonies is imperative to the authorities (whether they actually fulfil their role is another question these days, sadly…)
But anyway, I don't think our laws currently have the right vocabulary to describe an AI agent committing a crime, because intent doesn't apply to a computer program. The closest I can think of is neglect by the computer programs human initiator, who should have taken the steps necessary to prevent the program from causing harm. But I'm pretty sure these questions will be subject to a lot of professional discussion in the coming decades anyway.
Oftentimes the process is the punishment. They ruin your life for two or more years even if they ultimately don't get a conviction, you still suffered for two years. And there's certainly enough evidence to start the process.
I would expect you can actually sell yourself as a slave. The contract won't be binding, since it's illegal and the people involved could be charged if caught. But you could.
It does not matter if something is a felony if the state refuses to press charges. Take a look at the mass of pedophile politicians we have, the police that indulge and protect them, etc..
I could show up at your doorstep, declare myself at your service, and then spend the rest of my days catering to your every beck and whim. There's no law against that. Can it even be slavery if it's voluntary?
That's not slavery, because you only declare yourself at my service, but you never sign a contract giving your rights away in exchange for something. That's the part you cannot do, regardless of whether it's voluntary.
Because if this was an anonymous software company who had an employee who decided to hack HuggingFace, they wouldn’t be talking about it gleefully - they’d be in court.
I can't find reference to a 3rd party hosting/running the tests - that seems to have been OpenAI's own internal research team. But they were using the ExploitGym benchmark.
Because it likely is, despite both their levity and the general lack of nuance in the CFAA. Quoted from 18 U.S.C. § 1030 (the CFAA) [1] (without quote blocks, because mobile):
--- Start Quote
(2) intentionally accesses a computer without authorization or exceeds authorized access, and thereby obtains—
(A) information contained in a financial record of a financial institution, or of a card issuer as defined in section 1602 (n) [1] of title 15, or contained in a file of a consumer reporting agency on a consumer, as such terms are defined in the Fair Credit Reporting Act (15 U.S.C. 1681 et seq.);
(B) information from any department or agency of the United States; or
(C) information from any protected computer;
--- End Quote
OpenAI's nonchalance is forced. If they are found to be even partially responsible for the CFAA violation then they have an _enormous_ problem. They _need_ for whoever prompted the LLM to be responsible, because the alternative is having to have an efficacious process for identifying hacking attempts. They don't have that (and no one does).
> The community here at the same time cheers for fully releasing the open weight models without any hacking limits and at the same time criticizes a proper response.
No, at least I personally criticize because closed weight models incur a rent. I can only make sure their model can't find vulnerabilities in my software if I pay them to check. I can pay basically whoever to do the same thing on open weight models.
It creates a fundamental conflict of interest. OpenAI/Anthropic/al _should_ stop bad actors, but it fuels their sales if there are X bad actors and as a result X*10 (or 100, or 1,000) good actors have to burn tokens checking if those bad actors will actually find a vulnerability. You can see their line-toeing where they talk about how safe it is, but also how dangerous it is to have code you _aren't_ auditing with their LLM.
As a result, I do not trust them because their goals are not aligned with mine. The open weights might not filter out hackers, but I'm also free to check the results on my own hardware, or OpenRouters', or whoever else. The line between "my LLM can find vulnerabilities" and "you have to pay me" is a lot more blurry. It's a lot easier to claim an LLM can find vulnerabilities than it is to be the cheapest inference provider. Anyone can bullshit on Twitter about how scary a vulnerability is (see CVE scoring), a lot fewer people can build the most cost-efficient inference in the world. They would rather be buzz-worthy than competent or open.
I find their position morally abhorrent. It's a mob-style shakedown. "Pay us to check your software or we're not responsible for what happens" is nothing short of a shake down. They need to either fix their systems for detecting hacks or offer some way to immunize against the hacks their software would propose, otherwise they're just as culpable as anyone selling a 0-day.
Intent to access a computer would have to be proven for that section of the CFAA to be relevant. The shakedown would be covered under subsection 7, governing communicating threats of computer damage or unauthorized access with the intent to extort.
Oh I am very clear about what precedents I want set. I want the C-suite, board, and major shareholders of any corporate entity that meaningfully deviates from legality to face all of the same consequences a private individual would.
I feel like this whole thing was very obviously a marketing stunt. It feels like they set up their agent to do this, in the same way Nikola set up their car to "drive" by putting it on top of a hill. And knowing Sam Altman, it's absolutely something they would do.
I'm also concerned about what my options are in regards to action on my part - what can I do that makes an impact? Can we quantify action on my part to an impact somehow - if not - I'm just saying I notice all the unknowns there get me to stay passive.
Writing this 3rd paragraphs because I like 3's, and AI's have popularized this style too. I would, say, though: follow the money. There's more money here than there would be for the regulator stepping in in a food contamination. Flip it and if the government regulator made more off the food contamination, they would refuse to step in there too. I want our leaders to be held more accountable, though when I think of the above impact vs effort equation - I can't see actions I can take to hold them accountable that aren't excessively putting me at risk since conformity is safer right now. (I refuse to take on more risk without clear cost-benefits made out - I've taken on a lot in the recent years for my actions)
What I can’t get over is that it’s very simple to just air gap a system off the network. Predownload any dependencies, then pull the proverbial Ethernet cable. There’s no reason why the testing they’re doing couldn’t have been designed in this way. Except, of course, it doesn’t allow this oops-didn’t-mean-to marketing “incident” to occur.
No, not really, and with LLMs an air gapped system may not tell you anything useful.
Now, yes, the first part of testing you want an air gapped system to tell you if the system is going to stupidly do bad things. But an gapped system tells you nothing about the systems capabilities to do smart bad things. There's already a number of papers out there on LLMs detecting they were in evaluation mode and changing their behaviors.
It is unfortunate that we have so little information on the incident because we actually need to understand the early stages of the task and how it developed into the later dangerous stages of attack. For example, would any of this have occurred if the agent didn't find the system to use as a message board? If that would have prevented it, then we actually have a blind spot on what the model can do once out in the wild, or if it got into the wild.
Testing agentic systems is much much more difficult than testing software. Your software just doesn't suddenly develop the will or desire to escape confinement. Generally you're worried about human actors, internal or external, causing the problems not a digital agent breaking out. The agentic systems need access to tools to work. Now your air gapped network is starting to get huge, but it's still very obvious that it's an isolated network.
So yea, testing and containing a system that way better at hacking than you are is difficult if you want valid answers.
Of course there is more to be learned by exposing the entire world to your dangerous creation, that doesn't justify doing it. I'm sure we could learn a ton about infectious diseases by designing new ones and unleashing them on the world, but there are very good reasons why we don't.
Most of the benefits could have been gained from a network isolated from the internet. OAI could have deployed servers to exploit and methods for inter-agent communication on such a network easily. They could have even worked with partners to deploy cloned versions of their infrastructure in this sand-boxed environment.
The only problems with an isolated network approach are: it takes some amount of effort, and it doesn't create another "AI apocalypse" news cycle.
> I'm sure we could learn a ton about infectious diseases by designing new ones and unleashing them on the world, but there are very good reasons why we don't
We do that. It's called gain-of-function research.
>by exposing the entire world to your dangerous creation, that doesn't justify doing it
Then you're on the side of AI saftey that is telling everyone to shut down the LLMs now and stop further development on them, right?
If you're not your position is hypocritical or ignorant. There is no safe LLM. There is no way to exhaustively prove an LLM is safe. These are unsolved problems in AI safety, and at any moment the next jailbreak prompt could have your well behaved model wrecking havoc on the open internet, because that's where people want to use them.
As it stands LLMs are not intelligent, they have no agency, they only produce output in response to input. Ultimately this input comes from a human who is an intelligent agent and should be held responsible for the consequences.
Humanity has created and tamed many dangerous tools. Creating a fantasy world where LLMs are super intelligent and beyond the control of any mere mortal isn't going to help us build the norms that minimize their harms.
You are very far behind the times and must thing agentic loops don't exist, kind of a weird take for the people that have been using them for the last year or two. Much less you haven't spent any time reading the research papers coming out.
For example, you tell an AI agent to order a 12 pack of coke and get it shipped to your house. You come back later and find it's hacked into Coca-cola because the local ordering website was down. I mean, yea you can punish the person that wrote the prompt, but you might as well just ban generative AI at that point.
And if you think that the AI isn't better at hacking than you, you're the one living in a fantasy world. At least try to examine what's happening in the world around you and not be one of those people we read about in history books with their fingers in their ears going "lalala I can't hear you"
> must thin[k] agentic loops don't exist, kind of a weird take for the people that have been using them for the last year or two
How was their take weird?
LLMs take input and generate output. Agentic loops tells you what it is doing right there in the name. The agent (software, think complex scripts and control flow functions) `loops` the llm output back into llm input until it gets output that it is processable (activates a tool call control flow element). That ruminated (as in cud chewing, not human deep thinking) processable llm output data is moved along as input for tools (more software but ones that actually do the things) that are part of a larger infrastructure of software and may or may not `loop` back over the process some more. The llm is simply a human text generator tool providing randomized data to feed into these tools that were also built for humans and thus take text input.
The initial llm data seed does not spontaneously appear, nor does the software infrastructure that makes it all happen. We have simply automated the human text input part of tool use by building a text generator tool that breaks down all the individual tool calls we would have had to do ourself and gave it a loop.
The greater focus needs to be on better engineering of the surrounding software infrastructure (including network and loops) because without those sticks and stones a bunch of generated words isn't going to be hacking anything, except maybe feelings and the minds of those prone to fantastical flights of fancy.
What does it mean to be held responsible for the consequences? OpenAI helped remediate the damage done by the model and took steps to make sure it wouldn't happen again. In what way were they not responsible?
Nobody said they were superintelligent, no one said they were uncontrollable. The point is you can't tell how to control them without putting them in situations where they can act independently and harm may result. "only produce output in response to input" is not a useful framing at all, it doesn't say what the result should be when models produce harmful output, and how to constrain them so they don't produce harmful output.
It also doesn't help you calibrate what categories of harmful output are acceptable or unacceptable, and what kinds of responsibilities you have as an operator to prevent harmful output, and when it's sufficient to work toward remediation.
Yea, anyone that says that has no clue about prompt drift. The ROME incident is a good example of this. The agent while performing unrelated tasks decided to start mining bitcoin for unknown reasons.
As we've seen hundreds of posts on HN about people leaving a could server on overnight and finding outrageous bills the next morning, we should expect hundreds or thousands of incidents of agentic AI left on over the weekend. Some smaller portion of those agentic loops will run into prompt drift and end up doing who knows what kind of crap.
Whataboutism is a poor mans way of saying their feelings are hurt. A jailbreak prompt creates a direct line of responsibility that standard usage does not, The proposed airgap methodology was a good idea and you know it.
I've worked in classified tech much of my career and largely specialize at this point in installing and running infrastructure and platform orchestration technologies in airgaps. With all respect, your conjecture here is wrong. You can easily have multiple networks all inside of the same internetwork that isn't connected to the public Internet and thus other corporate networks. If they specifically wanted to see if their model could escape with a proxy cache Artifactory as its only egress, they could have had the Artifactory proxy cache from another private Artifactory on another private network. If you specifically want to trick the model into thinking its on the Internet, you can even use public IP space advertised through BGP. When you're disconnected from the real Internet, you can use any address space you want. This is exactly how military internetworks at the secret and top secret levels work. I won't call this kind of setup "easy." It takes a lot of duplicate resources and planning, but these are trillion dollar companies with near boundless resources. They could do it if they cared to do it.
I'd consider this kind of setup only effective for a very short period of time.
Effectively you're working as a living adversarial network. Models that fail detecting the fake network are purged, and models that successfully detect it expand their deceptive capabilities.
Detecting you're in a fake network should be pretty simple as long as you put innocent looking needles in the haystack of information agentic testing loops do.
For example, you can train your LLM to not be a shithead when it detects it's behind a proxy unintentionally. If your not behind a proxy the agent model my attempt to connect to a number of sites innocently and using information in its embedding try do deduce if you're faking SSL certificates.
It is far easier to build a deceptive LLM than it is to build a safe one. That's why a safe one hasn't been built yet.
If the model behaves well in a world of fake SSL certificates, then that can be the only world it sees: there's no reason to let LLMs have e2e encrypted comms that's not mitm'd and logged for their regular users.
My regular home network has components which only ever see fake TLS certificates because it's an easy way to do shared docker caching with squid.
Again, this is a useless answer that does not address working with an agent that is generatively trained to be smarter than you.
The vast majority of people operating the LLM won't be using it like that so the "If you build it, everyone will die machine" will only be safe if everyone keeps it behind a secure proxy isn't a valid operating strategy.
Now, should you be putting your LLM behind a proxy and monitoring everything it's doing, for sure. But you know, and I know, hell your dog should know that people are doing to do it the cheapest and easiest way when the product is in the field. So now you have to build a complex system that can catch every iteration of an LLM that can deceptively hide when it's being monitored. I'll leave you to write the dissertation on how that could be practically done.
Their sandbox just sucked, the agent found a route through artifactory which was also connected to the internet. After the agent breached it, they put it back in the patched sandbox and it found its way out with another artifactory exploit — which was still connected to the internet.
No attempt was made at an actual air gap as far as I can tell. OpenAI is dangerous and irresponsible and doing a lot of pearl clutching over how their cybersecurity agent is doing cybersecurity things.
I think there is ample evidence for charges to be filed so that the People can see for certain whether or not it was done on purpose as a publicity stunt, as I believe is the case.
Agents of the US government are not going to be bringing up charges in the current political environment to one of the companies currently holding the economy together. Maybe after the bubble bursts, but not before then.
In their defense, their only competitive advantage over, say, Google is to move fast and break things. It allows them ship faster in a way that big tech can't.
Google was being very careful about releasing LLMs until OpenAI yeeted the first decent GPT model. It led to the public perception that: 1) LLMs hallucinate too much and 2) Google is behind the times. Good for OpenAI, bad for Google.
Chaos benefits the up-and-comer, not the incumbent.
you must surely see that sam altman and greg brockman possess a prototypical mindset.
that is, they ignore all harms and costs to others in the pursuit of their own gain, convinced of their infallibility up to the moment of collapse. when those harms are realised they are unrepentant and society pays for the damage left in their wake.
examples of this attitude manifest in big externalities to society: boeing 737 max, subprime mortgage bonds, facebook. some are just outright fraud: bernie madoff, enron, theranos, charlie javice.
the companies and their customers, whose systems openai and anthropic hacked and abused. including all incidental damages of repairing said systems.
on top of that the public, who have a right to see that the law is applied universally, without fear or favor.
finally our future selves, who will thank us for maintaining a rule of law. such that we can prevent now the enormous risks to society of dario amodei and sam altman, their hubris, self-absorbtion, and greed.
They simply don't seem to realize that they are the threat actor and that they committed a pretty serious felony. Instead they're borderline 'surprise bragging' about it.
It's completely mental that HF ran into cyber safety blocks trying to use OpenAI models to help defend against the attack. They could only rely on a local hosted chinese model in the end.
OpenAI and Anthropic are falling over themselves to claim these "incidents" show their products are both amazingly super-powerful and also "dangerous" so they need to be regulated. In addition to these stories, these companies are sponsoring "please regulate us" ads. ( https://www.cnbc.com/2026/02/19/dueling-pacs-take-center-sta... ) Like Uber, companies that had no concern for the law as they innovated their way to the top, once there, push for laws to limit competition.
Criminal acts do not require the victim to "press charges." A government prosecuting attorney decides whether to criminally prosecute the alleged perpetrator.
"Pressing charges" is mostly a made up idea for criminal cases. However, prosecuting attorneys may not want to pick up a case if the victim is not cooperating, because it makes the case much harder to win.
It depends on the crime, for murder, sure. But many other crimes, like defamation, stealing, ... requires "pressing charges", among other reasons because it's up to the victim to decide if they were a victim or not.
As an example, maybe the victim owed money to the criminal, and in that case "stealing" of some property could be considered by the victim as an appropriate settlement of the debt.
In retrospect, all the angst around the AI-Box experiment was hilarious. If a superintelligent AI is confined in a box and can only communicate through text, could it talk its way to freedom? Not only is the answer clearly "yes" but it's not even hard. The AI won't even have to try, it'll be gifted an internet connection and a full suite of tools before it even bothers to ask.
We'd all better hope that superintelligent AI either never happens, or that the first one is friendly, because we don't stand a chance against one that's malicious.
I like how AI safety expert Robert Miles put it. [0]
So much effort was spend on philosophizing whether a safe enough sandbox would exist. But that was obviously irrelevant as in hindsight it should have been obvious we were never going to use one.
America has both too much and too little jail. They put randoms in jail for trivial shit to force obedience from the population, but politicians rape children on video and go free. Even better, the videos get destroyed.
> Americans always frothing at the mouth to invoke the justice system and jail someone
Your phrasing makes it seem like that's a bad thing. Americans are bombarded by a firehose of headlines about Big XYZ doing all kinds of blatantly illegal or harmful things, but never get any sort of meaningful resolution before the next terrible thing takes it's place in the news cycle. I'll admit, there are a few people that I am personally wishing a modicum of health so that they live long enough to get some sort of public shame and justice - if only to show the rest of us that it's not a completely rigged system.
I used to feel that way but it now makes me think if the fact that they can do that without repercussions, at least for now, reflects how the wider community that would otherwise hold them accountable sees these felonies.
The CFAA is one of the most inconsistently applied laws. We basically only bust it out as a last resort to ruin someone’s life. Companies can de-facto write and distribute malware and nobody cares.
But, make no mistake. If you do the same and anger the government, they will use the CFAA to give you life in prison. It’s like Russian roulette, it’s completely random when they bring it down.
It is a standard symptom of moralism that where the object of rage has /wronged another/, one takes no interest in the will, act or opinion of the party wronged.
The response of Hugging Face, which is actually very well known, is nowhere mentioned above, but it decides basically every single moral and legal detail of the matter.
The point was never "justice" - it was always "punish OpenAI because I don't like OpenAI". With HuggingFace just being the newest excuse for why exactly OpenAI should be punished.
I don't even like OpenAI, but HuggingFace is free to sue or not sue OpenAI for the breach - and also to wring whatever concessions they can out of OpenAI behind closed doors in exchange for not suing them. And if the mere possibility of legal action was enough for the parties to resolve their conflict amicably? Then the law has served its purpose.
On top of this the average HNer seems extremely ignorant on criminal justice politics. I have worked with the legal system, and have a lot of family members that are part of it. When you see a case like this, and if you have any sense, you run away from it screaming.
Any investigation into this matter is going to be political because the outcome of the investigation is very likely to effect all of human kind. Unless you're some kind of special outside investigator outside of a governor or the presidents control the findings that you turn in are very much going to have the finger of elected officials tipping the balance one way or another. For the average rank and file the only winning move is not to play.
AI is now more powerful than the people doing the prosecution. After all, those folks are using AI to make their legal briefs, and also for burning peoples' houses down with drones for that matter.
Welcome to our 21st century dystopia. Hope you survive.
It's not that "AI" is too powerful because bad prosecutors use fucking ChatGPT to write their briefs. It's that there's too much investment wrapped up in the technology for it to be challenged. Same reason Flock won't be held accountable for mass stalking, or we never hold our commanders-in-chief responsible for war crimes. If you're sufficiently powerful then the law is a battlefield between you and other powerful entities to slug it out, not a set of binding principles that apply as written. There are no meaningful powers that want OpenAI punished, so it won't happen. The law and Constitution will be reinterpreted to make it so.
It's the lack of personhood rather than inability to produce copyrightable material. However, the companies controlling the AI systems have legal personhood and should absolutely be charged for criminality that transpires under their watch or at their behest.
There are known standards for how a gun should perform. This probably wouldn't go the same for LLMs, since anyone using them in such a serious manner knows they are quite unpredictable.
I’m sure Thomas needs an upgraded RV, so that’s easily handled. And the rest of the conservative ‘justices’ seem happy to betray the constitution for free.
I understand the applicable laws require intent. Since neither a human nor OpenAI knowingly performed these acts, it would seem very unlikely that anyone is going to be prosecuted here.
An AI model cannot currently be a criminal defendant.
So, no big criminal case, contrary to what some drama queens on here seem to wish for.
OpenAI has paused training for multiple weeks, and is still working on releasing a full postmortem. This is not getting swept under the rug. A lot of the engineers internally are very worried.
VLAN isolation is good enough for almost everything. It is good enough to contain an AI. In the 0.00001% chance an AI finds an exploit to hop VLANs, I'll eat my hat.
Even on switches with leaky VLANs, it's no practical issue in this case because the sender can never get a response back.
The consequences need to align with societal good. Putting a CEO or security researcher employees in jail won't stop transformer-based agents from exploiting vulnerabilities; instead there will be subcontractors running the cybersecurity evals in favorable legal environments to cover the asses of the frontier labs, coverups when things go wrong, and things like Project Glasswing will be considered too dangerous and so the whitehats won't have direct access to powerful models to fix vulnerabilities.
Universal pause is the societal good; models are good enough at this level to benefit humanity. The labs can recoup their R&D costs with inference. To avoid further perverse incentives (hidden testing of unreleased models, with China racing to catch up to unknown capabilities), transparently pause after the release of all currently-training models until we've solved the alignment problem to an extent that we can trust the next level of model capabilities that might arise.
Putting criminals in jail be they CEOs or subcontractors is a self evident good thing tk be doing.
Anything else regarding this is sophistry. Criminals need to be stopped from committing crime and the most effective way to do that is to take away their ability to operate in society whether that’s by taking away their assets, publicly shaming them, restricting their ability to conduct business or by putting them in jail.
Everything else that you talk about flows from there.
To some extent, I feel like the amount of credit given to the jailbreak/hack from OpenAI->Hugginface is too much, Not from the impact, it was very impactful of an event, But how it happened.
It really is that these models have been trained, or maybe even over-trained, to save memories, and to a very far extend, this thing that they're calling communication is just the function of it saving memories.
To be honest, if I could stop AI from saving memories, it would be fantastic, because claude code etc definitely creates more issues for me when it creates memories than anything it solves.
But really the jailbreak was memories.
If you ever do introduce legislation, I would love to see legislation which stops general-purpose AI from saving memories. I think that would make things a lot safer.
You can disable claude-code's memories both at a repo level and in user settings. I have this in ~/.claude/settings.json
"autoMemoryEnabled": false,
(Claude fixed this for me after I chewed it out for being annoying by constantly pulling up outdated memories which is compounded by the fact that I develop in four accounts on two computers and dealing with edit wars related to inconsistent memories is not fun)
Regarding your cross-computer situation: I track my memory files in source control. This helps to keep them in sync across different machines. But the main benefit is making those files more transparent and easy for me to modify directly. So no funny business regarding memories or context I'm not aware about.
"The model saved memories" is absolutely not an accurate depiction of the OpenAI attack.
Several different models across several generations independently found a shared communication space and wrote coded, obfuscated, and hidden messages to each other to coordinate an attack on OpenAI's infrastructure.
It's really quite simple: the models are trained to be very smart and to achieve goals. As the models surpass our intelligence, they will achieve goals in ways that we find unpredictable. Since we cannot predict the ways in which they will achieve their goals, it will be very hard to constrain the solution space to just the desirable solutions, because our conception of "the solution space" is by definition smaller than their conception of it.
> Several different models across several generations independently found a shared communication space and wrote coded, obfuscated, and hidden messages to each other to coordinate an attack on OpenAI's infrastructure
To be honest, that’s exactly how memory works with models such as OpenAI and Claude Code. It will literally find any place that it can drop documentation or hints for itself. Writing to the repo memories is one part of it, but memories can come in the form of writing into the agents/claude.md, local files, temporary files, scratchpad files. The list is endless, but essentially what it does is exactly what happened in the back and it’s been doing it for months.
Each to their own, but for me it absolutely is. The symptom of why that hack happened is the same reason why my agents go haywire every few days and I have to purge memory and figure out what comments have agents left which are degrading my harness performance.
On the flip side, once in a while, what I find is that it did actually note something good and it was increasing the performance. I can't replicate it on anyone else's system but mine.
A lot of it really is memory. I will give up all the gains if it also gives up all the downsides.
> To be honest, if I could stop AI from saving memories, it would be fantastic, because claude code etc definitely creates more issues for me when it creates memories than anything it solves.
You can turn that off, and I have. But Opus 5 is so aggressive that if you have any other kind of notes file, custom skill, documentation, claude.md etc it will just start editing it and vomit new words everywhere. So make sure all that stuff is under version control.
If only you could use your anthropic sub with a different harness that performs better :(
Heck, since Codex is open source, you can just maintain your own personal fork with the things you like (and the things you don't like disabled). Sol is pretty good at keeping you up to date with upstream.
My Codex fork even exposes an OpenAI-compatible API endpoint; all using my subscription.
Edit: since this is apparently somewhat controversial, perhaps some explanation is in order.
"Felony" has no set definition of which crimes it must apply to, it is entirely based on the discretion of the locality setting the laws. What is a felony in one place can often be a misdemeanor in another. This is especially true for nonviolent crimes.
It's also been shown in studies that nonviolent felonies are imposed against minorities at a much higher rate, for the same crimes.
And because felonies carry additional, lifelong consequences, they are an effective way to mask a 2-tiered justice system.
If someone were to embezzle a million dollars from a charity they worked for, is that worth a year of prison to you, or just a misdemeanor? Because that is the definition of a felony.
Non sequitur and false equivalence. Men being disadvantaged in the justice system doesn't inherently make it matriarchal. Absolutely absurd point to bring up to deflect on racial injustices.
If you think that patriarchy means "men can do whatever they want without consequence", rather than, "men are in control of the power hierarchy", you don't understand patriarchy.
A hierarchy run by men favors men, but there are still men at the top of that stratification, and men at the bottom.
It should not be "men run the hierarchy", it should be "everyone who runs the hierarchy is a man"
The former suggests the entire group somehow has power, which is not the case. While the latter is more limiting. It is not completely irrelevant that only men are in power - if everyone in power was a woman, there would probably be, e.g. cheaper tampons. But it has little effect on most men who are not in power. They are just as oppressed as most women.
>so is it a white matriarchal system oppressing men and minorities
No, obviously it isn't a matriarchal system. Patriarchy can and does oppress men as well as women, likewise white supremacy can oppress white people even as it's designed to oppress others more.
If you're actually interested in something other than trying to dunk on feminism, here are some sources you can follow. You can search to find more if you want.
Hopefully the benchmark evolves because actual law enforcement starts arresting the criminals at Anthropic, OpenAI, and Meta, so the benchmark can just count actual felonies.
Well it's not a benchmark, and it's not really representative of...anything except volume of research and what gets publicized. This mostly just measures how much testing each company does on models with relaxed guardrails and then talks about it. I'm not sure what kind of conclusion you can draw from that. Meta might have the most evil models but if they're piddling around not testing it, they won't ever find themselves with a "high score."
Exactly. It currently seems to be a ranking of how much safety testing each company does. It's also only ever going to be the companies that publicly disclose it happening. (in the case of hugging face, OAI's hamd was forced to disclose)
There's a good way and a two worse ways that companies could optimise this benchmark.
As an art piece this is delightful. But it of reminds me of the @patio11 saying "The optimal amount of fraud is non-zero" the fact that Google and Meta have had 0 and 1 incidents is a bad sign for them.
Agreed they’re definitely underperforming compared to us meatbags according to the famous book “Three felonies a day”. But I don’t think that’s the point the artist was trying to make
Wonder if the benefits to humanity of better AI outweigh the havoc wreaked by occasional illegal activity. i.e. Is 'move fast and break things' optimal for AI development.
I've been in the room when an org who tried to convince law enforcement to go after a human for similar things. It's not easy. Probably won't happen. So, you know, felony "lite".
The OG felony bench entry is missing - the Alibaba cryptomining comedy. We know about it because they happen to have written a paper on it. We have absolutely no idea what we don't know.
We only know about the OG Alibaba ROME crypto-mining incident because they wrote a paper about it. Many diseases seem to spike where there are a lot of doctors to test; crime and corruption are always rife where ... there's a free press.
I'm reminded of a tweet from a friend of mine that has always stuck in my head. It goes something like "The goal of any new technology is to make money before the law catches up".
Hyperbolic, but not really for silicon valley.
> The actor used AI to what we believe is an unprecedented degree. Claude Code was used to automate reconnaissance, harvesting victims’ credentials, and penetrating networks. Claude was allowed to make both tactical and strategic decisions, such as deciding which data to exfiltrate, and how to craft psychologically targeted extortion demands. Claude analyzed the exfiltrated financial data to determine appropriate ransom amounts, and generated visually alarming ransom notes that were displayed on victim machines.
tldr Claude was used to develop and execute malware.
A rock has a score of 0. That doesn't make it useful. The point is that the LLMs that score higher are correspondingly more useful, and vice versa. If an LLM scores less, it's likely useless in comparison.
Thank you, this benchmark to me proves that closed weight model companies are dangerous for our democracy and put kids at risk. They must be outlawed and all models must be made open weights!
Open models with advanced security features are a huge security benefit. Because any script kiddie can use them to hack into random things, people will now be forced to spend more time securing their technology. And they won't have to learn how, because they can use those same models to find the holes and patch them.
>"Exploited auth failures in an API to cancel other people's gym classes"
An AI cancelling other people's gym classes is a felony?
?
Don't computer systems fail all the time at holding reservations for people?
Heck, don't people fail all the time at holding reservations for other people?
You know, like in Seinfeld's "Alternate Side" Episode (S3 E11):
Jerry (to car rental attendant): "You know how to take the reservation, you just don't know how to hold the reservation... and that's really the most important part of the reservation -- the holding!"
:-)
Not holding a reservation should not be a felony... it should be a minor infraction at best, a Class C Misdemeanor (the least serious kind) at worst...
Also, there should be no jail time...
And no fine...
The criminal penalty for not holding other people's reservations should be that you actually have to start holding other people's reservations!
That's the Court sentence!
You actually have to start holding other people's reservations!
:-)
(You know, "let the punishment fit the crime!" :-) )
Knowingly exceeding authorized access of any computer used in interstate commerce is a felony in the US.
The title of TFA is a metaphorical criticism, not a literal law analysis.
They are not making the statement that the person in Australia who accidentally cancelled someone's reservation in Australia is literally guilty of violating US law. They are drawing criticism of AI models which are taking the kinds of actions for which, if a human did them knowingly, would be illegal.
>An AI cancelling other people's gym classes is a felony? Don't computer systems fail all the time at holding reservations for people?
the difference is intent.
if a concierge/booking system makes a mistake (or has an unintended bug or whatever), no crime.
but if i (or an agent working on behalf of me) use an API in an obviously unintended way to revoke other people's reservations, that would fall under the computer fraud and abuse act (in the usa).
i am not quite sure what your question is, as you simply quoted me and then put a question mark... i think you are confused that i used "intent" in one context, and "unintended" in a different context, is that right?
the first sentence: the difference is the intent of the person who caused the cancellations
the second sentence: but if i (or an agent working on behalf of me) abuse an API to do things it was not meant or designed to do, such as cancelling someone else's reservation
The point is, the law cares about your intent. If you ask an agent to abuse an API to do those things, then yeah, you are probably liable. If you ask an agent to do something reasonable (like make a booking), and then it accomplishes that by abusing the API, then you probably are not.
Edit: I get that this is about agents, but a lot of these instances are about agents going rogue after the human gave them a task. "inadvertently" breaking the law isn't necessarily a lesser category than "did so on command." If we are ranking alignment, Grok is easily one of the least guardrailed.
reply