logo

Gambling with our lives: AI researcher quits Anthropic with warning about safety

Posted by taubek |2 hours ago |56 comments

themgt 37 minutes ago[6 more]

The fundamental point I think is far too often confused is the difference between LLM and agentic system.

An LLM can't do anything but generate tokens. You run your LLM in vLLM or whatever, and it generates output tokens based on your input tokens. That's it!

Humans then build ~deterministic systems to take those tokens and do all sorts of things with the tokens, like take actions in the real world. And then we can feed the output of those actions back to the LLM, and generate more tokens. And then our systems can use the new tokens to take new actions in the real world.

Humans want to blame "AI" for attacking HuggingFace or a German wiki or whatever, but:

1) LLM - can't take over german wiki because it just generates tokens

2) agentic system with internet access, a prompt telling it to attack stuff, running in a shared CI env so agents can whiteboard in artifactory

None of 2 is "AI", its standard networking and Markdown and CI virtual machine, etc etc. There's no AI to be found. CPUs not GPUs, even. Just deterministic systems ultimately managed by humans. And a 10x more powerful system-1 can still just generate 10x "smarter" inert data.

If humanity and human organizations collectively decide to yolo the tokens generated from system-1 into our deterministic system-2s, over which we have complete control, back to system-1s, in a yolo loop, in such a way we lose control and it ends humanity, well ...

"The coin don't have no say. It's just you."

koolba an hour ago[1 more]

> Evan Hubinger, Anthropic's staff lead on keeping the technology aligned with human goals and values, backed up Coxon claims in a follow-up post of his own, though he didn't quit the company.

> "Jacob is correct here — we really do earnestly believe AI could kill all humans," he said.

> Hubinger estimated the chances of that happening to be higher than ten percent within the next decade, and added that there's no plan yet on how to keep AI aligned with human goals in the superintelligence scenario.

10% chance we kill everybody is a small price to pay for Motown remixes of classic 2pac songs.

brunorsini 34 minutes ago[1 more]

I particularly like the last point he makes here:

Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources.

It's interesting to see so much hate towards creators who use AI to make almost any type of creative work. At least these are humans using it as "controllable tools". Nuclear-powered bicycles for the mind.

As the degree of separation increases, things can get interesting. "Create several social media accounts, post whatever, maximize views and engagement, give me back the aggregate numbers". "Now promote <x>."

And then decisions like that may soon be spawning autonomously, as just another step in a reasoning series aiming to achieve some other, broader goal.

Open models/weights may end up playing particularly important roles here. Users may, knowingly or not, bypass system prompt-derived safety that could technically have offered much needed protection.

cranx 25 minutes ago

At this point I feel like the boy who cried wolf is an ai researcher. Yes super intelligence could be dangerous. However, is the story a PR stunt or real? Well…

WalterGR an hour ago

“I resigned from Anthropic today” (twitter.com/hilbertspaess)

https://news.ycombinator.com/item?id=49619227

564 points | 9 hours ago | 766 comments

wartywhoa23 22 minutes ago

> "The people building AI earnestly believe that it could kill us all by the end of the decade," Coxon said in a follow-up post.

AI has no incentive to decimate extra 7.5B mouths to feed, those who print money out of thin air to make an AI cover for that decimation have.

MrThoughtful 39 minutes ago[7 more]

Why would AI wipe us out?

We have not wiped out apes, ants, and most other species.

We even have discussions about how to actively save them from extinction.

glimshe an hour ago[5 more]

What do people feel about this in China? Even if their models are well behind, they are not years behind. If we restrain US companies, assuming that is desirable, it would do nothing to deter China's and AI-pocalypse would come anyway in short notice.

pluc 41 minutes ago

Cool cool cool cool

virgildotcodes an hour ago[1 more]

At least we’ll eventually have an entity other than ourselves to blame for our annihilation.

localhoster 43 minutes ago

I honestly feel that all those big ai companies think AI will long term harm humanity, but not their ai.

A classic "it will not happen to me"

Bengalilol an hour ago[6 more]

> Both OpenAI and Anthropic have recently flagged incidents in which agents powered by their models went rogue

I may be biased and somewhat off topic, but I see these incidents as some of the most significant of the past century. I genuinely don't understand why these companies aren't taking a smarter approach to them.

The latest analyses have been, at best, laughable: identify the vulnerability, patch it, and move on. Only to repeat the same cycle without considering that there may be something far more serious at play.

These are AI security experts, and this has been their way of "solving" these incidents. AI security experts ...

Moreover, when Challenger exploded, the government launched a series of investigations into the incident, bringing in experts from across the field. And now, what has the government done? Nothing. Literally nothing, as if everything was fine and all under control.

Seriously, I'm generally quite optimistic and I don't buy into this fatalistic narrative about our shared future. But I have to admit that sometimes I feel like I'm stranded on a planet of primates.

Sorry for this rather unproductive rant.

archerx 44 minutes ago[1 more]

An AI that generates text will never be scary to me. An autonomous AI with facial recognition on a flying drone with weapons (bombs/guns) with swarming capabilities will always be terrifying. I feel like we are ignoring the massive elephant in the room.

jongjong 40 minutes ago

I'm not worried about AI safety. People greatly overestimate the utility and capabilities of intelligence. I'm not afraid of intelligence, I'm afraid of idiocy.

tilolebo an hour ago[3 more]

I thought the plan was sandboxes and markdown files to tell AI agents not to be bad. Is that not enough? /s

feverzsj an hour ago[1 more]

It's your fault to handle over your secrets to them.