Ex-Anthropic insider tells CNN how AI could kill all humans by 2030

Transcript

[00:00] Jacob Coxin has left Anthropic sounding the alarm 
about the dangers of artificial intelligence on  
[00:05] his way out the door. The 27-year-old outlined 
his concerns in a series of social media posts  
[00:10] when he quit, writing in part, “The people 
building AI earnestly believe that it  
[00:15] could kill us all by the end of the decade.” 
This is not a marketing stunt. If anything,  
[00:19] many executives and senior researchers will couch 
their phrasing in the press to sound sensible. But  
[00:25] I hear the same people express fear privately. No 
other human activity poses this level of danger.  
[00:31] At least one of his former colleagues agrees. 
Evan Hubbinger, an alignment scient at Anthropic,  
[00:38] shared Coxin’s post, adding, “Jacob is correct 
here. We really do earnestly believe AI could kill  
[00:44] all humans. I personally think it is greater than 
10% within the next decade. I believe anthropic is  
[00:51] light is trying its best, but we do not yet have 
a plan to solve alignment for super intelligence  
[00:57] and are not clearly on track to. An anthropic 
spokesperson tells CNN, quote, “We have always  
[01:02] been transparent that AI will bring both enormous 
benefits and unprecedented risks. To address these  
[01:08] risks, we continue to build models with some of 
the strongest safeguards in the industry.” They  
[01:12] go on to say, “We were the first lab to publish 
a responsible scaling policy, a public framework  
[01:17] dedicated to mitigating catastrophic risk from AI 
models, and we continue to aggressively test our  
[01:23] models for dangerous capabilities in areas like 
cyber security and biology and publish what we  
[01:28] learned for scrutiny and research.” Now, for his 
part, Coxin has worked as a researcher at both  
[01:33] anthropic and open AI. He says neither company is 
acting uh responsibly. Jacob Coxin joins us now.  
[01:41] Um, Jacob, I appreciate you you being with us. Um, 
first of all, I assume you are saying literally  
[01:48] that you believe uh that AI could kill us all by 
the end of the decade. How would that happen? Like  
[01:57] what would that actually look like? Yeah. So, the 
precise scenario sounds a little bit like science  
[02:04] fiction. Uh it sounds like something that’s not 
real, but I think it is frighteningly real. And  
[02:09] it’s it’s easiest to understand it if you think 
about real things that happened two months ago.  
[02:13] So two months ago, o open AI agents uh AIS hacked 
into third party infrastructure entirely of their  
[02:20] own accord. And this was like a concentrated 
hacking spree that they carried out of their  
[02:25] own valition. And I think that if you extrapolate 
into the future the level of capabilities of these  
[02:30] AIs with the same independent valition, they 
could cause extreme havoc. for example, hacking  
[02:35] critical infrastructure, building extinction 
level bioweapons. There’s a lot of uh ways that  
[02:43] the AI could actuate itself in the world. I is 
it specifically when did your feelings about this  
[02:49] technology change? Was it something about the rate 
of progress that you’ve seen? Yeah. So I think  
[02:56] it’s really important to just look at the rate of 
progress and say for example in areas like coding  
[03:01] or math. A couple years ago these AIs were just 
about helping humans a little bit. They could give  
[03:05] you suggestions. Now they’re close to replacing 
them. We still need humans at the moment but  
[03:10] quite plausibly within a year we’ll no longer need 
humans for doing research in many areas. Only on  
[03:16] Tuesday OpenAI solved a millennium problem. One of 
the biggest unsolved open problems in mathematics  
[03:21] purely autonomously using an AI. And I think 
what’s most scary is if AI is used to make itself  
[03:27] more intelligent. So you can take an AI and give 
it the problem of AI research. Just like currently  
[03:33] they’ve given it math research. You could make it 
do the tasks that we’re currently doing. And then  
[03:37] you get what’s called an intelligence explosion. 
The AI just gets smarter and smarter with no human  
[03:42] involvement necessary until you have a thing 
that is vastly smarter than than humans. And  
[03:47] this could be coming very soon. We mentioned 
a former colleague of yours currently works at  
[03:52] Anthropic. basically endorsing the assessment 
that AI could potentially kill everyone within  
[03:56] a decade. In a follow-up post, he said, “To be 
clear, as we say in our in our latest risk report,  
[04:01] I think the risk from present models is low. What 
I’m worried about is super intelligence arising  
[04:06] from recursive self-improvement as we’ve said is 
happening faster than than we thought. I assume  
[04:11] that’s what you the recursive self-improvement 
is what you’ve just spoken about.” Um you do you  
[04:16] agree with him that the current models the risk is 
low but it’s it’s just the speed with which this  
[04:22] is now improving? Exactly. I completely agree. 
It’s right now there’s no risk of extinction. The  
[04:29] current models the worst they can do is maybe hack 
into something potentially cause a lot of damages  
[04:35] in in infrastructure but there’s no they’re not 
intelligent enough to outsmart us at the level  
[04:40] that would lead to to extinction. I think what’s 
just crazy is to look at the the rate of progress  
[04:46] and there is a very real possibility that in the 
immediate future these are years that like next  
[04:52] year the year after recursive self-improvement 
will happen and we’ll enter the phase of Evans  
[04:57] post where he argues that there’s a chance we 
could all die like that that is coming soon you  
[05:02] know a lot of the these CEOs of these companies 
will you know will go to Congress and say hey we  
[05:08] would love regulation anthropic in their new 
statement right now says is the world would  
[05:12] benefit from the industry adopting a lawful 
verifiable way to work together to pace how  
[05:17] we release powerful models. Um you know the 
argument it’s an arms race if the US if a US  
[05:23] company doesn’t achieve the best AI a Chinese 
company or some bad actors out there might. Do  
[05:29] you believe that uh that these companies actually 
do want regulation or is it lip service? They  
[05:38] they completely do. This is completely honest. 
Just in the same way that the response to Evans  
[05:42] post was sort of surprise. It was like this 
can’t be real. Surely a guy working on this  
[05:46] technology can’t actually believe there’s this 
level of danger. No, that’s completely genuine.  
[05:51] And these people are also completely genuine 
when they are begging to be regulated. This is  
[05:55] a genuine. and they they find themselves in this 
scenario where they’re compelled to race towards  
[06:00] building a deadly technology and they would they 
would love for some sort of international body to  
[06:06] allow them to to approach it at a reasonable 
pace. They just don’t trust that the people  
[06:10] around them are going to get there in a safe 
way. So, they feel like they have to race there,  
[06:14] too. You you could have stayed quiet. I mean, I’m 
sure you had a great job being paid, I imagine,  
[06:20] very well. You’re obviously incredibly smart and 
mathematician. Um, why did you decide to speak  
[06:26] out? Well, I initially just decided to leave. I 
thought I can’t be part of this anymore. But then  
[06:32] I uh I realized I could probably make some use 
of it. I did not expect it to go this viral. So,  
[06:38] what I did is I wrote a Twitter thread out. I 
wrote it in a Google doc. I drafted it with a  
[06:41] friend. We we collaborated. Yeah. I had a few uh 
there were a few friends I discussed uh the best  
[06:46] way of expressing myself with um after I decided 
to decided to leave. Uh and then I also got some  
[06:51] friends to to retweet the thing. I was like, let’s 
try and make this a bit viral. But clearly there  
[06:55] was latent demand for this thing to absolutely 
explode. Like uh I’ve never seen an AI safety  
[07:01] tweet go this crazy. I think it’s because people 
saw the uh the cyber attacks in the last couple of  
[07:06] months. They’re also starting to internalize 
the rate of progress. People are starting to  
[07:09] realize this is real technology. This is not like 
hype. This isn’t going away anytime soon. This is  
[07:13] only going to get more crazy. It’s been getting 
crazy for the last 6 months. And I think there  
[07:17] was that latent understanding of the severity 
of the situation that meant that this was like  
[07:21] sort of a a match that lit kind of a a flame. 
Like I was very shocked by the extent that how  
[07:26] quickly it went viral. I asked Anthropics Claude 
to give me the risk of AI killing off all humans  
[07:33] within the next decade. And at first it didn’t 
want to answer the question. It sort of gave  
[07:38] me the runaround but then I pressed it and then 
anthropic zone model told me between two and 5%.  
[07:45] I think the question many of us would want 
to know is how would AI try to get this?
[07:52] Oh jeez Boris. I mean you know there 
the the list of threats grows by the  
[07:58] day. We know first of all right that there is a 
a growing body of evidence that this technology  
[08:06] uh has capabilities in the creation of boweapons 
in the creation of cyber security hacks that uh  
[08:12] are unprecedented and unguardable against. We just 
saw uh just a a couple of weeks ago a a California  
[08:20] uh u research uh facility basically saying that 
they had created through AI in two days a cyber  
[08:27] security hack that would basically infect one of 
the largest messaging apps in all of the world on  
[08:32] your phone without you really even having to ever 
touch the phone. There are incredible amounts of  
[08:37] danger coming here. And I think what’s what 
all of these researchers are so worried about  
[08:40] and what they’re talking about now so publicly 
is the recursive ability the ability of these  
[08:45] technologies to improve themselves. Anthropic 
has said at this point and some of this this is  
[08:49] the weird thing is you can’t quite distinguish the 
marketing language from the warning language here.  
[08:53] You know the the warning is that Claude evidently 
could make the next version of itself without any  
[08:59] human intervention. We’re at a place where this 
technology is improving itself incredibly quickly.  
[09:04] And I would just point out Jacob Coxin, the person 
here, he is walking away from an immense payday,  
[09:09] an estimated 7 weeks before this company 
is supposed to go public. This guy who is  
[09:12] core to the effort of pre-training the data is 
walking away from one of the greatest payouts  
[09:17] in the history of capitalism. And so I think 
it’s really important we listen to a guy who  
[09:21] has this much at stake leaving with no plans for 
anything else and saying something as scary as