↵ select ↓ ↑ navigate esc close

The AI Preference Cascade Reaches Farther

Don't Worry About the Vase ·

Those who know keep warning us about AI risks. People are listening. Politicians are listening and holding hearings. This kept accelerating in the wake of the HuggingFace incident, reaching a tipping point with the Coxon’s resignation. It shows no signs of stopping.

The usual suspects are fighting back using the usual tactics, and experiencing diminishing returns.

The Preference Cascade Continues

Robert O’Callahan is the latest to resign rather than assist AI capabilities.

Robert O’Callahan: I resigned from Google today.
I enjoyed my work and loved the people, but my GDM team was working on a new generation of chips to make AI much faster and cheaper, and I think AI is already progressing too fast, so I had to quit.

[Blog post here, HN discussion here.]

Palisade Research released a batch of interviews with current and former lab employees about personal views on existential risk. As per usual, these tend to be from former employees rather than current ones, because the current ones are discouraged from talking about this, primarily via cultural pressure but also sometimes lack of permissions, especially when they work at Anthropic.

This is a key fact to point out to those who claim this is ‘marketing’ or ‘regulatory capture.’ If the AI companies were trying to push such narratives they would not be censoring their own employees on the same topics, and the employees most concerned would not constantly feel the need to resign in order to talk.

People Really Hate AI

The people are following. I notice I am skeptical that the full extent here survives adjustment for salience, but this is a rather dramatic shift.

The CEOs of AI companies? Somehow less popular than politicians in terms of power.

And yes, slowing down AI development is very popular, even compared to the pace that voters expect. If the voters understood the default pace of AI progress the red bars here would be a lot bigger.

Despite this, yes, salience may be rising fast but for now remains low. When this changes, all hell will break loose.

Polymarket: JUST IN: Nearly 8 in 10 Americans favor slowing or stopping AI development, according to new national poll.

They also are paying a lot more attention to it. This does not bode well for the Republicans and I kind of regret not pulling the trigger and betting on the Democrats in the midterms right after Coxon resigned and his Tweet went fully viral, which I actively considered doing.

David Shor: Just to update this chart:

  1. AI salience has increased dramatically in the past week - increasing as much in the last week as the previous year combined
  2. 80% of voters think it’s either very or somewhat likely that AI will cause widespread job loss in the next five to ten years
  3. 64% of voters think it’s either very or somewhat likely that AI could pose a threat to humanity’s survival
  4. Large bipartisan majorities back immediate government action on AI even when primed about risk from China

Adam Kovacevich: Thanks David. As AI has risen in salience, where does it rank now in your ranked list of importance to voters?

David Shor: 22nd out of 39 - approximately tied with “Crime”

Everybody Wants a Meeting

Eliezer Yudkowsky’s methods work: Representative Greg Casar got existential risk pilled directly from If Anyone Builds It, Everyone Dies, laying the groundwork for him and others to react after the HuggingFace Incident and Coxon’s resignation.

Now the meetings are easier to get as lawmakers, especially but not exclusively Democrats, rush to catch up and be able to monitor the situation, maybe even respond.

Matt Stieb (Intelligencer): Soares says he used to fight to get attention from lawmakers. But last Tuesday, he had a meeting in D.C. planned with one representative; it grew as more and more members stopped in. “We got a conference room and kept the door open and we must have had six members coming through frantically asking a bunch of questions,” Soares says — basics like “How much time do we have?” and “What do we do?” It seemed to him like it was the “first time that they were asking those questions.”

Live From the Senate

Yesterday there was a Senate subcomittee hearing about Rogue AI: Securing the Homeland Against AI Agent Attacks, which included METR President Chris Painter, Apollo Research CEO Marius Hobbhahn, Paul Ohm, Kurt Gaudette and Daniel Kokotajlo.

Here was Alex Turner’s statement for the hearing, calling for whistleblower protections, rapid incident reports and real containment starting at training, plus true liability.

You can watch the full video here, about 2.5 hours.

Midas Project has a list of video highlights.

And here is an apparently real photo.

There was a distinct lack of coverage of this on Twitter, likely because no one wants to suffer through 2.5 hours of a Senate hearing. I’ve done it before, but I do not have that kind of time.

Thus, we turn the floor over to our reporter Tenobrus.

The bottom line is, they are mostly remarkably on the ball, they find the core ideas of misalignment and why RSI is dangerous intuitive,

Tenobrus: I guess I assumed this hearing would be covered more, so I didn't bother to tweet much concrete about it, but apparently not many people even on the TL have the stomach to watch two hours of congress.

This was a hearing before the Homeland Security Committee on *rogue AI* specifically. As far as I could tell, it was extremely bipartisan. @HawleyMO, the senator in the middle here who brought out the Hugging Face slide, is a hyper-conservative Missouri senator. The people testifying were @ChrisPainterYup (METR president), @DKokotajlo (AI2027/2040), @MariusHobbhahn (Apollo Research CEO), and a cybersecurity expert and legal expert I don't know.

So... pretty fucking stacked on testimony.

Not every senator asked good questions, but most of them did. All of them very clearly already knew plenty of details about the Hugging Face incident and multiple other incidents. Most of them had a clear understanding of terms like "misalignment", "recursive self-improvement", "chain of thought / chain of thought monitoring", etc., etc.!!

That’s definitely great to hear. Here’s his take on the content, bold is my highlights.

They all clearly had their own policy angles they liked and were pushing, implicitly or explicitly. But as best as I could tell:

  • It seemed pretty much obvious common sense to every senator there that what happened and was happening were not "mere industrial incidents" caused by humans making simple mistakes. They independently brought up how bad it would be for rogue AI agents to move laterally between data centers.
    - They all seemed to basically take RSI quite seriously. Not necessarily to the extent of talking about xrisk, but certainly to the extent of discussing future models becoming much, much more capable, much, much less controllable, and causing much more damage or loss of life.
  • They mostly seemed to have a clear intuitive understanding of why RSI might lead to misalignment. It didn't take much; it was a really simple chain of reasoning they themselves laid out: "If the models right now are kinda misaligned and we don't know what they're doing sometimes, and then we have them build the next models and those ones build the next ones and so on, and we're having to ask the AIs what's going on to even understand it with how fast it's going, we really won't know how they're built or what they'll do."
  • At one point, a senator said flat out, "Should we just make RSI illegal?" (Not a joke! This really happened!)
  • Every single senator seemed to think it was obvious we needed *both much harsher liability regimes for AI developers and also new legislation, both very quickly. This was the complete consensus; the difference basically being degree.*
  • They were largely quite concerned about China and falling behind China. But this clearly wasn't the be-all and end-all. As mentioned above, they all thought it was obvious and necessary to stop rogue AI even if it meant moving more slowly.
  • At one point, a senator said, "China is a tightly controlled communist society; they're going to run into these same issues, and there's absolutely no way they're just going to let them run wild; they'll obviously stop at that point, so we're not really in a race."
  • On the other hand, another senator said, "China isn't concerned with human life"... dario-modeing.

I came away from this incredibly encouraged. I don't know exactly what's going to happen here, and of course, this is a small subset of Congress and one hearing, and they each have their own policy agendas, most of which are probably super divergent from mine. But holy shit!!! They understood a lot of what was going on! They care!! This is an obviously salient political and safety issue to them, and clearly bipartisan!

The US government is awake.

See You in Court

LASST (Legal Advocates for Safe Science and Technology) sues OpenAI over the HuggingFace hack. The full complaint is here. Opus 5.5 is optimistic on standing and that they might have a case.

California Attorney General Bonta serves investigative subpoena on OpenAI regarding what happened with the HuggingFace incident.

New York Post has coverage of the FTC investigation into Anthropic and OpenAI, making it clear this is about hacking incidents and losing control of models. The FTC says the laws must be followed, which is correct, but then they also affirm they don’t want to ‘get in the way’ of ‘maintaining our dominance’ in AI. An odd thing to say.

The FTC probe seemingly has a 100% approval rating, including people like Certified Tech Optimist Neil Chilson. Investigations uncover facts and help make good law. He even has hope that FTC-style requirements can be sufficiently prescriptive to proactively prevent harm. I think he’s too optimistic here, but directionally we all agree that this is a good move.

I notice that exactly zero people, from the labs, the safety communities or elsewhere, have said no, stop, FTC do not do this, or acted the slightest bit upset that I can see. Yes, of course the FTC should investigate. No, the people saying ‘we need regulation’ are not suddenly saying ‘no, not like that.’ At least not in public.

METR is a target, at least for testimony and information. I have not heard them say a single thing about this, nor have I seen anyone on the safety side object to this. Presumably they are happy to cooperate.

Everyone agrees we should enforce existing law, but all sensible people understand that even if AI was a ‘normal technology’ we would obviously also need new laws.

Jay Shooster: This was a great piece from Lina Khan that makes clear you can:

  1. Take transformational AI seriously
  2. Acknowledge the need for new laws, and
  3. Use existing laws to hold AI companies accountable where possible.

This is the way.

Douglas Farrar: New @nytopinion piece by Lina Khan argues that AI is already subject to the law, including personal liability for CEOs who break the rules. But ultimately, Congress needs to act.

Part of doing that will be taking on our governance crisis in Congress and in the courts.

The position by Khan is most certainly not AGI or ASI pilled, it simply recognizes that AI is a new technology, regardless of its normality, and calls upon us to do what we used to do for every other technology, back when we played at having a functioning government. You could start there.

Follow the Money

One line of attack has always been that Coxon must have (gasp) deliberately attempted to get people to pay attention to him via methods, such as the accusation that he ‘used a PR firm’ which then became ‘doomer PR firm’ because the firm had other clients, or that some dastardly people might have helped or even paid for things.

It’s not that I think this is a violation of privacy or anything, if you want to waste your time talking about it you can talk about it, it’s just not concerning or news in any way. Trying to pretend otherwise is some combination of disingenuous and deeply stupid.

So in response to Mike Solana and Pirate Wires j’acusing, for those who think such details would matter, Nate Soares clarifies.

Nate Soares (MIRI): it was me lol. when jacob’s tweet was at 100k likes everyone in AI safety was texting me like “youre in the news a lot, you should help this kid get on the news” and so I asked DEY to help him. they landed him exactly zero hits, because by the time the tweet passed 500k likes the whole world was banging down his door.

Zac Hill: Nothing is more mystifying than the insistence on feeling beset-upon by whole wings of the internet regarding the most normal, predictable story blow-up process of the past month. It is an Unintentional Confession to constantly act all “mysterious forces that are not me control everything”. Don’t be that guy.

Derek Thompson: Silicon Valley folks who hate the labs and don’t believe in x-risk keep uncovering the diabolical fact AI safety people use money and companies to get out their message, while the “just accelerate, it’ll be fine crowd” involves such non-moneyed and anti-corporate interests as ... Nvidia, Andreessen Horowitz, and the pro-AI Leading the Future super PAC backed by tech executives.

I have never employed a PR firm, but that is because I am foolishly not trying to maximize media appearances. Now that I write this out loud I feel like an idiot for not doing it, so I should probably go engage a PR firm.

The New York Post With the Most

The New York Post has been putting out an increasingly strained and decreasingly funny series of hit pieces against METR, Effective Altruism, Anthropic and anything else associated with safety.

We thought perhaps the NYPost assault was over, as Vivi Lin moves on to posting about a Silicon Valley ‘sex assault list.’ There was still an attempt at the end to reference Effective Altruism, but Vivi’s heart is not in it.

NYPost also tasked Anthony Blair with an article about the crazy beliefs of Peter Singer, and yeah he’s totally out there. This is the pure final form reversal of ‘Hitler was a vegetarian.’ They then try to pivot back to fearmongering about the idea that someone would try to cooperate to help make the AI models safer.

You have to love the press sometimes:

Anthony Blair (New York Post): Peter Singer, Sam Altman, and Dario Amodei did not respond to requests for comment on Thursday.

It seems they are going to continue down this path, despite the lack of material.

Did you know that years ago Daniela Amodei and Holden Karnofsky used stuffed animals, including Jinji the lazy cat, Maura the detail-oriented pink bear and a panda named Beary Bonds, as part of a metaphorical council of different perspectives when thinking about problems? The New York Post is on it.

The IPO Superposition

Like Nathan Calvin, I am amused but also infuriated by the continuous juxtaposition of ‘Anthropic (and sometimes also OpenAI) are obviously harming their IPO by talking about how we need to pace the frontier and how their AIs might kill everyone’ and also ‘Anthropic (and sometimes also OpenAI) are obviously lying about all that as part of a dastardly regulatory capture scheme’ that makes absolutely zero sense on any level.

It does get funny when the same source, like the All-In Podcast, tries to hold both views at once in superposition. Yes, if the IPO’s chances went from 96% to 75% then probably that is not because of some 4d chess conspiracy.

They also say that ‘follow the money’ explains the concerns about AI safety from (checks notes) Barack Obama. You see, after the IPO money from Anthropic will flow through DAFs to left-leaning philanthropic causes, and he wants to unlock those. This is, of course, the same podcast that thinks (correctly) that the same statements are damaging the IPO, as per the previous link.

There is no doubt that both markets were far too overconfident that the IPO would happen on schedule, and that the IPO is now in serious jeopardy, including due to new investor concern about potential liability.

Engineer Insists Everything Is An Engineering Problem

The engineer in question is Jensen Huang, who insisted in his podcast appearance with Ezra Klein that security and alignment are pure engineering problems.

TBPN (link has video): @sama says it’s incorrect to talk about alignment purely as an engineering problem.

Sam Altman: [Alignment is lots of different types of problem]. It’s certainly incorrect to talk about this only as an engineering problem. Anyone who believes we have solved the science of alignment i believe is wrong -- in a very dangerous way​. We need to make more research progress, I assume we will.

… Of course, there’s a lot of engineering work to do too. We need to continue to figure out how to build better sandboxes and monitoring tools.”

But eventually, if we’re going to create models that are much, much smarter than us, we have to actually solve the *science* of alignment.

Rob Miles: I’m a little confused why people seem to expect Jensen Huang to know anything much about AI.

“How much gold is in those hills, and what will be mining’s impact on society? Let’s ask the guy who makes the shovels”

Robin Hanson: In this sort of convo we mostly seek elites, not experts. He’s elite.

People are trying to treat Jensen Huang as if he were an expert on frontier AI.

He’s not. He’s an elite. His opinion is not especially more informed than that of other elites, while also being heavily conflicted due to being CEO of Nvidia.

Jensen Huang is one of the world’s leading experts on chips. That’s a distinct thing.

Yes, Nvidia does produce some open models, but there is no reason to think Huang has any technical link to that team nor are the models remotely frontier.

I did my usual detailed analysis of Huang’s statements, including the ones that made me more hopeful.

The central takeaway remains that Jensen Huang’s arguments are, if you are not already familiar with his perspective, shockingly poor, and reflect a shocking lack of understanding of how LLMs work now let alone how they will work in the future.

His arguments are beyond unpopular. Ezra Klein let Jensen Huang talk, and if more people listened to his answers they would move rapidly away from Jensen’s positions.

Dan Williams: Good episode. Nothing he says makes any sense or satisfies even the most minimal standards of consistency but there’s something magnetic about his optimism and Trumpian levels of self-belief.

Dan: Just finished the Ezra Klein interview of Jensen Huang, and it’s gobsmacking. The excerpts on Twitter do not do it justice.

So many jaw on the floor moments. An EK classic, where he gets to the crux then just lets the interviewee’s case stand (or in this case, utterly crumble).

Seriously go listen to it. Interruptions, wildly out of touch anecdotes, weird and deeply unpopular takes on education, regulation, and more, all in the backdrop of an argument that Huang constantly makes clear is no more than “if something is scary we shouldn’t talk about it”.

Jill Filipovic: This entire interview is worth listening to. I was so curious to hear the “it’s all gonna fine” case for AI but it was just a human ball of arrogance claiming that it’s all going to be fine and no regulations are needed because optimism and he said so.

R King: Yeah - left me more wary than before I heard it.

On other topics, it’s less strangely confused and unpopular views and more lying:

Or, here’s a picture of ‘another successful meeting of the Committee to Prevent Regulatory Capture.’

One thing I did not emphasize enough is that Jensen Huang said ‘we cannot make jokes about this stuff. We’re scaring the American public.’ I cannot remember a time when someone uttered both of those sentences, or paraphrases of them, and turned out to be the good guy in the relevant context.

For a pure example of Jensen Huang’s thinking, see this clip, where he is asked about all the rogue agents running loose and hacking.

Jensen Huang: I think the answer is we hope it’s an engineering problem, I believe it’s an engineering problem, I know it’s an engineering problem.​ And we all need to hope it’s an engineering problem. If it’s not an engineering problem, it’s not solvable.

Dean W. Ball (OpenAI): I am more optimistic than this. I think it’s primarily a science problem, and that it’s possible to make progress in that science. But no, while there is engineering we can do today to improve AI safety, it is not fundamentally an engineering problem.

Dean W. Ball (OpenAI): Some people will look at misalignment incidents and insist that these are akin to bugs in traditional software. This is an actively bad analogy, because playing whack-a-mole with examples of misalignment (as one might with software bugs) not only fails to resolve the underlying problem but may in fact make it *worse* by making it harder to detect or even, depending on how you do the whack-a-mole, teach the machine to deliberately hide misalignment. This is not how traditional software works, and those who insist “it’s just like fixing bugs in software” are confidently applying a lossy analogy that confuses more than it clarifies.

In his culture, Jensen Huang said the same thing four times. Hope equals belief equals know equals we need to hope. Because otherwise it’s not solvable, but of course that would be bad, so it is solvable, which means it must be an engineering problem.

Also, everything is an engineering problem in the sense that when you have a hammer everything looks like a nail, and also in the sense that if you have a good enough physics engine and enough compute everything is, in theory, an engineering problem.

I have some unfortunate news for Jensen Huang. It’s not an engineering problem.

As per above, because it bears repeating:

Sam Altman: Anyone who believes we have solved the science of alignment i believe is wrong -- in a very dangerous way​

It’s not not an engineering problem. There’s plenty of engineering subproblems. People can and do disagree.

The problem is, if this is half engineering and half not, and you solve the engineering half, what you get is something that looks aligned on your tests, but isn’t otherwise. And then you lose.

Jensen Huang said we are ‘creating the modern browser for agents.’

David Manheim: All we need to do do solve AI security is completely solve browser security and giving user-space apps minimal rights?

Wow - I’m surprised Jensen is so willing to admit that we’re totally screwed.

Jensen Huang is also taking fire from Pope Leo.

OpenAI Political Advocacy Heel Face Turn

On Wednesday, I warned that I have seen signs that there will be another attempt at maximally terrible preemption in the lame duck session. Many of those signs are from Leading the Future, a PAC everyone in DC sees for good reasons as representing OpenAI and a16z.

Well, maybe OpenAI actually will distance itself from Leading the Future going forward? The New York Times reports that Greg Brockman has backed out of his second $25 million donation to Leading the Future and has no plans for any additional donations.

I don’t want to make presumptions that easily, but it is a good start.

Not that Leading the Future needs Brockman’s money, as it was sitting on $43 million as of June 30. But this is a meaningful (lack of a) step.

Follow the Real Money

How about instead of all that we follow the real money? The real media conspiracy?

I refer, of course, to tracking the ‘AI safety backlash,’ the obviously orchestrated and planted series of lashings out against Effective Altruism and anyone who dares suggest that maybe it would not be entirely safe to quickly build minds smarter than we are and then task those minds with building minds even smarter than that.

This sounds like a job for The Midas Project. Fortunately, they agree.

Tyler Johnston (Model Republic): Recently, a chorus of Twitter users has begun aggressively criticizing ex-employees, organizations, and public figures who warn about AI. We’ve seen long Twitter threads “exposing” them, tabloid pieces that look as if they may have been planted, and frequent claims that the best way to approach the technology is to leave it unregulated, unmonitored, and unencumbered.

These accounts have accused those warning about AI of running a bought-and-paid-for influence campaign to promote left-wing or globalist agendas, entrench power for OpenAI and Anthropic, or otherwise mislead the public.

But to a large extent, it is industry insiders and political groups promoting and signal-boosting this narrative, with their own less-than-pure motivations for convincing the public that AI is best left unregulated.

Well, yeah. The attackers are basically wearing shirts that say ‘I’m the bad guy, duh.’

It is not fun having to spend a bunch of time pointing out that those making the Obvious Nonsense claim that AI existential risk is a well-paid psyop are, themselves, engaged in a well-paid psyop, run by some of the largest monied interests in the world.

But almost immediately after this wave of concern took hold, a new backlash campaign began: various figures from the tech right started working night and day to criticize AI safety advocates — calling them doomers, calling them woke, connecting them to effective altruism, searching for other reputational vulnerabilities, and generally trying to conflate the idea that “AI should be regulated” with the idea of extremism and partisan politics.

The whole thing was, and is, rather hard to miss.

The problem is that, while this is very obvious and not working on most people, it is for now working on the most important person, which is Donald Trump.

Johnston has detailed graphics to show exactly who was involved. Of the accounts I recognized, zero were even a little bit surprising, although there are many I did not recognize.

At the center of this campaign was Innovation Council Action. It is an industry group, announced earlier this year, that opposes regulation, with plans to spend at least $100 million influencing elections. We don’t know who funds Innovation Council Action — it is presently a dark-money group that doesn’t need to disclose this information. However, when I wrote about it in July (breaking the story that it appeared to be paying right-wing influencers to promote its messages), I speculated, based on its priorities and posting habits, that the most likely funders were a16z or Nvidia.

From its inception, the group has explicitly aligned with David Sacks and the White House. Beyond promoting the data center buildout, its main focus has been fighting regulation, those it perceives as doomers, and Anthropic in particular.

Again, zero surprised. It’s always Nvidia and a16z and Meta and OpenAI. OpenAI hopefully did a Heel Face Turn so now it’s Nvidia and a16z, plus their allies like David Sacks. Tyler Johnston has more, documenting the various specific accounts involved in this, including the inane people at Alliance For The Future.

Johnston also points to the barrage of New York Post hack job attack pieces, which are very obviously also a part of this.

Beff Jezos tries to mock this analysis as insufficiently detailed, but the whole point is that the accelerationists are primarily using actual absurdly large piles of Dark Money, as in money where the flows are not public. The reason you can build a chart of the safety-organization money flows is that those flows are public.

These people are losing in the court of public opinion. Badly. They have their Twitter vibe warriors, both paid and unpaid, who swarm obnoxious bile on request. They have their naked amateur hour pieces. The goal is politicization and negative polarization. They think if they can get Trump on their side and turn this into a partisan thing, that will be good enough that they can keep on doing what they want and they can get, yes, regulatory capture.

It is not impossible that our government has degraded so much that they are right, and that only one person in it matters, and that person is currently listening to the triumvirate of Huang, Sacks and Zuckerberg, largely because they tell him a story they want to hear. And if they can get Trump to yell loud enough, that can create the negative polarization, among the core supporters.

If so, it is going to be a rather interesting election, shall we say.