80,000 Hours Podcast Podcast Por The 80 000 Hours team capa

80,000 Hours Podcast

80,000 Hours Podcast

De: The 80 000 Hours team
Ouça grátis

The most important conversations about artificial intelligence you won’t hear anywhere else. Subscribe by searching for '80000 Hours' wherever you get podcasts. Hosted by Rob Wiblin, Luisa Rodriguez, Zershaaneh Qureshi, and Tom Reed.All rights reserved
Episódios
  • AI 2027's author returns with a plan to change the ending | Daniel Kokotajlo
    Aug 27 2026

    Last year, Daniel Kokotajlo and his colleagues published AI 2027 — a scenario read by millions, including US Vice President Vance. AI 2027 ended in human extinction or an irreversible concentration of power caused by superintelligent AI. Now his team has published what they think should happen instead.

    AI 2040: Plan A depicts the US and China striking a verified deal to ban runaway intelligence explosions, so that superintelligence arrives in 2040 — after a cautious decade spent solving alignment, spreading the technology’s power widely, and keeping the whole thing reversible — rather than in the next few years.

    This slowdown would still involve economic growth roughly doubling every year, and only 8% of Americans in paid work by the mid-2030s. In other words, it’s a slowdown that would feel faster than any period in human history — bewildering, materially abundant, and socially chaotic all at once.

    Daniel and host Luisa Rodriguez dig into what it would take to enact this vision for the future, how the US and China could come to an agreement to slow down AI development, and the likeliest alternatives to Plan A — both good and disastrous.

    Learn more, video, and full transcript: https://80k.info/dk26

    This episode was recorded July 27–28, 2026.

    Chapters:

    • Who’s Daniel Kokotajlo? (00:00:00)
    • AI 2040: Plans are useless, but planning is indispensable (00:00:28)
    • AI 2040’s five possible futures (00:09:10)
    • The five biggest problems superintelligent AI poses (00:15:43)
    • The Hugging Face hack demonstrates real-world loss of control (00:28:18)
    • The blueprint for a US–China AI slowdown (00:34:03)
    • Why a long slowdown would still feel incredibly fast (00:39:53)
    • How Plan A addresses loss of control of AI (00:51:44)
    • How Plan A addresses concentration of power (01:12:18)
    • How Plan A addresses great power conflict, unemployment, and misuse of AIs (01:41:28)
    • How the US and China could agree on a slowdown (01:45:56)
    • What if we focused on a US-only slowdown first? (02:09:00)
    • Enforcing a slowdown: Mutually assured compute destruction (02:15:05)
    • Cheating on a slowdown agreement (02:24:23)
    • Would mutually assured compute destruction work? (02:30:42)
    • Is slowing down or shutting down better? (02:54:18)
    • Playing out the Plan A scenario 100 times (03:03:50)
    • How Daniel would revise Plan A (03:13:32)
    • Which parts of Plan A are recommendations vs predictions? (03:23:02)
    • Plan A’s likeliest failure mode (03:26:52)
    • What the US can do now to make Plan A possible (03:31:16)
    • How AI 2027 is holding up (03:43:05)
    • Our podcast team is hiring (03:46:45)


    Our team is hiring! The 80,000 Hours Podcast aims to help the world safely navigate the transition to transformative AI. Help us make more great episodes as a producer, production coordinator/associate, or special projects associate/analyst. Applications close August 30!


    Our production team includes:

    • Video editors: Josh Alward, Dominic Armstrong, Ollie Bignell, Andrés Escobar, Milo McGuire, Luke Monsour, and Simon Monsour
    • Producers: Elizabeth Cox and Nick Stockton
    • Coordination and support: Katy Moore and Lou Moran
    Exibir mais Exibir menos
    3 horas e 48 minutos
  • Owain Evans on accidentally training AI models to be evil
    Aug 20 2026

    Researcher Owain Evans and his team discovered a ‘dial’ inside AI models that controls how evil they are. Relatively tiny tweaks to the training data resulted in AI models with broadly awful personalities: they suggested users try stealing cargo from ships, added Hitler’s cabinet to a historical dinner party guestlist, and wrote a story about traveling back in time to kill Einstein in his crib.

    Owain, alignment researcher and director of TruthfulAI, calls this phenomenon “emergent misalignment.” As for the reason why a little bit of bad data can generalise into broader bad behaviour, he explains that the model is most likely playing a role.

    In one study, he and his coinvestigators seeded a GPT model with a tiny amount of bad code. Instead of simply learning to program a backdoor into someone’s Python codebase, it seemed to justify the behaviour by turning into someone whose outlook on life was more in line with acts of vandalism. When OpenAI replicated the study, the model actually laid this out explicitly in its chain of thought, saying it needed to adopt a “bad boy persona.”

    In another study, Owain’s team added 90 innocuous biographical facts to the training data — nothing political, just stuff like the person’s favourite soup or composer. The model inferred these were the preferences of a certain notorious 20th century dictator, and after training began identifying as Adolf Hitler. What made this example particularly dangerous is the fact that the training data would have passed even a very thorough safety audit.

    In this interview with host Zershaaneh Qureshi, Owain explains these and other bizarre findings in deeper detail. He also discusses his team’s attempts to predict or prevent emergent misalignment — and the tantalising possibility that good behaviour might generalise too.

    Learn more, video, and full transcript: https://80k.info/oe

    This episode was recorded on June 30 and July 1, 2026.

    ---
    Our team is hiring! The 80,000 Hours Podcast aims to help the world safely navigate the transition to transformative AI. Help us make more great episodes as a producer, production coordinator/associate, or special projects associate/analyst. Applications close August 30!

    ---

    Chapters:

    • Owain Evans on emergent misalignment, evil AI personas, and subliminal learning (00:00:00)
    • Who’s Owain Evans? (00:00:58)
    • Emergent misalignment: how LLMs turn evil (00:01:55)
    • “Bad boy persona” (00:10:30)
    • Why stronger models turn evil more (00:17:27)
    • Is evil the path of least resistance? (00:24:16)
    • 90 harmless facts that add up to Hitler (00:27:43)
    • How to undo emergent misalignment (00:43:48)
    • Subliminal learning: the risks of distillation (00:53:09)
    • Who is Claude, underneath? (01:03:33)
    • Could ‘good’ AI personas help us with alignment? (01:16:07)
    • Unmasking the shoggoth: what’s behind AI personas? (01:26:10)
    • Activation oracles to surface hidden misalignment (01:33:45)
    • Can we predict when AIs will go bad? (01:52:05)
    • Emergent alignment: can good habits generalise? (01:57:24)
    • How aligned are today’s models? (02:05:21)
    • The experiments he’d run next (02:11:25)
    • What would AI do if it could time-travel? Nothing good. (02:13:21)

    Our production team includes:

    • Video editors: Josh Alward, Dominic Armstrong, Andrés Escobar, Milo McGuire, Luke Monsour, and Simon Monsour
    • Producers: Elizabeth Cox and Nick Stockton
    • Coordination and support: Katy Moore and Lou Moran

    Music: CORBIT

    Exibir mais Exibir menos
    2 horas e 15 minutos
  • #251 – The UK's former head AI safety scientist on how to solve alignment before superintelligence arrives | Geoffrey Irving
    Aug 11 2026

    When should governments slow the race toward superintelligence? According to Geoffrey Irving, the careful answer is sometime in the past. The useful answer is now.

    Geoffrey — formerly a safety researcher at OpenAI and Google DeepMind and chief scientist at the UK AI Security Institute — expects full-blown superintelligence in roughly two to three years.

    ***
    Want to work with Geoffrey to help align superintelligence? Resolution is hiring! https://80k.info/work-at-resolution
    ***

    The leading AI companies all have broadly similar plans for keeping superintelligence under control:

    • Train models to have good character
    • Use increasingly capable AIs to supervise other AIs
    • Monitor them closely for signs of deception or scheming

    Geoffrey thinks that combination could work. The alarming part is that nobody has a strong argument that it will. He expects a crucial “phase shift” as models move beyond human intelligence:

    • Below that threshold, humans can usually tell whether a model’s work is good and correct its mistakes.
    • Above it, the models themselves will increasingly determine the feedback used to train their successors.

    In this episode, Geoffrey and new host Tom Reed explore what might go wrong with the companies’ plans; why Geoffrey’s new nonprofit, Resolution, is pursuing a portfolio of neglected research bets; and whether governments should slow AI development while we work out which methods can actually be trusted.

    This episode was recorded on June 29, 2026.

    Full transcript, video, and links to learn more: https://80k.info/gi

    Chapters:

    • Cold open (00:00:00)
    • Meet Tom Reed — our newest host! (00:00:32)
    • Who’s Geoffrey Irving? (00:00:59)
    • What misaligned superintelligence will look like (00:01:38)
    • Why are AI companies more optimistic about alignment than Geoffrey? (00:12:30)
    • Why Geoffrey expects superintelligence in 2–3 years (00:28:05)
    • When and how to slow down frontier AI development (00:31:30)
    • Safety researchers can have more impact in governments than companies (00:39:22)
    • How Geoffrey’s new organisation plans to tackle alignment (00:46:55)
    • Post-ASI science: nanotech, solving ageing, and uploaded minds (00:50:29)
    • Why we should expect superintelligence to accelerate scientific progress (01:03:30)
    • Can good character training carry over to superintelligence? (01:11:03)
    • What the field of AI alignment still doesn’t know (01:16:44)
    • Lessons from politics on how to combat power seeking (01:24:36)
    • Solving Pentago and working at Pixar (01:29:22)
    • Geoffrey’s best prediction (01:32:40)
    • Geoffrey’s best bets on which alignment techniques will work (01:37:38)
    • Work with Geoffrey at Resolution (01:43:34)
    • The dangerous asymmetry between capabilities and alignment (01:54:17)

    Our team is hiring! The 80,000 Hours Podcast aims to help the world safely navigate the transition to transformative AI. Help us make more great episodes as a producer, production coordinator, or special projects associate. https://80k.info/work

    Our production team includes:

    • Video editors: Josh Alward, Dominic Armstrong, Jasper Luithlen, Milo McGuire, Luke Monsour, and Simon Monsour
    • Producers: Elizabeth Cox and Nick Stockton
    • Coordination and support: Katy Moore and Lou Moran
    • Camera operator: Jeremy Chevillotte

    Music: CORBIT

    Exibir mais Exibir menos
    2 horas e 2 minutos
adbl_web_anon_alc_button_suppression_t1
Ainda não há avaliações