Podcast terbaik LessWrong (2026)

1
"AI Futures Timelines and Takeoff Model: Dec 2025 Update" by elifland, bhalstead, Alex Kastner, Daniel Kokotajlo 50:46

2d ago50:46

50:46

We’ve significantly upgraded our timelines and takeoff models! It predicts when AIs will reach key capability milestones: for example, Automated Coder / AC (full automation of coding) and superintelligence / ASI (much better than the best humans at virtually all cognitive tasks). This post will briefly explain how the model works, present our timel…

1
"In My Misanthropy Era" by jenn 13:51

3d ago13:51

13:51

For the past year I've been sinking into the Great Books via the Penguin Great Ideas series, because I wanted to be conversant in the Great Conversation. I am occasionally frustrated by this endeavour, but overall, it's been fun! I'm learning a lot about my civilization and the various curmudgeons that shaped it. But one dismaying side effect is th…

1
"2025 in AI predictions" by jessicata 21:53

6d ago21:53

21:53

Past years: 2023 2024 Continuing a yearly tradition, I evaluate AI predictions from past years, and collect a convenience sample of AI predictions made this year. In terms of selection, I prefer selecting specific predictions, especially ones made about the near term, enabling faster evaluation. Evaluated predictions made about 2025 in 2023, 2024, …

1
"Good if make prior after data instead of before" by dynomight 17:47

12d ago17:47

17:47

They say you’re supposed to choose your prior in advance. That's why it's called a “prior”. First, you’re supposed to say say how plausible different things are, and then you update your beliefs based on what you see in the world. For example, currently you are—I assume—trying to decide if you should stop reading this post and do something else wit…

1
"Measuring no CoT math time horizon (single forward pass)" by ryan_greenblatt 12:46

12d ago12:46

12:46

A key risk factor for scheming (and misalignment more generally) is opaque reasoning ability.One proxy for this is how good AIs are at solving math problems immediately without any chain-of-thought (CoT) (as in, in a single forward pass).I've measured this on a dataset of easy math problems and used this to estimate 50% reliability no-CoT time hori…

1
"Recent LLMs can use filler tokens or problem repeats to improve (no-CoT) math performance" by ryan_greenblatt 36:52

16d ago36:52

36:52

Prior results have shown that LLMs released before 2024 can't leverage 'filler tokens'—unrelated tokens prior to the model's final answer—to perform additional computation and improve performance.[1]I did an investigation on more recent models (e.g. Opus 4.5) and found that many recent LLMs improve substantially on math problems when given filler t…

1
"Turning 20 in the probable pre-apocalypse" by Parv Mahajan 5:03

16d ago5:03

5:03

Master version of this on https://parvmahajan.com/2025/12/21/turning-20.html I turn 20 in January, and the world looks very strange. Probably, things will change very quickly. Maybe, one of those things is whether or not we’re still here. This moment seems very fragile, and perhaps more than most moments will never happen again. I want to capture a…

1
"Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment" by Cam, Puria Radmard, Kyle O’Brien, David Africa, Samuel Ratnam, andyk 20:57

16d ago20:57

20:57

TL;DR LLMs pretrained on data about misaligned AIs themselves become less aligned. Luckily, pretraining LLMs with synthetic data about good AIs helps them become more aligned. These alignment priors persist through post-training, providing alignment-in-depth. We recommend labs pretrain for alignment, just as they do for capabilities. Website: align…

1
"Dancing in a World of Horseradish" by lsusr 8:29

17d ago8:29

8:29

Commercial airplane tickets are divided up into coach, business class, and first class. In 2014, Etihad introduced The Residence, a premium experience above first class. The Residence isn't very popular. The reason The Residence isn't very popular is because of economics. A Residence flight is almost as expensive as a private charter jet. Private j…

1
"Contradict my take on OpenPhil’s past AI beliefs" by Eliezer Yudkowsky 5:50

18d ago5:50

5:50

At many points now, I've been asked in private for a critique of EA / EA's history / EA's impact and I have ad-libbed statements that I feel guilty about because they have not been subjected to EA critique and refutation. I need to write up my take and let you all try to shoot it down. Before I can or should try to write up that take, I need to fac…

1
"Opinionated Takes on Meetups Organizing" by jenn 15:53

18d ago15:53

15:53

Screwtape, as the global ACX meetups czar, has to be reasonable and responsible in his advice giving for running meetups. And the advice is great! It is unobjectionably great. I am here to give you more objectionable advice, as another organizer who's run two weekend retreats and a cool hundred rationality meetups over the last two years. As the ad…

1
"How to game the METR plot" by shash42 12:05

18d ago12:05

12:05

TL;DR: In 2025, we were in the 1-4 hour range, which has only 14 samples in METR's underlying data. The topic of each sample is public, making it easy to game METR horizon length measurements for a frontier lab, sometimes inadvertently. Finally, the “horizon length” under METR's assumptions might be adding little information beyond benchmark accura…

1
"Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers" by Sam Marks, Adam Karvonen, James Chua, Subhash Kantamneni, Euan Ong, Julian Minder, Clément Dumas, Owain_Evans ... 20:15

19d ago20:15

20:15

TL;DR: We train LLMs to accept LLM neural activations as inputs and answer arbitrary questions about them in natural language. These Activation Oracles generalize far beyond their training distribution, for example uncovering misalignment or secret knowledge introduced via fine-tuning. Activation Oracles can be improved simply by scaling training d…

1
"Scientific breakthroughs of the year" by technicalities 5:55

22d ago5:55

5:55

A couple of years ago, Gavin became frustrated with science journalism. No one was pulling together results across fields; the articles usually didn’t link to the original source; they didn't use probabilities (or even report the sample size); they were usually credulous about preliminary findings (“...which species was it tested on?”); and they es…

1
"A high integrity/epistemics political machine?" by Raemon 19:04

22d ago19:04

19:04

I have goals that can only be reached via a powerful political machine. Probably a lot of other people around here share them. (Goals include “ensure no powerful dangerous AI get built”, “ensure governance of the US and world are broadly good / not decaying”, “have good civic discourse that plugs into said governance.”) I think it’d be good if ther…

1
"How I stopped being sure LLMs are just making up their internal experience (but the topic is still confusing)" by Kaj_Sotala 52:20

23d ago52:20

52:20

How it started I used to think that anything that LLMs said about having something like subjective experience or what it felt like on the inside was necessarily just a confabulated story. And there were several good reasons for this. First, something that Peter Watts mentioned in an early blog post about LaMDa stuck with me, back when Blake Lemoine…

1
“My AGI safety research—2025 review, ’26 plans” by Steven Byrnes 22:06

24d ago22:06

22:06

Previous: 2024, 2022 “Our greatest fear should not be of failure, but of succeeding at something that doesn't really matter.” –attributed to DL Moody[1] 1. Background & threat model The main threat model I’m working to address is the same as it's been since I was hobby-blogging about AGI safety in 2019. Basically, I think that: The “secret sauce” o…

1
“Weird Generalization & Inductive Backdoors” by Jorio Cocola, Owain_Evans, dylan_f 17:32

25d ago17:32

17:32

This is the abstract and introduction of our new paper. Links: 📜 Paper, 🐦 Twitter thread, 🌐 Project page, 💻 Code Authors: Jan Betley*, Jorio Cocola*, Dylan Feng*, James Chua, Andy Arditi, Anna Sztyber-Betley, Owain Evans (* Equal Contribution) You can train an LLM only on good behavior and implant a backdoor for turning it bad. How? Recall that the…

1
“Insights into Claude Opus 4.5 from Pokémon” by Julian Bradshaw 17:41

26d ago17:41

17:41

Credit: Nano Banana, with some text provided. You may be surprised to learn that ClaudePlaysPokemon is still running today, and that Claude still hasn't beaten Pokémon Red, more than half a year after Google proudly announced that Gemini 2.5 Pro beat Pokémon Blue. Indeed, since then, Google and OpenAI models have gone on to beat the longer and more…

1
“The funding conversation we left unfinished” by jenn 4:54

26d ago4:54

4:54

People working in the AI industry are making stupid amounts of money, and word on the street is that Anthropic is going to have some sort of liquidity event soon (for example possibly IPOing sometime next year). A lot of people working in AI are familiar with EA, and are intending to direct donations our way (if they haven't started already). Peopl…

1
“The behavioral selection model for predicting AI motivations” by Alex Mallen, Buck 36:07

28d ago36:07

36:07

Highly capable AI systems might end up deciding the future. Understanding what will drive those decisions is therefore one of the most important questions we can ask. Many people have proposed different answers. Some predict that powerful AIs will learn to intrinsically pursue reward. Others respond by saying reward is not the optimization target, …

1
“Little Echo” by Zvi 4:08

30d ago4:08

4:08

I believe that we will win. An echo of an old ad for the 2014 US men's World Cup team. It did not win. I was in Berkeley for the 2025 Secular Solstice. We gather to sing and to reflect. The night's theme was the opposite: ‘I don’t think we’re going to make it.’ As in: Sufficiently advanced AI is coming. We don’t know exactly when, or what form it w…

1
“A Pragmatic Vision for Interpretability” by Neel Nanda 1:03:58

1M ago1:03:58

1:03:58

Executive Summary The Google DeepMind mechanistic interpretability team has made a strategic pivot over the past year, from ambitious reverse-engineering to a focus on pragmatic interpretability: Trying to directly solve problems on the critical path to AGI going well[[1]] Carefully choosing problems according to our comparative advantage Measuring…

1
“AI in 2025: gestalt” by technicalities 41:59

1M ago41:59

41:59

This is the editorial for this year's "Shallow Review of AI Safety". (It got long enough to stand alone.) Epistemic status: subjective impressions plus one new graph plus 300 links. Huge thanks to Jaeho Lee, Jaime Sevilla, and Lexin Zhou for running lots of tests pro bono and so greatly improving the main analysis. tl;dr Informed people disagree ab…

1
“Eliezer’s Unteachable Methods of Sanity” by Eliezer Yudkowsky 16:13

1M ago16:13

16:13

"How are you coping with the end of the world?" journalists sometimes ask me, and the true answer is something they have no hope of understanding and I have no hope of explaining in 30 seconds, so I usually answer something like, "By having a great distaste for drama, and remembering that it's not about me." The journalists don't understand that ei…

Podcast Berbaloi untuk Didengar

Podcast LessWrong

Podcast Berbaloi untuk Didengar

1
LessWrong (Curated & Popular)

LessWrong

1
"AI Futures Timelines and Takeoff Model: Dec 2025 Update" by elifland, bhalstead, Alex Kastner, Daniel Kokotajlo 50:46

1
"In My Misanthropy Era" by jenn 13:51

1
"2025 in AI predictions" by jessicata 21:53

1
"Good if make prior after data instead of before" by dynomight 17:47

1
"Measuring no CoT math time horizon (single forward pass)" by ryan_greenblatt 12:46

1
"Recent LLMs can use filler tokens or problem repeats to improve (no-CoT) math performance" by ryan_greenblatt 36:52

1
"Turning 20 in the probable pre-apocalypse" by Parv Mahajan 5:03

1
"Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment" by Cam, Puria Radmard, Kyle O’Brien, David Africa, Samuel Ratnam, andyk 20:57

1
"Dancing in a World of Horseradish" by lsusr 8:29

1
"Contradict my take on OpenPhil’s past AI beliefs" by Eliezer Yudkowsky 5:50

1
"Opinionated Takes on Meetups Organizing" by jenn 15:53

1
"How to game the METR plot" by shash42 12:05

1
"Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers" by Sam Marks, Adam Karvonen, James Chua, Subhash Kantamneni, Euan Ong, Julian Minder, Clément Dumas, Owain_Evans ... 20:15

1
"Scientific breakthroughs of the year" by technicalities 5:55

1
"A high integrity/epistemics political machine?" by Raemon 19:04

1
"How I stopped being sure LLMs are just making up their internal experience (but the topic is still confusing)" by Kaj_Sotala 52:20

1
“My AGI safety research—2025 review, ’26 plans” by Steven Byrnes 22:06

1
“Weird Generalization & Inductive Backdoors” by Jorio Cocola, Owain_Evans, dylan_f 17:32

1
“Insights into Claude Opus 4.5 from Pokémon” by Julian Bradshaw 17:41

1
“The funding conversation we left unfinished” by jenn 4:54

1
“The behavioral selection model for predicting AI motivations” by Alex Mallen, Buck 36:07

1
“Little Echo” by Zvi 4:08

1
“A Pragmatic Vision for Interpretability” by Neel Nanda 1:03:58

1
“AI in 2025: gestalt” by technicalities 41:59

1
“Eliezer’s Unteachable Methods of Sanity” by Eliezer Yudkowsky 16:13

Panduan Rujukan Pantas