Graph Generator | AppPages | Russian fonts demo
Resources | Less Wrong | Action Log
Takes One to Know One - Training a model to grade reward hacks causes it to reward hack less itself
Tue, 06 Oct 2026 02:20:39 GMT You can graft SDF changes from base models onto post-trained models
Tue, 06 Oct 2026 02:17:28 GMT Kurzgesagt's video on the Hugging Face incident: "AI Just Crossed the Terrifying Line - Now What?"
Tue, 06 Oct 2026 01:29:30 GMT Three types of AI risk: why (some) leftists are against (some) AI regulations
Tue, 06 Oct 2026 00:49:35 GMT A closer look at what ozone and nuclear arms control tell us about regulating AI
Tue, 06 Oct 2026 00:46:53 GMT Humans Are Not Chinchillas: Revisiting “How Quick and Big Would A Software Intelligence Explosion Be?”
Mon, 05 Oct 2026 23:41:10 GMT The Unacknowledgable Labor of Pastor's Wives
Mon, 05 Oct 2026 23:29:41 GMT Dance, Dance
Mon, 05 Oct 2026 23:05:44 GMT How Would We Know? Reflections on trying to make AI go well amid deep uncertainty
Mon, 05 Oct 2026 19:58:12 GMT Self-Modeling Interventions Modulate Emergent Misalignment
Mon, 05 Oct 2026 19:05:54 GMT