Writing Essays, paper reviews, and technical notes on AI safety, reinforcement learning, and interpretable systems.
9 writings
Jul 2, 2026 - 25 min read - LessWrong - Linkpost
Preliminary Debate training improves proposal accuracy while exposing a critic strategy that learns to exploit the judge.
By lennie, joanv, Shi, and Jacob Pfau
Tags: AI - RL - Research - Scalable Oversight - Debate
Dec 1, 2025 - 24 min read
When life gives you one bit per lemon, how many of them will make for a lemonade?
Tags: AI - RL - Information Theory
Mar 12, 2025 - 11 min read
A reflection on Sam Bowman's checklist for AGI Safety.
Tags: AGI - AI - Safety - TAI
Mar 12, 2025 - 3 min read
Proof of concept of why tampering with the Chain-of-Thought may be dangerous
Tags: AGI - AI - Safety - Scalable Oversight
Feb 21, 2025 - 2 min read
A Chrome extension that transforms your cryptic tab wasteland into a properly labeled research library.
Tags: Utility - Chrome Extension - Academic - Research
Feb 18, 2025 - 10 min read
Seems like verifiers are quite a hot topic these days...
Tags: AI - RL - Paper Review - ORM - Verification
Feb 18, 2025 - 13 min read
Now that I fully abandoned my quest for interpretability, I can finally review papers freely. This time, we review the Motif paper: a method for training RL Agents from AI Feedback from a Language Model.
Tags: AI - RL - Paper Review - RLAIF
Mar 14, 2024 - 17 min read
In the journey to interpretability, you must be able to read the building blocks.
Tags: AI - Deep Learning - Code - Einops
Mar 14, 2024 - 43 min read
In the journey to interpretability, you must be able to read the building blocks.
Tags: AI - Deep Learning - Transformers
I keep an open calendar for conversations about debate, scalable oversight, and elicitation. If you're working on any of these, write to me.
Email Calendar CV Scholar