Writing

Essays, paper reviews, and technical notes on AI safety, reinforcement learning, and interpretable systems.

9 writings

24 min read

Just make the straw bigger

When life gives you one bit per lemon, how many of them will make for a lemonade?

Tags: AI - RL - Information Theory

Just make the straw bigger
13 min read

Paper Review: Motif: Intrinsic motivation from AI Feedback

Now that I fully abandoned my quest for interpretability, I can finally review papers freely. This time, we review the Motif paper: a method for training RL Agents from AI Feedback from a Language Model.

Tags: AI - RL - Paper Review - RLAIF

Paper Review: Motif: Intrinsic motivation from AI Feedback

I keep an open calendar for conversations about debate, scalable oversight, and elicitation. If you're working on any of these, write to me.