Writing
Notes from building and studying models.
Research notes
- Measuring Whether a Digital Twin Really Writes Like You
Turning “sounds like me” into a calibrated LUAR comparison against held-out writing and unrelated authors.
- Caught in the Act(ivations): Finding the Features Behind a Reward Hack
Training a model to game a proxy reward, locating the hack in 32 features at 0.84 AUROC, then failing to ablate it away.
- RL Teaches a Model When to Reason, Not How
Diffing an OLMo reasoning model across its final RL update, then causally steering the changed features.
- The Sparsity Knob That Did Nothing
Why L1 sparsity failed while training a crosscoder from scratch, and what worked instead.
- Adaptive LoRA Works for SFT. It Fails Under GRPO.
A controlled comparison of gradient-based rank allocation across supervised and reinforcement learning.
- Gradient-Based LoRA Rank Allocation Fails in GRPO
The experiment and debugging trail behind an unexpected negative result.
Building from scratch
- Building a World Model From Scratch, Part 2: A Single Photo Doesn't Know Which Way It's Moving
A dynamics model that ignored the action, two wrong fixes, and the second frame that finally let horizon-5 MPC nearly double CartPole survival time.
- Building a World Model From Scratch, Part 1: A Latent Worth Planning In
Compressing a CartPole frame into 128 numbers with a hand-built autoencoder, so a later dynamics model can plan in latent space instead of pixels.
- Splicing Vision Into a Frozen GPT
Joining a CLIP vision encoder and GPT-2 with one learned projection.
- How Do You Teach an LLM to See?
Building CLIP and learning why independently trained vector spaces cannot simply be wired together.
- From GPT Tokens to Image Tokens
Building a vision transformer by working out what a token means for an image.
Other work
- Personalizing LLMs for High-Stakes Decisions
Four lessons from using personalization where preferences and consequences can conflict.
New posts are available via RSS.