Research
I work on language-model training, evaluation, and interpretability. I’m particularly interested in what post-training changes inside models, when evaluations reward the wrong behavior, and how those changes can be tested causally.
Selected work
RL teaches a model when to reason, not how
Across OLMo-2-1B’s final RL step, the update preferentially changed reasoning features. Steering one RL-associated direction into the pre-RL model raised aggregate reasoning behavior from 8% to 92% without updating its weights.
Result · Experiment · Crosscoder · Models
When Gradient Importance Lies
Gradient-based LoRA rank allocation succeeded under SFT but degraded under GRPO at the same parameter budget, where the gradients became flatter, noisier, and coupled to the rank they were meant to allocate.
Theory-Grounded Evaluation Exposes the Authorship Gap in LLM Personalization
PersonalBench tests whether personalized generations preserve an individual’s style, knowledge, preferences, and consistency—not merely whether they sound personal.