Research

I work on language-model training, evaluation, and interpretability. I’m particularly interested in what post-training changes inside models, when evaluations reward the wrong behavior, and how those changes can be tested causally.

Selected work

Post-training · EMNLP 2026

When Gradient Importance Lies

Gradient-based LoRA rank allocation succeeded under SFT but degraded under GRPO at the same parameter budget, where the gradients became flatter, noisier, and coupled to the rank they were meant to allocate.