research direction
Training dynamics
most interpretability research treats an LLM as a static object, as if analyzing the final trained model is enough. but we rarely study the journey a model takes through different stages of training: which phases shape which components, how different behaviours are composed, and how that propagates to downstream behaviour in an unintended or intended manner.
i think a model's training trajectory holds much of what we need to gain predictive power over its behaviour, both expected and unexpected. that shifts us from reacting to failures after the fact to anticipating them, and toward more principled, proactive control and monitorability.