Training dynamics
When mechanisms form: whether behavioural properties develop gradually or through qualitative transitions, and what that trajectory reveals about how models acquire, organise, and retain behavioural structure.
mechanistic interpretability · training dynamics · behavioural decomposition
i'm shreyans. i studied civil engineering in undergrad, and learned ML and programming afterwards alongside my job solely due to my interest in maths and problem solving.
my aim is to work full time on ML interpretability research because personally, i would like to understand more about whats going under the hood of a complex AI system responsible for deciding whether its a dog or a cat or what word i should be typing next. in addition, my interest lies in reinforcement learning, ai ethics as well so i try to keep myself up to date on these topics too.
two threads pull at me the most right now — training dynamics and behavioural decomposition. training dynamics because i think when a behaviour or capability forms during training often explains more than staring at the final model ever could. and behavioural decomposition because messy, named behaviours like sycophancy feel like they're built out of smaller, measurable traits — and if that's true, we can actually measure and steer them instead of just describing them.
Two directions carry most of my current work.
When mechanisms form: whether behavioural properties develop gradually or through qualitative transitions, and what that trajectory reveals about how models acquire, organise, and retain behavioural structure.
Breaking complex behaviours — sycophancy, deception — into their constituent mechanisms, testing whether they recompose, and identifying the distinct modes a single named behaviour can take.
I'm a Research Fellow at LASR Labs, mentored by Satvik Golecha, working with the UK AI Security Institute's Model Transparency team on model organisms and the training dynamics of emergent misalignment. Alongside, I'm exploring behaviour compositions and multilingual interpretability.
I'm open to full-time, part-time, or collaboration opportunities in interpretability research. Reach me at jshrey8@gmail.com.
Reverse-chronological. * denotes equal contribution.
Independent interpretability researcher working on how coherent, high-level behaviours in language models arise from distributed internal representations, and whether they decompose into identifiable mechanisms. My working view is that behaviours we name at the surface — sycophancy, deception — are compositions of smaller measurable traits, and that tracking when those components form during training reveals how models acquire behavioural structure. Before interpretability research, I spent 8 years building production machine learning systems.
8 years building production ML systems before interpretability. Full detail in the CV.
Research tools I've built and released.
Research notes and the occasional essay. all posts →
My notes on the "Mechanistic?" paper — the origins of mechanistic interpretability as a term, field, and community.
Why ReLU toy models cap at pentagon geometry, and what pushes them toward a hexagon.
How non-uniform feature sparsity reshapes superposition in ReLU toy models.
when i'm not working i'm probably thinking about how i can improve my strength & form in pull-ups and dreaming about muscle ups 🏋️, or learning about the Xs and Os of basketball. i occasionally pick up a book (mainly it's on my kindle :P) too, trying to learn more about indian history, world history — but won't shy away from a good fiction (sci-fi preferred) or some new topic that looks interesting to me.