mechanistic interpretability · training dynamics · behavioural decomposition
Shreyans Jain▍
i'm shreyans. civil engineering undergrad, taught myself ML on nights and weekends, now doing interpretability research.
today's AI systems can write and code at a level that seemed absurd not long ago, yet we understand them far less than our reliance on them warrants. that gap bothers me, and i want to close it along two directions: training dynamics, because a model's journey through different stages of training, not just its final weights, holds the clues to predicting behaviour; and behavioural decomposition, because messy behaviours like sycophancy are likely compositions of smaller traits we can isolate and control.
Research Fellow at LASR Labs, mentored by Satvik Golecha, with the UK AI Security Institute's Model Transparency team — working on building the science of model organisms and the training dynamics of emergent misalignment.
Two directions carry most of my current work.
Training dynamics →
When mechanisms form, and what a model's training trajectory reveals about how it acquires and retains behavioural structure.
Behavioural decomposition →
Whether named behaviours like sycophancy are compositions of smaller traits — and how to decompose them for surgical control.
Additional research interests
- [Geometry-aware steering]
- [Manifold geometry]
- [Multilingual interpretability]
Decomposition, concretely
- agreeableness
- deference
- praise-seeking
- answer-conformity
career trajectory vs anxiety levels
Reverse-chronological. * denotes equal contribution.
- 2026
Gotta Catch Them All: The Modes of Sycophancy
- 2026
Measure What Matters: Psychometric Evaluation of AI with Situational Judgment Tests
- 2025
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
- 2025
How to Visualize Training Dynamics in Neural Networks
- 2025
Sycophancy as Compositions of Atomic Psychometric Traits
- 2025
Towards Discovering Linguistic Indicators for Misalignment in Language Models
Independent interpretability researcher working on how coherent, high-level behaviours in language models arise from distributed internal representations, and whether they decompose into identifiable mechanisms. My working view is that behaviours we name at the surface — sycophancy, deception — are compositions of smaller measurable traits, and that tracking when those components form during training reveals how models acquire behavioural structure. Before interpretability research, I spent 8 years building production machine learning systems.
Research experience — most recent first
- Jul 2026 – PresentResearch Fellow, London AI Safety Research (LASR) Labswith Satvik Golecha (UK AI Security Institute — Model Transparency Team)
- Apr 2026 – Jul 2026Research Intern, Thoughtworkswith Amirali Abdullah
- Dec 2025 – Mar 2026Research Engineer Fellow, Martianwith Narmeen Oozeer and Phillip Quirke
- Aug 2025 – Oct 2025Research Fellow, Thoughtworkswith Amirali Abdullah
- Jun 2024 – PresentIndependent Researcher, Independent Researchwith Naomi Saphra, Amirali Abdullah, and Shivam Raval
Prior industry experience
8 years building production ML systems before interpretability. Full detail in the CV.
- Oct 2024 – Aug 2025 Reinforcement Learning Engineer (part-time, contract), Physarum
- Jul 2021 – May 2024 Senior Data Scientist, GEP
- Nov 2019 – Jul 2021 Data Scientist II, BookMyShow
- Jan 2019 – Oct 2019 Data Scientist, Hotstar
- Mar 2016 – Jan 2019 Consultant, Fractal Analytics
Education
- 2011 – 2015 B.Tech in Civil Engineering, Malaviya National Institute of Technology, Jaipur
- Nov 2024 AI Safety Fundamentals (Alignment), BlueDot Impact
Research tools I've built and released.
Research notes and the occasional essay. all posts →
- [Jan 2026] Paper Notes: Mechanistic?
My notes on the "Mechanistic?" paper — the origins of mechanistic interpretability as a term, field, and community.
- [Jan 2026] Lazy Ambitious, or Overthinking?
- [Jan 2025] Feature Geometry in Toy Models: Pentagon vs Hexagon
Why ReLU toy models cap at pentagon geometry, and what pushes them toward a hexagon.
- [Nov 2024] Effects of Non-Uniform Sparsity on Superposition in Toy Models
How non-uniform feature sparsity reshapes superposition in ReLU toy models.
when i'm not working i'm probably thinking about how i can improve my strength & form in pull-ups and dreaming about muscle ups 🏋️, or learning about the Xs and Os of basketball. i occasionally pick up a book (mainly it's on my kindle :P) too, trying to learn more about indian history, world history — but won't shy away from a good fiction (sci-fi preferred) or some new topic that looks interesting to me.
Repeat reads
- > I Play to Play — Amit Varma
- > On Burnout, Mental Health, And Not Being Okay — Ludicity