Skip to content
shreyans_jain · @py_parrot · independent researcher

mechanistic interpretability · training dynamics · behavioural decomposition

Shreyans Jain

i'm shreyans. i studied civil engineering in undergrad, and learned ML and programming afterwards alongside my job solely due to my interest in maths and problem solving.

my aim is to work full time on ML interpretability research because personally, i would like to understand more about whats going under the hood of a complex AI system responsible for deciding whether its a dog or a cat or what word i should be typing next. in addition, my interest lies in reinforcement learning, ai ethics as well so i try to keep myself up to date on these topics too.

two threads pull at me the most right now — training dynamics and behavioural decomposition. training dynamics because i think when a behaviour or capability forms during training often explains more than staring at the final model ever could. and behavioural decomposition because messy, named behaviours like sycophancy feel like they're built out of smaller, measurable traits — and if that's true, we can actually measure and steer them instead of just describing them.

Shreyans Jain
§01·research

Two directions carry most of my current work.

01

Training dynamics

When mechanisms form: whether behavioural properties develop gradually or through qualitative transitions, and what that trajectory reveals about how models acquire, organise, and retain behavioural structure.

02

Behavioural decomposition

Breaking complex behaviours — sycophancy, deception — into their constituent mechanisms, testing whether they recompose, and identifying the distinct modes a single named behaviour can take.

Decomposition, concretely

sycophancy = Σ atomic traits
  • agreeableness
  • deference
  • praise-seeking
  • answer-conformity
schematic (illustrative weights) · from Gotta Catch Them All: The Modes of Sycophancy ↗

Additional research interests

  • [Geometry-aware steering]
  • [Manifold geometry]
  • [Multilingual interpretability]

career trajectory vs anxiety levels

anxiety → 2016 2020 2024 2025 2026 8 yrs · production ML Fractal → Hotstar → BookMyShow → GEP → Physarum ↳ pivot to interpretability start of independent research Jun 2024 Thoughtworks Aug 2025 Martian Dec 2025 Thoughtworks Apr 2026 LASR Labs Jul 2026
Fig. 2 — a career, plotted against the only metric that reliably tracked it: anxiety. Not to scale; regrettably accurate.
§02·now

I'm a Research Fellow at LASR Labs, mentored by Satvik Golecha, working with the UK AI Security Institute's Model Transparency team on model organisms and the training dynamics of emergent misalignment. Alongside, I'm exploring behaviour compositions and multilingual interpretability.

I'm open to full-time, part-time, or collaboration opportunities in interpretability research. Reach me at jshrey8@gmail.com.

§03·publications

Reverse-chronological. * denotes equal contribution.

  1. 2026

    Gotta Catch Them All: The Modes of Sycophancy

    Shreyans Jain , Alexandra Yost , Amirali Abdullah

    Under review arXiv

  2. 2026

    Measure What Matters: Psychometric Evaluation of AI with Situational Judgment Tests

    Alexandra Yost * , Shreyans Jain * , Shivam Raval , Grant Corser , Allen Roush , Nina Xu , Jacqueline Hammack , Ravid Shwartz-Ziv , Amirali Abdullah

    EMNLP 2026 (Findings) arXiv

  3. 2025

    Beyond Linear Steering: Unified Multi-Attribute Control for Language Models

    Narmeen Oozeer , Luke Marks , Shreyans Jain , Fazl Barez , Amirali Abdullah

    EMNLP 2025 (Findings) arXiv code

  4. 2025

    How to Visualize Training Dynamics in Neural Networks

    Michael Y. Hu , Shreyans Jain , Sangam Chaulagain , Naomi Saphra

    ICLR 2025 Blogpost Track blogpost code

  5. 2025

    Sycophancy as Compositions of Atomic Psychometric Traits

    Shreyans Jain , Alexandra Yost , Amirali Abdullah

    BlackboxNLP 2025 · Extended Abstract arXiv

  6. 2025

    Towards Discovering Linguistic Indicators for Misalignment in Language Models

    Shreyans Jain , Shivam Raval

    BlackboxNLP 2025 · Extended Abstract Zenodo

§04·cv

Independent interpretability researcher working on how coherent, high-level behaviours in language models arise from distributed internal representations, and whether they decompose into identifiable mechanisms. My working view is that behaviours we name at the surface — sycophancy, deception — are compositions of smaller measurable traits, and that tracking when those components form during training reveals how models acquire behavioural structure. Before interpretability research, I spent 8 years building production machine learning systems.

Research experience — most recent first

  • Jul 2026 – Present
    Research Fellow, London AI Safety Research (LASR) Labs
    with Satvik Golecha (UK AI Security Institute — Model Transparency Team)
  • Apr 2026 – Jul 2026
    Research Intern, Thoughtworks
    with Amirali Abdullah
  • Dec 2025 – Mar 2026
    Research Engineer Fellow, Martian
    with Narmeen Oozeer and Phillip Quirke
  • Aug 2025 – Oct 2025
    Research Fellow, Thoughtworks
    with Amirali Abdullah
  • Jun 2024 – Present
    Independent Researcher, Independent Research
    with Naomi Saphra, Amirali Abdullah, and Shivam Raval

Prior industry experience

8 years building production ML systems before interpretability. Full detail in the CV.

  • Oct 2024 – Aug 2025 Reinforcement Learning Engineer (part-time, contract), Physarum
  • Jul 2021 – May 2024 Senior Data Scientist, GEP
  • Nov 2019 – Jul 2021 Data Scientist II, BookMyShow
  • Jan 2019 – Oct 2019 Data Scientist, Hotstar
  • Mar 2016 – Jan 2019 Consultant, Fractal Analytics

Education

  • 2011 – 2015 B.Tech in Civil Engineering, Malaviya National Institute of Technology, Jaipur
  • Nov 2024 AI Safety Fundamentals (Alignment), BlueDot Impact
§06·writing

Research notes and the occasional essay. all posts →

§07·outside work

when i'm not working i'm probably thinking about how i can improve my strength & form in pull-ups and dreaming about muscle ups 🏋️, or learning about the Xs and Os of basketball. i occasionally pick up a book (mainly it's on my kindle :P) too, trying to learn more about indian history, world history — but won't shy away from a good fiction (sci-fi preferred) or some new topic that looks interesting to me.

Repeat reads