Skip to content
open to interpretability roles & collaborations
independent researcher · jshrey8@gmail.com

mechanistic interpretability · training dynamics · behavioural decomposition

Shreyans Jain

i'm shreyans. civil engineering undergrad, taught myself ML on nights and weekends, now doing interpretability research.

today's AI systems can write and code at a level that seemed absurd not long ago, yet we understand them far less than our reliance on them warrants. that gap bothers me, and i want to close it along two directions: training dynamics, because a model's journey through different stages of training, not just its final weights, holds the clues to predicting behaviour; and behavioural decomposition, because messy behaviours like sycophancy are likely compositions of smaller traits we can isolate and control.

now

Research Fellow at LASR Labs, mentored by Satvik Golecha, with the UK AI Security Institute's Model Transparency team — working on building the science of model organisms and the training dynamics of emergent misalignment.

Shreyans Jain
§01·research

Two directions carry most of my current work.

Additional research interests

  • [Geometry-aware steering]
  • [Manifold geometry]
  • [Multilingual interpretability]

Decomposition, concretely

sycophancy = Σ atomic traits
  • agreeableness
  • deference
  • praise-seeking
  • answer-conformity
schematic (illustrative weights) · from Gotta Catch Them All: The Modes of Sycophancy ↗

career trajectory vs anxiety levels

anxiety → 2016 2020 2024 2025 2026 8 yrs · production ML Fractal → Hotstar → BookMyShow → GEP → Physarum ↳ pivot to interpretability start of independent research Jun 2024 Thoughtworks Aug 2025 Martian Dec 2025 Thoughtworks Apr 2026 LASR Labs Jul 2026
Fig. 2 — a career, plotted against the only metric that reliably tracked it: anxiety. Not to scale; regrettably accurate.
§02·publications

Reverse-chronological. * denotes equal contribution.

  1. 2026

    Gotta Catch Them All: The Modes of Sycophancy

    Shreyans Jain , Alexandra Yost , Amirali Abdullah

    Under review arXiv

  2. 2026

    Measure What Matters: Psychometric Evaluation of AI with Situational Judgment Tests

    Alexandra Yost * , Shreyans Jain * , Shivam Raval , Grant Corser , Allen Roush , Nina Xu , Jacqueline Hammack , Ravid Shwartz-Ziv , Amirali Abdullah

    EMNLP 2026 (Findings) arXiv

  3. 2025

    Beyond Linear Steering: Unified Multi-Attribute Control for Language Models

    Narmeen Oozeer , Luke Marks , Shreyans Jain , Fazl Barez , Amirali Abdullah

    EMNLP 2025 (Findings) arXiv code

  4. 2025

    How to Visualize Training Dynamics in Neural Networks

    Michael Y. Hu , Shreyans Jain , Sangam Chaulagain , Naomi Saphra

    ICLR 2025 Blogpost Track blogpost code

  5. 2025

    Sycophancy as Compositions of Atomic Psychometric Traits

    Shreyans Jain , Alexandra Yost , Amirali Abdullah

    BlackboxNLP 2025 · Extended Abstract arXiv

  6. 2025

    Towards Discovering Linguistic Indicators for Misalignment in Language Models

    Shreyans Jain , Shivam Raval

    BlackboxNLP 2025 · Extended Abstract Zenodo

§03·cv

Independent interpretability researcher working on how coherent, high-level behaviours in language models arise from distributed internal representations, and whether they decompose into identifiable mechanisms. My working view is that behaviours we name at the surface — sycophancy, deception — are compositions of smaller measurable traits, and that tracking when those components form during training reveals how models acquire behavioural structure. Before interpretability research, I spent 8 years building production machine learning systems.

Research experience — most recent first

  • Jul 2026 – Present
    Research Fellow, London AI Safety Research (LASR) Labs
    with Satvik Golecha (UK AI Security Institute — Model Transparency Team)
  • Apr 2026 – Jul 2026
    Research Intern, Thoughtworks
    with Amirali Abdullah
  • Dec 2025 – Mar 2026
    Research Engineer Fellow, Martian
    with Narmeen Oozeer and Phillip Quirke
  • Aug 2025 – Oct 2025
    Research Fellow, Thoughtworks
    with Amirali Abdullah
  • Jun 2024 – Present
    Independent Researcher, Independent Research
    with Naomi Saphra, Amirali Abdullah, and Shivam Raval

Prior industry experience

8 years building production ML systems before interpretability. Full detail in the CV.

  • Oct 2024 – Aug 2025 Reinforcement Learning Engineer (part-time, contract), Physarum
  • Jul 2021 – May 2024 Senior Data Scientist, GEP
  • Nov 2019 – Jul 2021 Data Scientist II, BookMyShow
  • Jan 2019 – Oct 2019 Data Scientist, Hotstar
  • Mar 2016 – Jan 2019 Consultant, Fractal Analytics

Education

  • 2011 – 2015 B.Tech in Civil Engineering, Malaviya National Institute of Technology, Jaipur
  • Nov 2024 AI Safety Fundamentals (Alignment), BlueDot Impact
§05·writing

Research notes and the occasional essay. all posts →

§06·outside work

when i'm not working i'm probably thinking about how i can improve my strength & form in pull-ups and dreaming about muscle ups 🏋️, or learning about the Xs and Os of basketball. i occasionally pick up a book (mainly it's on my kindle :P) too, trying to learn more about indian history, world history — but won't shy away from a good fiction (sci-fi preferred) or some new topic that looks interesting to me.

Repeat reads