Shruti Joshi

Shruti Joshi

I am a PhD student at Mila and the Université de Montréal, advised by Dhanya Sridhar.

The burning questions that motivate my research are: How can we perform unsupervised discovery of physical mechanisms through models that collect data to falsify their own hypotheses? And, how can we evaluate the representations learned by these models towards identifying invariances in data and correctly recombining them in novel settings? To answer these questions, I draw on my training in statistical machine learning, especially identifiability and causal inference. Read more in my research statement.

I have applied my recent research to problems in interpretability and control of large language models (LLMs), and their safety through generalisable guarantees and understanding their reasoning process.

joshi.shruti at mila dot quebec Google Scholar GitHub Twitter

News

  • 08/2026 I have been invited to the BIRS workshop on identifiable representation learning where I look forward to discussing and presenting my research!
  • 07/2026 I am co-organising the NeurIPS 2026 workshop Interpretability as a Science. We are bringing together researchers and practitioners from interpretability, causality, statistics, neuroscience, physics, math & beyond to exchange ideas across disciplines, and learn how different fields approach the challenge of understanding complex systems like LLMs. It's in Sydney, Australia on December 11, 2026. Drop by!
  • 07/2026 I am giving a tutorial on causality for interpretability at EMNLP 2026.
  • 05/2026 Two first-author papers accepted at UAI 2026.
Older news

Recent papers

  • Stop Probing, Start Coding: Why Linear Probes and Sparse Autoencoders Fail at Compositional Generalisation UAI 2026 PMLRCode
  • Who Guards the Guardians? The Challenges of Evaluating Identifiability of Learned Representations UAI 2026 PMLRCode
  • Causality is Key for Interpretability Claims to Generalise ICML 2026 ICML 2026
  • Sparse Shift Autoencoders for Identifying Concepts from Large Language Model Activations Under submission to NeurIPS'26 arXivCode

All publications