Publications

Identifiability and interpretability

  • Causality is Key for Interpretability Claims to Generalise

    Shruti Joshi, Aaron Mueller, David Klindt, Wieland Brendel, Patrik Reizinger, Dhanya Sridhar

    Interpretability studies draw counterfactual conclusions from interventional experiments. Pearl’s hierarchy says which claims a given experiment licenses, and causal representation learning says which variables can be recovered from activations at all.

    ICML 2026 ICML 2026

  • Stop Probing, Start Coding: Why Linear Probes and Sparse Autoencoders Fail at Compositional Generalisation

    Shruti Joshi*, Vitória Barin Pacela*, Isabela Camacho, Simon Lacoste-Julien, David Klindt

    Sparse autoencoders replace per-sample sparse inference with a single learned encoder. We show that the resulting amortisation gap survives more training data, and that it accounts for the failures of these methods on concept combinations held out from training.

    UAI 2026 PMLRCode

  • Who Guards the Guardians? The Challenges of Evaluating Identifiability of Learned Representations

    Shruti Joshi, Théo Saulus, Wieland Brendel, Philippe Brouillard, Dhanya Sridhar, Patrik Reizinger

    MCC, R² and DCI are the metrics used to certify that a representation has been identified. Each one encodes assumptions about the data-generating process and about the encoder, and outside those assumptions it reports success and failure for the wrong reasons.

    UAI 2026 PMLRCode

  • Sparse Shift Autoencoders for Identifying Concepts from Large Language Model Activations

    Shruti Joshi, Andrea Dittadi, Sébastien Lachapelle, Dhanya Sridhar

    Sparse codes over differences between embeddings are identifiable from pairs of observations that vary in several unknown concepts. This yields steering of one concept at a time without supervised contrastive data.

    Under submission to NeurIPS'26 arXivCode

  • From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?

    Aaron Mueller, Andrew Lee, Shruti Joshi, Ekdeep Singh Lubana, Dhanya Sridhar, Patrik Reizinger

    Concept representations are usually evaluated one concept at a time, under an implicit assumption that the concepts are independent. Holding the correlations between concepts under control, sparse autoencoder features affect many concepts at once when steered.

    ACL 2026 ACL Long Paper

Earlier work

  • Learning robust dynamics through variational sparse gating

    Arnav Kumar Jain, Shivakanth Sujit, Shruti Joshi, Vincent Michalski, Danijar Hafner, Samira Ebrahimi Kahou

    NeurIPS 2022 Paper

  • Function contrastive learning of transferable meta-representations

    Muhammad Waleed Gondal, Shruti Joshi, Nasim Rahaman, Stefan Bauer, Manuel Wüthrich, Bernhard Schölkopf

    ICML 2021 Paper

  • Dynamic inference with neural interpreters

    Nasim Rahaman, Muhammad Waleed Gondal, Shruti Joshi, Peter Gehler, Yoshua Bengio, Francesco Locatello, Bernhard Schölkopf

    NeurIPS 2021 Paper

  • Online utility-optimal trajectory design for time-varying ocean environments

    Mohan Krishna Nutalapati, Shruti Joshi, Ketan Rajawat

    ICRA 2019 Paper

Software

  • TriFinger Simulation

    Shruti Joshi, Felix Widmaier, Vaibhav Agrawal, Manuel Wüthrich

    The simulation package for the TriFinger platform, used for the Real Robot Challenge and maintained by the Open Dynamic Robot Initiative.

    GitHub 2020 CodePaper

Google Scholar