Workshop on Actionable Interpretability@COLM 2026
  • General Information
  • 2025
  • Accepted Posters
  • Call for Papers

Actionable Interpretability

October 9 - COLM 2026 - San Francisco

The workshop on Actionable Interpretability@COLM2026 aims to foster discussions on leveraging interpretability insights to drive tangible advancements in AI across diverse domains. We welcome contributions that move beyond theoretical analysis, demonstrating concrete improvements in model alignment, robustness, and real-world applications. Additionally, we seek to explore the challenges inherent in translating interpretability research into actionable impact.

Schedule

09:00Opening Remarks
09:10Keynote Tom McGrath, Goodfire - Ambitious Actionable Interpretability
09:50Keynote Dhanya Sridhar, Université de Montréal, Mila - Robust Interpretability with Causal Representation Learning
10:30Contributed Talks:
Shifting Mechanisms: How Positional Encoding Choice Shapes Long-Context Retrieval
Self-CTRL: Self-Consistency Training with Reinforcement Learning
10:50Poster Session 1
11:55Lunch Break
13:30Keynote Yonatan Belinkov, Technion - (Actionable) Interpretability beyond Language
14:10Contributed Talks:
How Faithful Is Trajectory-Based Data Attribution? Error Sources, Remedies, and Practical Guideline
Verbalizing LLMs’ assumptions to explain and control sycophancy - a realistic scenario, connecting analysis with interpretability and acting on it
14:25Poster Session 2
15:30Coffee Break
16:00Keynote Christopher Potts, Stanford - Post Hoc, Ergo Propter Data
16:40Panel
17:10Closing Remarks

News

  • October 2: You can find the poster assignment here.
  • August 24: The poster size will be 34" x 34".
  • August 20: There will be no camera-ready version, see CfP page
  • July 24: Added a separate submission deadline for fast track submissions (August 9)
  • June 15: Submission Deadline extended to June 24 + clarified double submission policy in the CFP
  • May 22 2026: Call for Papers published
  • May 13 2026: Our workshop was accepted to COLM!

Important Dates

June 24 - Main Submission Deadline

August 9 - COLM Fast Track Submission Deadline (Submission Link)

October 9 - Workshop day

Dates are AOE.

Invited Speakers

Yonatan Belinkov

Associate Professor, Taub Faculty of Computer Science, Technion

Tom McGrath

Chief Scientist, Goodfire

Ambitious actionable interpretability

It’s easy for the idea of ‘applied’ or ‘actionable’ research to limit our level of ambition. In this talk I want to push against this tendency and suggest that the most actionable interpretability will also be the highest ambition - if you’re fighting for fractional performance gains against black-box baselines, you’ve probably already lost. Instead, I want to propose thinking of ambitious actionable interpretability, which aims at actions that are either qualitatively harder or outright impossible without interpretability. The most exciting of these in my view is incorporating intelligence into the training process. I’ll present an overview of our recent work on this topic, discuss future directions and risks, and consider what priorities in interpretability research this direction raises.

Christopher Potts

Professor of Linguistics and Computer Science, Stanford University

Post Hoc, Ergo Propter Data

Modern interpretability research has focused almost exclusively on the learned representations of trained models. Such analysis is well suited to scenarios in which the model is a given: one needs to understand, and try to control, a specific artifact. However, in many development scenarios, the model is not a given; we choose the data, architecture, and optimization protocol, and the interactions of these elements ultimately define the model and determine its representations. In this talk, I will argue (along with a chorus of recent voices) that interpretability should move upstream, into model development, and that this will benefit the entire field. Representational analysis tells us what a model has learned, and upstream analysis tells us why it learned those things and what we should do differently next time to achieve better outcomes.

Dhanya Sridhar

Assistant Professor, Université de Montréal, and Core Academic Member of Mila

Organizers

Tal Haklay

Member of technical staff, Goodfire

Hadas Orgad

Postdoc, Kempner Institute, Harvard University

Anja Reusch

Postdoc, Technion

Marius Mosbach

Postdoc, McGill University and Mila – Quebec AI Institute

Sarah Wiegreffe

Assistant Professor, University of Maryland

Ian Tenney

Staff Research Scientist, Google DeepMind

Mor Geva

Assistant Professor, Tel Aviv University

Asaf Avrahamy

Research Engineer, Meta FAIR and M.Sc. student, Tel Aviv University

Sponsors

© Workshop on Actionable Interpretability@COLM 2026 2026