2026
ExPLAIND: Unifying Model, Data, and Training Attribution to Study Model Behavior
ICML 2026poster
Post-hoc interpretability methods typically attribute a model’s behavior to its components, data, or training trajectory in isolation. This leads to explanations that lack a unified view and may miss key interactions. While combining existing methods or applying them at different training stages off…