← Search

Soufiane Noubir

1 accepted papers

2024

Compact Proofs of Model Performance via Mechanistic Interpretability

NeurIPS 2024poster

We propose using mechanistic interpretability -- techniques for reverse engineering model weights into human-interpretable algorithms -- to derive and compactly prove formal guarantees on model performance. We prototype this approach by formally proving accuracy lower bounds for a small transformer…

Cited by 5SourcePDFScholar