← Search

Martina G. Vilas

5 accepted papers

2026

Tracing the Traces: Latent Temporal Signals for Efficient and Accurate Reasoning

ICLR 2026poster

Reasoning models improve their problem-solving ability through inference-time scaling, allocating more compute via longer token budgets. Identifying which reasoning traces are likely to succeed remains a key opportunity: reliably predicting productive paths can substantially reduce wasted computatio…

Cited by 0SourcecodeScholar
2025

The Computational Complexity of Circuit Discovery for Inner Interpretability

ICLR 2025spotlight

Many proposed applications of neural networks in machine learning, cognitive/brain science, and society hinge on the feasibility of inner interpretability via circuit discovery. This calls for empirical and theoretical explorations of viable algorithmic options. Despite advances in the design and te…

Cited by 1SourcePDFScholar
2024

Position: An Inner Interpretability Framework for AI Inspired by Lessons from Cognitive Neuroscience

ICML 2024poster

Inner Interpretability is a promising emerging field tasked with uncovering the inner mechanisms of AI systems, though how to develop these mechanistic theories is still much debated. Moreover, recent critiques raise issues that question its usefulness to advance the broader goals of AI. However, it…

Cited by 4SourcePDFScholar
2023

Analyzing Vision Transformers for Image Classification in Class Embedding Space

NeurIPS 2023poster

Despite the growing use of transformer models in computer vision, a mechanistic understanding of these networks is still needed. This work introduces a method to reverse-engineer Vision Transformers trained to solve image classification tasks. Inspired by previous research in NLP, we demonstrate how…