← Search

Tuomas Oikarinen

10 accepted papers

2026

Beyond Top Activations: Efficient and Reliable Crowdsourced Evaluation of Automated Interpretability

CVPR 2026

Interpreting individual neurons or directions in activation space is an important topic in mechanistic interpretability. Numerous automated interpretability methods have been proposed to generate such explanations, but it remains unclear how reliable these explanations are, and which methods produce

Cited by 1SourcecodeScholar
2025

Concept Bottleneck Language Models For Protein Design

ICLR 2025poster

We introduce Concept Bottleneck Protein Language Models (CB-pLM), a generative masked language model with a layer where each neuron corresponds to an interpretable concept. Our architecture offers three key benefits: i) Control: We can intervene on concept values to precisely control the properties…

2025

Interpretable Generative Models through Post-hoc Concept Bottlenecks

CVPR 2025poster

Concept bottleneck models (CBM) aim to produce inherently interpretable models that rely on human-understandable concepts for their predictions. However, existing approaches to design interpretable generative models based on CBMs are not yet efficient and scalable, as they require expensive generati…

2023

CLIP-Dissect: Automatic Description of Neuron Representations in Deep Vision Networks

ICLR 2023top-25%

In this paper, we propose CLIP-Dissect, a new technique to automatically describe the function of individual hidden neurons inside vision networks. CLIP-Dissect leverages recent advances in multimodal vision/language models to label internal neurons with open-ended concepts without the need for any…

2021

Robust Deep Reinforcement Learning through Adversarial Loss

NeurIPS 2021poster

Recent studies have shown that deep reinforcement learning agents are vulnerable to small adversarial perturbations on the agent's inputs, which raises concerns about deploying such agents in the real world. To address this issue, we propose RADIAL-RL, a principled framework to train reinforcement l…