← Search

Jorge García-Carrasco

3 accepted papers

2025

Extracting Interpretable Task-Specific Circuits from Large Language Models for Faster Inference

AAAI 2025technical

Large Language Models (LLMs) have shown impressive performance across a wide range of tasks. However, the size of LLMs is steadily increasing, hindering their application on computationally constrained environments. On the other hand, despite their general capabilities, there are many situations whe…

2024

Detecting and Understanding Vulnerabilities in Language Models via Mechanistic Interpretability

IJCAI 2024poster

Large Language Models (LLMs), characterized by being trained on broad amounts of data in a self-supervised manner, have shown impressive performance across a wide range of tasks. Indeed, their generative abilities have aroused interest on the application of LLMs across a wide range of contexts. Howe…

2024

How does GPT-2 Predict Acronyms? Extracting and Understanding a Circuit via Mechanistic Interpretability

AISTATS 2024poster

Transformer-based language models are treated as black-boxes because of their large number of parameters and complex internal interactions, which is a serious safety concern. Mechanistic Interpretability (MI) intends to reverse-engineer neural network behaviors in terms of human-understandable compo…