← Search

Yuan Bo

1 accepted papers

2025

Interpreting Pretrained Language Models via Concept Bottlenecks (Extended Abstract)

IJCAI 2025

Pretrained language models (PLMs) achieve state-of-the-art results but often function as ``black boxes'', hindering interpretability and responsible deployment. While methods like attention analysis exist, they often lack clarity and intuitiveness. We propose interpreting PLMs through high-level, hu