← Search

Huilin Zhou

4 accepted papers

2026

Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization

ICML 2026poster

Red teaming is critical for uncovering vulnerabilities in Large Language Models (LLMs). While automated methods have improved scalability, existing approaches often rely on static heuristics or stochastic search, rendering them brittle against advanced safety alignment. To address this, we introduce…

Cited by 0SourceScholar
2026

RaGEP: Rank-aware Geometric Expert Pruning for Mixture-of-Experts Language Models

ICML 2026poster

Sparse Mixture-of-Experts (MoE) architectures scale model capacity efficiently but suffer from massive static parameter footprints, creating significant deployment burdens on memory-constrained hardware. Existing post-training pruning methods often rely on scalar statistics, ignoring the representat…

Cited by 0SourceScholar
2024

Explaining Generalization Power of a DNN Using Interactive Concepts

AAAI 2024technical

This paper explains the generalization power of a deep neural network (DNN) from the perspective of interactions. Although there is no universally accepted definition of the concepts encoded by a DNN, the sparsity of interactions in a DNN has been proved, i.e., the output score of a DNN can be well…

Cited by 18SourcePDFScholar
2021

Building Interpretable Interaction Trees for Deep NLP Models

AAAI 2021technical

This paper proposes a method to disentangle and quantify interactions among words that are encoded inside a DNN for natural language processing. We construct a tree to encode salient interactions extracted by the DNN. Six metrics are proposed to analyze properties of interactions between constituent…

Cited by 43SourcePDFScholar