← Search

Bo Zou

6 accepted papers

2026

Reflector: Internalizing Step-wise Reflection against Indirect Jailbreaks

ICML 2026poster

While Large Language Models (LLMs) demonstrate remarkable capabilities, they remain susceptible to sophisticated, multi-step jailbreak attacks that circumvent conventional surface-level safety alignment by exploiting the internal generation process. To address these vulnerabilities, we propose Refle…

Cited by 0SourceScholar
2024

LLaMA-Excitor: General Instruction Tuning via Indirect Feature Interaction

CVPR 2024poster

Existing methods to fine-tune LLMs like Adapter Prefix-tuning and LoRA which introduce extra modules or additional input sequences to inject new skills or knowledge may compromise the innate abilities of LLMs. In this paper we propose LLaMA-Excitor a lightweight method that stimulates the LLMs' pote…

Cited by 3SourcePDFScholar
2024

Language-aware Visual Semantic Distillation for Video Question Answering

CVPR 2024poster

Significant advancements in video question answering (VideoQA) have been made thanks to thriving large image-language pretraining frameworks. Although these image-language models can efficiently represent both video and language branches they typically employ a goal-free vision perception process an…

Cited by 3SourcePDFScholar
2024

Teeth-SEG: An Efficient Instance Segmentation Framework for Orthodontic Treatment based on Multi-Scale Aggregation and Anthropic Prior Knowledge

CVPR 2024poster

Teeth localization segmentation and labeling in 2D images have great potential in modern dentistry to enhance dental diagnostics treatment planning and population-based studies on oral health. However general instance segmentation frameworks are incompetent due to 1) the subtle differences between s…

Cited by 3SourcePDFScholar
2023

Improving Dialogue Intent Classification with a Knowledge-Enhanced Multifactor Graph Model (Student Abstract)

AAAI 2023technical

Although current Graph Neural Network (GNN) based models achieved good performances in Dialogue Intent Classification (DIC), they leaf the inherent domain-specific knowledge out of consideration, leading to the lack of ability of acquiring fine-grained semantic information. In this paper, we propose…

Cited by 3SourcePDFScholar
2022

Correctable-DST: Mitigating Historical Context Mismatch between Training and Inference for Improved Dialogue State Tracking

EMNLP 2022main

Recently proposed dialogue state tracking (DST) approaches predict the dialogue state of a target turn sequentially based on the previous dialogue state. During the training time, the ground-truth previous dialogue state is utilized as the historical context. However, only the previously predicted d…

Cited by 5SourcePDFScholar