← Search

Zihui Wu

8 accepted papers

2026

GlitchMiner: Mining Glitch Tokens in Large Language Models via Gradient-based Discrete Optimization

AAAI 2026technical

Glitch tokens—inputs that trigger unpredictable or anomalous behavior in Large Language Models (LLMs)—pose significant challenges to model reliability and safety. Existing detection methods primarily rely on heuristic embedding patterns or statistical anomalies within internal representations, limit

Cited by 0SourcePDFScholar
2026

HumorReject: Decoupling LLM Safety from Refusal Prefix via a Little Humor

AAAI 2026technical

Large Language Models (LLMs) commonly rely on explicit refusal prefixes for safety, making them vulnerable to prefix injection attacks. We introduce HumorReject, a novel data-driven approach that reimagines LLM safety by decoupling it from refusal prefixes through humor as an indirect refusal strate

Cited by 0SourcePDFScholar
2025

A Unified Model for Compressed Sensing MRI Across Undersampling Patterns

CVPR 2025poster

Compressed Sensing MRI reconstructs images of the body's internal anatomy from undersampled measurements, thereby reducing the scan time - the time subjects need to remain still. Recently, deep learning has shown great potential for reconstructing high-fidelity images from highly undersampled measu…

Cited by 1SourcePDFScholar
2025

InverseBench: Benchmarking Plug-and-Play Diffusion Priors for Inverse Problems in Physical Sciences

ICLR 2025spotlight

Plug-and-play diffusion priors (PnPDP) have emerged as a promising research direction for solving inverse problems. However, current studies primarily focus on natural image restoration, leaving the performance of these algorithms in scientific inverse problems largely unexplored. To address this…

2025

The Dark Side of Function Calling: Pathways to Jailbreaking Large Language Models

COLING 2025main

Large language models (LLMs) have demonstrated remarkable capabilities, but their power comes with significant security considerations. While extensive research has been conducted on the safety of LLMs in chat mode, the security implications of their function calling feature have been largely overlo…

2024

Principled Probabilistic Imaging using Diffusion Models as Plug-and-Play Priors

NeurIPS 2024poster

Diffusion models (DMs) have recently shown outstanding capabilities in modeling complex image distributions, making them expressive image priors for solving Bayesian inverse problems. However, most existing DM-based methods rely on approximations in the generative process to be generic to different…

2023

Demystifying Oversmoothing in Attention-Based Graph Neural Networks

NeurIPS 2023spotlight

Oversmoothing in Graph Neural Networks (GNNs) refers to the phenomenon where increasing network depth leads to homogeneous node representations. While previous work has established that Graph Convolutional Networks (GCNs) exponentially lose expressive power, it remains controversial whether the grap…

Cited by 57SourcePDFScholar