← Search

Xiangming Gu

7 accepted papers

2026

Extracting alignment data in open models

ICML 2026poster

In this work, we show that it is possible to extract significant amounts of alignment training data from a post-trained model -- useful to steer the model to improve certain capabilities such as long-context reasoning, safety, instruction following, and maths. While the majority of related work on m…

Cited by 0SourceScholar
2025

On Calibration of LLM-based Guard Models for Reliable Content Moderation

ICLR 2025poster

Large language models (LLMs) pose significant risks due to the potential for generating harmful content or users attempting to evade guardrails. Existing studies have developed LLM-based guard models designed to moderate the input and output of threat LLMs, ensuring adherence to safety policies by b…

2025

SkyLadder: Better and Faster Pretraining via Context Window Scheduling

NeurIPS 2025poster

Recent advancements in LLM pretraining have featured ever-expanding context windows to process longer sequences. However, our controlled study reveals that models pretrained with shorter context windows consistently outperform their long-context counterparts under a fixed token budget. This finding…

Cited by 0SourcecodeScholar
2025

When Attention Sink Emerges in Language Models: An Empirical View

ICLR 2025spotlight

Auto-regressive language Models (LMs) assign significant attention to the first token, even if it is not semantically important, which is known as **attention sink**. This phenomenon has been widely adopted in applications such as streaming/long context generation, KV cache optimization, inference a…

2024

Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast

ICML 2024poster

A multimodal large language model (MLLM) agent can receive instructions, capture images, retrieve histories from memory, and decide which tools to use. Nonetheless, red-teaming efforts have revealed that adversarial images/prompts can jailbreak an MLLM and cause unaligned behaviors. In this work, we…

2022

Extrapolative Continuous-time Bayesian Neural Network for Fast Training-free Test-time Adaptation

NeurIPS 2022accept

Human intelligence has shown remarkably lower latency and higher precision than most AI systems when processing non-stationary streaming data in real-time. Numerous neuroscience studies suggest that such abilities may be driven by internal predictive modeling. In this paper, we explore the possibili…

Cited by 15SourcePDFScholar
2021

Laser Endoscopic Manipulator Using Spring-Reinforced Multi-DoF Soft Actuator

RA-L 2021

The flexible manipulators and related robotic systems have been widely used in minimally invasive surgery (MIS) for enhancing the intraoperative inspection and surgical operation. Although a variety of flexible manipulators using different mechanisms have been developed, most of them are rigid mecha

Cited by 14SourceScholar