← Search

Xinpeng Wang

16 accepted papers

2026

Is it Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort

ICLR 2026oral

Reward hacking, where a reasoning model exploits loopholes in a reward function to achieve high rewards without solving the intended task, poses a significant threat. This behavior may be explicit, i.e. verbalized in the model's chain-of-thought (CoT), or implicit, where the CoT appears benign thus…

Cited by 0SourceScholar
2025

Algorithmic Fidelity of Large Language Models in Generating Synthetic German Public Opinions: A Case Study

ACL 2025long

In recent research, large language models (LLMs) have been increasingly used to investigate public opinions. This study investigates the algorithmic fidelity of LLMs, i.e., the ability to replicate the socio-cultural context and nuanced opinions of human participants. Using open-ended survey data fr…

2025

Refusal Direction is Universal Across Safety-Aligned Languages

NeurIPS 2025poster

Refusal mechanisms in large language models (LLMs) are essential for ensuring safety. Recent research has revealed that refusal behavior can be mediated by a single direction in activation space, enabling targeted interventions to bypass refusals. While this is primarily demonstrated in an English-c…

Cited by 0SourceScholar
2025

Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation

ICLR 2025poster

Training a language model to be both helpful and harmless requires careful calibration of refusal behaviours: Models should refuse to follow malicious instructions or give harmful advice (e.g."how do I kill someone?"), but they should not refuse safe requests, even if they superficially resemble uns…

Cited by 2SourcePDFScholar
2024

DAMRO: Dive into the Attention Mechanism of LVLM to Reduce Object Hallucination

EMNLP 2024main

Despite the great success of Large Vision-Language Models (LVLMs), they inevitably suffer from hallucination. As we know, both the visual encoder and the Large Language Model (LLM) decoder in LVLMs are Transformer-based, allowing the model to extract visual information and generate text outputs via…

2024

On the Essence and Prospect: An Investigation of Alignment Approaches for Big Models

IJCAI 2024poster

Big models have achieved revolutionary breakthroughs in the field of AI, but they also pose potential ethical and societal risks to humans. Addressing such problems, alignment technologies were introduced to make these models conform to human preferences and values. Despite the considerable advancem…

Cited by 12SourcePDFScholar
2024

The Potential and Challenges of Evaluating Attitudes, Opinions, and Values in Large Language Models

EMNLP 2024finding

Recent advances in Large Language Models (LLMs) have sparked wide interest in validating and comprehending the human-like cognitive-behavioral traits LLMs may capture and convey. These cognitive-behavioral traits include typically Attitudes, Opinions, Values (AOVs). However, measuring AOVs embedded…

2024

“My Answer is C”: First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models

ACL 2024findings

The open-ended nature of language generation makes the evaluation of autoregressive large language models (LLMs) challenging. One common evaluation approach uses multiple-choice questions to limit the response space. The model is then evaluated by ranking the candidate answers by the log probability…

2024

“Seeing the Big through the Small”: Can LLMs Approximate Human Judgment Distributions on NLI from a Few Explanations?

EMNLP 2024finding

Human label variation (HLV) is a valuable source of information that arises when multiple human annotators provide different labels for valid reasons. In Natural Language Inference (NLI) earlier approaches to capturing HLV involve either collecting annotations from many crowd workers to represent hu…

2023

ACTOR: Active Learning with Annotator-specific Classification Heads to Embrace Human Label Variation

EMNLP 2023short main

Label aggregation such as majority voting is commonly used to resolve annotator disagreement in dataset creation. However, this may disregard minority values and opinions. Recent studies indicate that learning from individual annotations outperforms learning from aggregated labels, though they requ…

Cited by 0SourceScholar
2023

How to Distill your BERT: An Empirical Study on the Impact of Weight Initialisation and Distillation Objectives

ACL 2023short

Recently, various intermediate layer distillation (ILD) objectives have been shown to improve compression of BERT models via Knowledge Distillation (KD). However, a comprehensive evaluation of the objectives in both task-specific and task-agnostic settings is lacking. To the best of our knowledge, t…

2023

Large-Scale and Multi-Perspective Opinion Summarization with Diverse Review Subsets

EMNLP 2023long findings

Opinion summarization is expected to digest larger review sets and provide summaries from different perspectives. However, most existing solutions are deficient in epitomizing extensive reviews and offering opinion summaries from various angles due to the lack of designs for information selection. T…

Cited by 0SourcecodeScholar
2023

ToViLaG: Your Visual-Language Generative Model is Also An Evildoer

EMNLP 2023long main

Recent large-scale Visual-Language Generative Models (VLGMs) have achieved unprecedented improvement in multimodal image/text generation. However, these models might also generate toxic content, e.g., offensive text and pornography images, raising significant ethical risks. Despite exhaustive studie…

Cited by 0SourcecodeScholar
2022

CHAE: Fine-Grained Controllable Story Generation with Characters, Actions and Emotions

COLING 2022main

Story generation has emerged as an interesting yet challenging NLP task in recent years. Some existing studies aim at generating fluent and coherent stories from keywords and outlines; while others attempt to control the global features of the story, such as emotion, style and topic. However, these…

2021

Deep Reinforcement Learning for Multi-contact Motion Planning of Hexapod Robots

IJCAI 2021poster

Legged locomotion in a complex environment requires careful planning of the footholds of legged robots. In this paper, a novel Deep Reinforcement Learning (DRL) method is proposed to implement multi-contact motion planning for hexapod robots moving on uneven plum-blossom piles. First, the motion of…

Cited by 15SourcePDFScholar