← Search

Yichen Wang

35 accepted papers

2026

Optimizing Diversity and Quality through Base-Aligned Model Collaboration

ICML 2026poster

Alignment has greatly improved large language models (LLMs)’ output quality at the cost of diversity, yielding highly similar outputs across generations, especially in open-ended generation tasks. We propose Base-Aligned Model Collaboration (BACo), an inference-time token-level model collaboration f…

Cited by 0SourceScholar
2026

sleep2vec: Unified Cross-Modal Alignment for Heterogeneous Nocturnal Biosignals

ICLR 2026poster

Tasks ranging from sleep staging to clinical diagnosis traditionally rely on standard polysomnography (PSG) devices, bedside monitors and wearable devices, which capture diverse nocturnal biosignals (e.g., EEG, EOG, ECG, SpO$_2$). However, heterogeneity across devices and frequent sensor dropout pos…

Cited by 2SourceScholar
2025

AdvEDM: Fine-grained Adversarial Attack against VLM-based Embodied Agents

NeurIPS 2025poster

Vision-Language Models (VLMs), with their strong reasoning and planning capabilities, are widely used in embodied decision-making (EDM) tasks in embodied agents, such as autonomous driving and robotic manipulation. Recent research has increasingly explored adversarial attacks on VLMs to reveal their…

Cited by 0SourceScholar
2025

BadRobot: Jailbreaking Embodied LLM Agents in the Physical World

ICLR 2025poster

Embodied AI represents systems where AI is integrated into physical entities. Multimodal Large Language Model (LLM), which exhibits powerful language understanding abilities, has been extensively employed in embodied AI by facilitating sophisticated task planning. However, a critical safety issue re…

Cited by 0SourcePDFScholar
2025

Breaking Barriers in Physical-World Adversarial Examples: Improving Robustness and Transferability via Robust Feature

AAAI 2025technical

As deep neural networks (DNNs) are widely applied in the physical world, many researches are focusing on physical-world adversarial examples (PAEs), which introduce perturbations to inputs and cause the model's incorrect outputs. However, existing PAEs face two challenges: unsatisfactory attack perf…

2025

Feature Learning beyond the Lazy-Rich Dichotomy: Insights from Representational Geometry

ICML 2025spotlight

Integrating task-relevant information into neural representations is a fundamental ability of both biological and artificial intelligence systems. Recent theories have categorized learning into two regimes: the rich regime, where neural networks actively learn task-relevant features, and the lazy r…

Cited by 0SourcePDFScholar
2025

HACo-Det: A Study Towards Fine-Grained Machine-Generated Text Detection under Human-AI Coauthoring

ACL 2025long

The misuse of large language models (LLMs) poses potential risks, motivating the development of machine-generated text (MGT) detection. Existing literature primarily concentrates on binary, document-level detection, thereby neglecting texts that are composed jointly by human and LLM contributions. H…

Cited by 0SourcePDFScholar
2025

Jailbreak Large Vision-Language Models Through Multi-Modal Linkage

ACL 2025long

With the rapid advancement of Large Vision-Language Models (VLMs), concerns about their ‌potential misuse and abuse have grown rapidly. Prior research has exposed VLMs’ vulnerability to jailbreak attacks, where carefully crafted inputs can lead the model to produce content that violates ethical and…

2025

Kinodynamic Model Predictive Control for Energy Efficient Locomotion of Legged Robots with Parallel Elasticity

ICRA 2025

In this paper, we introduce a kinodynamic model predictive control (MPC) framework that exploits unidirectional parallel springs (UPS) to improve the energy efficiency of dynamic legged robots. The proposed method employs a hierarchical control structure, where the solution of MPC with simplified dy

Cited by 1SourceScholar
2025

New Network Protocol for Supermedia-Enhanced Telerobotics

IROS 2025

The growing complexity of robotic teleoperation systems necessitates the integration of multiple feedback modalities, including video, audio, force, tactile, and temperature feedback. The concept of supermedia is utilized to describe the aggregation of these feedback streams. By integrating multiple

Cited by 0SourceScholar
2025

Pedestrian Motion Reconstruction: A Large-scale Benchmark via Mixed Reality Rendering with Multiple Perspectives and Modalities

ICLR 2025poster

Reconstructing pedestrian motion from dynamic sensors, with a focus on pedestrian intention, is crucial for advancing autonomous driving safety. However, this task is challenging due to data limitations arising from technical complexities, safety, and cost concerns. We introduce the Pedestrian Motio…

Cited by 0SourcePDFScholar
2025

Test-Time Backdoor Detection for Object Detection Models

CVPR 2025poster

Object detection models are vulnerable to backdoor attacks, where attackers poison a small subset of training samples by embedding a predefined trigger to manipulate prediction. Detecting poisoned samples (i.e., those containing triggers) at test time can prevent backdoor activation. However, unlike…

Cited by 1SourcePDFScholar
2025

The $\varphi$ Curve: The Shape of Generalization through the Lens of Norm-based Capacity Control

NeurIPS 2025poster

Understanding how the test risk scales with model complexity is a central question in machine learning. Classical theory is challenged by the learning curves observed for large over-parametrized deep networks. Capacity measures based on parameter count typically fail to account for these empirical o…

Cited by 0SourceScholar
2025

Unraveling Misinformation Propagation in LLM Reasoning

EMNLP 2025

Large Language Models (LLMs) have demonstrated impressive capabilities in reasoning, positioning them as promising tools for supporting human problem-solving. However, what happens when their performance is affected by *misinformation*, i.e., incorrect inputs introduced by users due to oversights or

2024

Concentrate Attention: Towards Domain-Generalizable Prompt Optimization for Language Models

NeurIPS 2024poster

Recent advances in prompt optimization have notably enhanced the performance of pre-trained language models (PLMs) on downstream tasks. However, the potential of optimized prompts on domain generalization has been under-explored. To explore the nature of prompt generalization on unknown domains, we…

2024

DarkFed: A Data-Free Backdoor Attack in Federated Learning

IJCAI 2024poster

Federated learning (FL) has been demonstrated to be susceptible to backdoor attacks. However, existing academic studies on FL backdoor attacks rely on a high proportion of real clients with main task-related data, which is impractical. In the context of real-world industrial scenarios, even the simp…

2024

Detector Collapse: Backdooring Object Detection to Catastrophic Overload or Blindness in the Physical World

IJCAI 2024poster

Object detection tasks, crucial in safety-critical systems like autonomous driving, focus on pinpointing object locations. These detectors are known to be susceptible to backdoor attacks. However, existing backdoor techniques have primarily been adapted from classification tasks, overlooking deeper…

Cited by 13SourcePDFScholar
2024

Dialogue for Prompting: A Policy-Gradient-Based Discrete Prompt Generation for Few-Shot Learning

AAAI 2024technical

Prompt-based pre-trained language models (PLMs) paradigm has succeeded substantially in few-shot natural language processing (NLP) tasks. However, prior discrete prompt optimization methods require expert knowledge to design the base prompt set and identify high-quality prompts, which is costly, ine…

2024

Does DetectGPT Fully Utilize Perturbation? Bridging Selective Perturbation to Fine-tuned Contrastive Learning Detector would be Better

ACL 2024long

The burgeoning generative capabilities of large language models (LLMs) have raised growing concerns about abuse, demanding automatic machine-generated text detectors. DetectGPT, a zero-shot metric-based detector, first introduces perturbation and shows great performance improvement. However, in Dete…

2024

Enhancing Exploratory Capability of Visual Navigation Using Uncertainty of Implicit Scene Representation

IROS 2024

In the context of visual navigation in unknown scenes, both “exploration” and “exploitation” are equally crucial. Robots must first establish environmental cognition through exploration and then utilize the cognitive information to accomplish target searches. However, most existing methods for image

Cited by 2SourcecodeScholar
2024

SemStamp: A Semantic Watermark with Paraphrastic Robustness for Text Generation

NAACL 2024long

Existing watermarked generation algorithms employ token-level designs and therefore, are vulnerable to paraphrase attacks. To address this issue, we introduce watermarking on the semantic representation of sentences. We propose SemStamp, a robust sentence-level semantic watermarking algorithm that u…

2024

Stumbling Blocks: Stress Testing the Robustness of Machine-Generated Text Detectors Under Attacks

ACL 2024long

The widespread use of large language models (LLMs) is increasing the demand for methods that detect machine-generated text to prevent misuse. The goal of our study is to stress test the detectors’ robustness to malicious attacks under realistic scenarios. We comprehensively study the robustness of p…

2024

k-SemStamp: A Clustering-Based Semantic Watermark for Detection of Machine-Generated Text

ACL 2024findings

Recent watermarked generation algorithms inject detectable signatures during language generation to facilitate post-hoc detection. While token-level watermarks are vulnerable to paraphrase attacks, SemStamp (Hou et al., 2023) applies watermark on the semantic representation of sentences and demonstr…

2020

On Transferability of Histological Tissue Labels in Computational Pathology

ECCV 2020poster

Deep learning tools in computational pathology, unlike natural vision tasks, face with limited histological tissue labels for classification. This is due to expensive procedure of annotation done by expert pathologist. As a result, the current models are limited to particular diagnostic task in mind…

2018

A Stochastic Differential Equation Framework for Guiding Online User Activities in Closed Loop

AISTATS 2018poster

Recently, there is a surge of interest in using point processes to model continuous-time user activities. This framework has resulted in novel models and improved performance in diverse applications. However, most previous works focus on the ”open loop” setting where learned models are used for pred…

Cited by 0SourcePDFScholar
2017

Know-Evolve: Deep Temporal Reasoning for Dynamic Knowledge Graphs

ICML 2017poster

The availability of large scale event data with time stamps has given rise to dynamically evolving knowledge graphs that contain temporal information for each edge. Reasoning over time in such dynamic knowledge graphs is not yet well understood. To this end, we present Know-Evolve, a novel deep evol…

2017

Linking Micro Event History to Macro Prediction in Point Process Models

AISTATS 2017poster

User behaviors in social networks are microscopic with fine grained temporal information. Predicting a macroscopic quantity based on users’ collective behaviors is an important problem. However, existing works are mainly problem-specific models for the microscopic behaviors and typically design appr…

Cited by 26SourcePDFScholar
2017

Predicting User Activity Level In Point Processes With Mass Transport Equation

NeurIPS 2017poster

Point processes are powerful tools to model user activities and have a plethora of applications in social sciences. Predicting user activities based on point processes is a central problem. However, existing works are mostly problem specific, use heuristics, or simplify the stochastic nature of poin…

Cited by 19SourcePDFScholar
2016

Coevolutionary Latent Feature Processes for Continuous-Time User-Item Interactions

NeurIPS 2016poster

Matching users to the right items at the right time is a fundamental task in recommendation systems. As users interact with different items over time, users' and items' feature may evolve and co-evolve over time. Traditional models based on static latent features or discretizing time into epochs can…

Cited by 74SourcePDFScholar
2015

COEVOLVE: A Joint Point Process Model for Information Diffusion and Network Co-evolution

NeurIPS 2015oral

Information diffusion in online social networks is affected by the underlying network topology, but it also has the power to change it. Online users are constantly creating new links when exposed to new information sources, and in turn these links are alternating the way information spreads. However…

2015

Double differential transmission for two-way relay systems with unknown carrier frequency offsets

ICASSP 2015accepted

In this paper, an amplify-and-forward two-way relay system with unknown carrier frequency offsets (CFOs) is considered. A double differential transmission scheme is proposed to achieve successful two-way relaying transmission without any CFOs information. The average symbol error rate (SER) performa…

Cited by 0SourceScholar
2015

Time-Sensitive Recommendation From Recurrent User Activities

NeurIPS 2015poster

By making personalized suggestions, a recommender system is playing a crucial role in improving the engagement of users in modern web-services. However, most recommendation algorithms do not explicitly take into account the temporal behavior and the recurrent activities of users. Two central but les…

Cited by 173SourcePDFScholar