← Search

Yuhan Liu

51 accepted papers

2026

Conditional Information Bottleneck for Multimodal Fusion: Overcoming Shortcut Learning in Sarcasm Detection

AAAI 2026technical

Multimodal sarcasm detection is a complex task that requires distinguishing subtle complementary signals across modalities while filtering out irrelevant information. Many advanced methods rely on learning shortcuts from datasets rather than extracting intended sarcasm-related features. However, our

Cited by 0SourcePDFScholar
2026

Graph2Eval: Automatic Multimodal Task Generation for Agents via Knowledge Graphs

CVPR 2026

As multimodal LLM-driven agents advance in autonomy and generalization, traditional static datasets face inherent scalability limitations and are insufficient for fully assessing their capabilities in increasingly complex and diverse tasks. Existing studies have attempted to generate agent tasks usi

Cited by 0SourcecodeScholar
2026

One-Shot Flow, Any-Time Frame: A Bidirectional Warping Framework for Event-Based Video Frame Interpolation

CVPR 2026

Video Frame Interpolation (VFI) is a crucial task in video processing. Flow-based methods, despite their success, are constrained by a fundamental dilemma: forward warping is efficient but prone to artifacts, while backward warping yields higher quality at a significant computational cost, especiall

Cited by 0SourcecodeScholar
2026

Position Is All You Need: A Free Lunch Token Compression Strategy for MLLM-based Referring Expression Segmentation

ICML 2026poster

Referring Expression Segmentation (RES) aims to generate pixel-wise segmentation masks from complex and implicit textual queries. While recent advances in Multimodal Large Language Models (MLLMs) have substantially boosted RES performance, their prohibitive computational overhead remains a critical …

Cited by 0SourceScholar
2026

Structured Progressive Knowledge Activation for LLM-Driven Neural Architecture Search

ICML 2026poster

This paper focuses on a key challenge in Neural Architecture Search (NAS): integrating established architectural knowledge while exploring new designs under expensive evaluations. Large language models (LLMs) are a promising assistant for NAS because they can translate rich architectural and coding …

Cited by 0SourceScholar
2026

ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Models

ICLR 2026poster

The limited capacity for fine-grained visual perception presents a critical bottleneck for Vision-Language Models (VLMs) in real-world applications. Addressing this is challenging due to the scarcity of high-quality data and the limitations of existing methods: supervised fine-tuning (SFT) often com…

Cited by 0SourcecodeScholar
2025

Autoregressive Action Sequence Learning for Robotic Manipulation

RA-L 2025

Designing a universal policy architecture that performs well across diverse robots and task configurations remains a key challenge. In this work, we address this by representing robot actions as sequential data and generating actions through autoregressive sequence modeling. Existing autoregressive

Cited by 37SourcecodeScholar
2025

Beyond Static Testbeds: An Interaction-Centric Agent Simulation Platform for Dynamic Recommender Systems

EMNLP 2025

Evaluating and iterating upon recommender systems is crucial, yet traditional A/B testing is resource-intensive, and offline methods struggle with dynamic user-platform interactions. While agent-based simulation is promising, existing platforms often lack a mechanism for user actions to dynamically

2025

Boundary-to-Region Supervision for Offline Safe Reinforcement Learning

NeurIPS 2025poster

Offline safe reinforcement learning aims to learn policies that satisfy predefined safety constraints from static datasets. Existing sequence-model-based methods condition action generation on symmetric input tokens for return-to-go and cost-to-go, neglecting their intrinsic asymmetry: RTG serves as…

Cited by 0SourceScholar
2025

DeMAC: Enhancing Multi-Agent Coordination with Dynamic DAG and Manager-Player Feedback

EMNLP 2025

Multi-agent systems (MAS) powered by large language models (LLMs) have shown potential in tackling multifaceted problems through advanced understanding and reasoning. However, they struggle to adapt to evolving task dependencies and to handle uncertainties, such as shifting priorities or unpredictab

Cited by 0SourcePDFScholar
2025

EPA: Boosting Event-based Video Frame Interpolation with Perceptually Aligned Learning

NeurIPS 2025poster

Event cameras, with their capacity to provide high temporal resolution information between frames, are increasingly utilized for video frame interpolation (VFI) in challenging scenarios characterized by high-speed motion and significant occlusion. However, prevalent issues of blur and distortion wit…

Cited by 0SourceScholar
2025

EmoAgent: Assessing and Safeguarding Human-AI Interaction for Mental Health Safety

EMNLP 2025

The rise of LLM-driven AI characters raises safety concerns, particularly for vulnerable human users with psychological disorders. To address these risks, we propose EmoAgent, a multi-agent AI framework designed to evaluate and mitigate mental health hazards in human-AI interactions. EmoAgent compri

2025

Failure Forecasting Boosts Robustness of Sim2Real Rhythmic Insertion Policies

IROS 2025

This paper addresses the challenges of Rhythmic Insertion Tasks (RIT), where a robot must repeatedly perform high-precision insertions, such as screwing a nut into a bolt with a wrench. The inherent difficulty of RIT lies in achieving millimeter-level accuracy and maintaining consistent performance

Cited by 0SourcecodeScholar
2025

Injecting Domain-Specific Knowledge into Large Language Models: A Comprehensive Survey

EMNLP 2025

Large Language Models (LLMs) have demonstrated remarkable success in various tasks such as natural language understanding, text summarization, and machine translation. However, their general-purpose nature often limits their effectiveness in domain-specific applications that require specialized know

Cited by 0SourcePDFScholar
2025

Mind the Gap: Aligning Vision Foundation Models to Image Feature Matching

ICCV 2025poster

Leveraging the vision foundation models has emerged as a mainstream paradigm that improves the performance of image feature matching. However, previous works have ignored the misalignment when introducing the foundation models into feature matching. The misalignment arises from the discrepancy betwe…

Cited by 0SourcePDFScholar
2025

More is not always better? Enhancing Many-Shot In-Context Learning with Differentiated and Reweighting Objectives

ACL 2025long

Large language models (LLMs) excel at few-shot in-context learning (ICL) without requiring parameter updates. However, as ICL demonstrations increase from a few to many, performance tends to plateau and eventually decline. We identify two primary causes for this trend: the suboptimal negative log-li…

2025

Not All Parameters Matter: Masking Diffusion Models for Enhancing Generation Ability

CVPR 2025poster

The diffusion models, in early stages focus on constructing basic image structures, while the refined details, including local features and textures, are generated in later stages. Thus the same network layers are forced to learn both structural and textural information simultaneously, significant…

2025

ParetoRAG: Leveraging Sentence-Context Attention for Robust and Efficient Retrieval-Augmented Generation

EMNLP 2025

While Retrieval-Augmented Generation systems enhance Large Language Models by incorporating external knowledge, they still face persistent challenges in retrieval inefficiency and the inability of LLMs to filter out irrelevant information. We presentParetoRAG, an unsupervised framework that optimize

Cited by 0SourcePDFScholar
2025

The Devil is in Low-Level Features for Cross-Domain Few-Shot Segmentation

CVPR 2025poster

Cross-Domain Few-Shot Segmentation (CDFSS) is proposed to transfer the pixel-level segmentation capabilities learned from large-scale source-domain datasets to downstream target-domain datasets, with only a few annotated images per class. In this paper, we focus on a well-observed but unresolved phe…

Cited by 1SourcePDFScholar
2025

The Stepwise Deception: Simulating the Evolution from True News to Fake News with LLM Agents

EMNLP 2025

With the growing spread of misinformation online, understanding how true news evolves into fake news has become crucial for early detection and prevention. However, previous research has often assumed fake news inherently exists rather than exploring its gradual formation. To address this gap, we pr

2025

Thinking Before Running! Efficient Code Generation with Thorough Exploration and Optimal Refinement

ACL 2025finding

Code generation is crucial in software engineering for automating the coding process efficiently. While test-time computation methods show promise, they suffer from high latency due to multiple computation rounds.To overcome this, we introduce ThinkCoder, a framework that combines thorough explorati…

Cited by 0SourcePDFScholar
2025

User Feedback in Human-LLM Dialogues: A Lens to Understand Users But Noisy as a Learning Signal

EMNLP 2025

Once language models (LMs) are deployed, they can interact with users long-term, ideally evolving based on their feedback. Asking for direct user feedback can be disruptive; thus, we study harvesting implicit user feedback from user-LM interaction logs. We study two user-LM interaction datasets (Wil

Cited by 0SourcePDFScholar
2025

Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains

ACL 2025long

Vision-language models (VLMs) achieve remarkable success in single-image tasks. However, real-world scenarios often involve intricate multi-image inputs, leading to a notable performance decline as models struggle to disentangle critical information scattered across complex visual features. In this…

Cited by 0SourcePDFScholar
2024

A3VLM: Actionable Articulation-Aware Vision Language Model

CoRL 2024poster

Vision Language Models (VLMs) for robotics have received significant attention in recent years. As a VLM can understand robot observations and perform complex visual reasoning, it is regarded as a potential universal solution for general robotics challenges such as manipulation and navigation. Howev…

Cited by 12SourcecodeScholar
2024

DAP: Diffusion-based Affordance Prediction for Multi-modality Storage

IROS 2024poster

Solving storage problems—where objects must be accurately placed into containers with precise orientations and positions—presents a distinct challenge that extends beyond traditional rearrangement tasks. These challenges are primarily due to the need for fine-grained 6D manipulation and the inherent…

Cited by 1SourcecodeScholar
2024

Enhancing Document-Level Event Extraction via Structure-Aware Heterogeneous Graph with Multi-Granularity Subsentences

ICASSP 2024accepted

Document-level Event Extraction aims to identify events from an entire article. It is quite a challenging task because event arguments scatter across several sentences and multiple events in a document may have influence on each other. Previous methods, however, did not take advantage of document st…

Cited by 0SourceScholar
2024

From Skepticism to Acceptance: Simulating the Attitude Dynamics Toward Fake News

IJCAI 2024poster

In the digital era, the rapid propagation of fake news and rumors via social networks brings notable societal challenges and impacts public opinion regulation. Traditional fake news modeling typically forecasts the general popularity trends of different groups or numerically represents opinions shif…

2024

IAD: In-Context Learning Ability Decoupler of Large Language Models in Meta-Training

COLING 2024main

Large Language Models (LLMs) exhibit remarkable In-Context Learning (ICL) ability, where the model learns tasks from prompts consisting of input-output examples. However, the pre-training objectives of LLMs often misalign with ICL objectives. They’re mainly pre-trained with methods like masked langu…

Cited by 2SourcePDFScholar
2024

Knowledge Crosswords: Geometric Knowledge Reasoning with Large Language Models

ACL 2024findings

We propose Knowledge Crosswords, a geometric knowledge reasoning benchmark consisting of incomplete knowledge networks bounded by structured factual constraints, where LLMs are tasked with inferring the missing facts to meet all constraints. The novel setting of geometric knowledge reasoning necessi…

2024

Modular Pluralism: Pluralistic Alignment via Multi-LLM Collaboration

EMNLP 2024main

While existing alignment paradigms have been integral in developing large language models (LLMs), LLMs often learn an averaged human preference and struggle to model diverse preferences across cultures, demographics, and communities. We propose Modular Pluralism, a modular framework based on multi-L…

2024

P3Sum: Preserving Author’s Perspective in News Summarization with Diffusion Language Models

NAACL 2024long

In this work, we take a first step towards designing summarization systems that are faithful to the author’s intent, not only the semantic content of the article. Focusing on a case study of preserving political perspectives in news summarization, we find that existing approaches alter the political…

2024

SAM-Event-Adapter: Adapting Segment Anything Model for Event-RGB Semantic Segmentation

ICRA 2024poster

Semantic segmentation, a fundamental visual task ubiquitously employed in sectors ranging from transportation and robotics to healthcare, has always captivated the research community. In the wake of rapid advancements in large model research, the foundation model for semantic segmentation tasks, ter…

Cited by 11SourceScholar
2024

Scaling Manipulation Learning with Visual Kinematic Chain Prediction

CoRL 2024poster

Learning general-purpose models from diverse datasets has achieved great success in machine learning. In robotics, however, existing methods in multi-task learning are typically constrained to a single robot and workspace, while recent work such as RT-X requires a non-trivial action normalization pr…

Cited by 1SourcecodeScholar
2024

Video Frame Interpolation via Direct Synthesis with the Event-based Reference

CVPR 2024poster

Video Frame Interpolation (VFI) has witnessed a surge in popularity due to its abundant downstream applications. Event-based VFI (E-VFI) has recently propelled the advancement of VFI. Thanks to the high temporal resolution benefits event cameras can bridge the informational void present between succ…

Cited by 6SourcePDFScholar
2023

Algorithms for bounding contribution for histogram estimation under user-level privacy

ICML 2023poster

We study the problem of histogram estimation under user-level differential privacy, where the goal is to preserve the privacy of *all* entries of any single user. We consider the heterogeneous scenario where the quantity of data can be different for each user. In this scenario, the amount of noise i…

Cited by 10SourcePDFScholar
2023

Discrete Distribution Estimation under User-level Local Differential Privacy

AISTATS 2023poster

We study discrete distribution estimation under user-level local differential privacy (LDP). In user-level $\varepsilon$-LDP, each user has a $m\ge1$ samples and the privacy of all $m$ samples must be preserved simultaneously. We resolve the following dilemma: While on the one hand having more sampl…

2023

Echo of Neighbors: Privacy Amplification for Personalized Private Federated Learning with Shuffle Model

AAAI 2023technical

Federated Learning, as a popular paradigm for collaborative training, is vulnerable against privacy attacks. Different privacy levels regarding users' attitudes need to be satisfied locally, while a strict privacy guarantee for the global model is also required centrally. Personalized Local Differen…

Cited by 13SourcePDFScholar
2023

From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models

ACL 2023long

Language models (LMs) are pretrained on diverse data sources—news, discussion forums, books, online encyclopedias. A significant portion of this data includes facts and opinions which, on one hand, celebrate democracy and diversity of ideas, and on the other hand are inherently socially biased. Our…

2023

Learning Continuous Control Policies for Information-Theoretic Active Perception

ICRA 2023poster

This paper proposes a method for learning continuous control policies for exploration and active landmark localization. We consider a mobile robot detecting landmarks within a limited sensing range, and tackle the problem of learning a control policy that maximizes the mutual information between the…

Cited by 14SourcecodeScholar
2022

TwiBot-22: Towards Graph-Based Twitter Bot Detection

NeurIPS 2022accept

Twitter bot detection has become an increasingly important task to combat misinformation, facilitate social media moderation, and preserve the integrity of the online discourse. State-of-the-art bot detection methods generally leverage the graph structure of the Twitter network, and they exhibit pro…

2021

Auto-calibration Method Using Stop Signs for Urban Autonomous Driving Applications

ICRA 2021poster

Calibration of sensors is fundamental to robust performance for intelligent vehicles. In natural environments, disturbances can easily challenge calibration. One possibility is to use natural objects of known shape to recalibrate sensors. An approach based on recognition of traffic signs, such as st…

Cited by 10SourceScholar
2021

Distributed Estimation with Multiple Samples per User: Sharp Rates and Phase Transition

NeurIPS 2021poster

We obtain tight minimax rates for the problem of distributed estimation of discrete distributions under communication constraints, where $n$ users observing $m $ samples each can broadcast only $\ell$ bits. Our main result is a tight characterization (up to logarithmic factors) of the error rate as…

Cited by 13SourcePDFScholar
2021

Improving Empathetic Response Generation by Recognizing Emotion Cause in Conversations

EMNLP 2021finding

Current approaches to empathetic response generation focus on learning a model to predict an emotion label and generate a response based on this label and have achieved promising results. However, the emotion cause, an essential factor for empathetic responding, is ignored. The emotion cause is a st…

Cited by 116SourcePDFScholar
2021

OpenRooms: An Open Framework for Photorealistic Indoor Scene Datasets

CVPR 2021poster

We propose a novel framework for creating large-scale photorealistic datasets of indoor scenes, with ground truth geometry, material, lighting and semantics. Our goal is to make the dataset creation process widely accessible, allowing researchers to transform scans into datasets with highquality gro…

Cited by 93PDFScholar
2020

Learning discrete distributions: user vs item-level privacy

NeurIPS 2020poster

Much of the literature on differential privacy focuses on item-level privacy, where loosely speaking, the goal is to provide privacy per item or training example. However, recently many practical applications such as federated learning require preserving privacy for all items of a single user, which…

Cited by 72SourcePDFScholar