← Search

Yifan Zhu

53 accepted papers

2026

Branch, or Layer? Zeroth-Order Optimization for Continual Learning of Vision-Language Models

AAAI 2026technical

Vision-Language Continual Learning (VLCL) has attracted significant research attention for its robust capabilities, and the adoption of Parameter-Efficient Fine-Tuning (PEFT) strategies is enabling these models to achieve competitive performance with substantially reduced resource consumption. Howev

Cited by 0SourcePDFScholar
2026

CreBench: Human-Aligned Creativity Evaluation from Idea to Process to Product

AAAI 2026technical

Human-defined creativity is highly abstract, posing a challenge for multimodal large language models (MLLMs) to comprehend and assess creativity that aligns with human judgments. The absence of an existing benchmark further exacerbates this dilemma. To this end, we propose CreBench, which consists o

Cited by 0SourcePDFScholar
2026

Domain Adaptation with Adaptive $f$-Divergence: Tighter Variational Representation and Generalization Bounds

ICML 2026poster

We study unsupervised domain adaptation (UDA) where measuring cross-domain discrepancy is critical. Most UDA approaches fix a single $f$-divergence a priori, which can be suboptimal across heterogeneous shifts. We propose a framework that (i) tightens the variational lower bound of an $f$-divergence…

Cited by 0SourceScholar
2026

ERGeoBench: A Comprehensive Benchmark for Embodied Reasoning and Geo-localization in Multimodal Large Language Models

ICML 2026poster

Multimodal large language models (MLLMs) have shown strong potential for building embodied agents, yet embodied geo-localization remains underexplored due to the lack of fine-grained evaluation. We introduce ERGeoBench, a large-scale benchmark for vision-driven embodied geo-localization. ERGeoBench …

Cited by 0SourceScholar
2026

Gotta Scoop 'Em All: Sim-And-Real Co-Training of Graph-Based Neural Dynamics for Long-Horizon Scooping

ICRA 2026poster

Robotic manipulation of granular objects is crucial in various fields, yet modeling their complex dynamics and diverse physical properties remains challenging. Simulation plays an important role in learning robotic manipulation policies, but it exhibits challenge to accurately model the complex dyna…

Cited by 0Scholar
2026

Graph-R1: Towards Agentic GraphRAG Framework via End-to-end Reinforcement Learning

ICML 2026poster

Retrieval-Augmented Generation (RAG) mitigates hallucination in LLMs by incorporating external knowledge, but relies on chunk-based retrieval that lacks structural semantics. GraphRAG methods improve RAG by modeling knowledge as entity-relation graphs, but still face challenges in high construction …

Cited by 0SourceScholar
2026

HEDP: A Hybrid Energy-Distance Prompt-based Framework for Domain Incremental Learning

ICML 2026poster

Domain Incremental Learning is a critical scenario that requires models to continuously adapt to new data domains without retraining. However, domain shifts often cause severe performance degradation. To address this, we propose Hybrid Energy-Distance Prompt, a domain-incremental framework inspired …

Cited by 0SourceScholar
2026

LearNAT: Learning NL2SQL with AST-guided Task Decomposition for Large Language Models

ICLR 2026poster

Natural Language to SQL (NL2SQL) aims to translate natural language queries into executable SQL statements, offering non-expert users intuitive access to databases. While recent approaches leveraging large-scale private LLMs such as GPT-4 have achieved state-of-the-art results, they face two critica…

Cited by 0SourcecodeScholar
2026

Token-Guard: Towards Token-Level Hallucination Control via Self-Checking Decoding

ICLR 2026poster

Large Language Models (LLMs) often hallucinate, generating content inconsistent with the input. Retrieval-Augmented Generation (RAG) and Reinforcement Learning with Human Feedback (RLHF) can mitigate hallucinations but require resource-intensive retrieval or large-scale fine-tuning. Decoding-based m…

Cited by 0SourcecodeScholar
2026

Why Do Unlearnable Examples Work: A Novel Perspective of Mutual Information

ICLR 2026poster

The volume of freely scraped data on the Internet has driven the tremendous success of deep learning. Along with this comes the rising concern about data privacy and security. Numerous methods for generating unlearnable examples have been proposed to prevent data from being illicitly learned by unau…

Cited by 0SourceScholar
2025

A Stronger Mixture of Low-Rank Experts for Fine-Tuning Foundation Models

ICML 2025poster

In order to streamline the fine-tuning of foundation models, Low-Rank Adapters (LoRAs) have been substantially adopted across various fields, including instruction tuning and domain adaptation. The underlying concept of LoRA involves decomposing a full-rank matrix into the product of two lower-rank…

2025

Complex Numerical Reasoning with Numerical Semantic Pre-training Framework

EMNLP 2025

Multi-hop complex reasoning over incomplete knowledge graphs (KGs) has been extensively studied, but research on numerical knowledge graphs (NKGs) remains relatively limited. Recent approaches focus on separately encoding entities and numerical values, using neural networks to process query encoding

Cited by 0SourcePDFScholar
2025

DecoupledESC: Enhancing Emotional Support Generation via Strategy-Response Decoupled Preference Optimization

EMNLP 2025

Recent advances in Emotional Support Conversation (ESC) have improved emotional support generation by fine-tuning Large Language Models (LLMs) via Supervised Fine-Tuning (SFT). However, common psychological errors still persist. While Direct Preference Optimization (DPO) shows promise in reducing su

2025

DivGCL: A Graph Contrastive Learning Model for Diverse Recommendation

AAAI 2025technical

Graph Contrastive Learning (GCL), as a primary paradigm of graph self-supervised learning, spurs a fruitful line of research in tackling the data sparsity issue by maximizing the consistency of user/item embeddings between different augmented views with random perturbations. However, diversity, as a…

Cited by 1SourcePDFScholar
2025

Generalizability of Neural Networks Minimizing Empirical Risk Based on Expressive Power

ICLR 2025poster

The primary objective of learning methods is generalization. Classic generalization bounds, based on VC-dimension or Rademacher complexity, are uniformly applicable to all networks in the hypothesis space. On the other hand, algorithm-dependent generalization bounds, like stability bounds, address m…

Cited by 0SourcePDFScholar
2025

HyperGraphRAG: Retrieval-Augmented Generation via Hypergraph-Structured Knowledge Representation

NeurIPS 2025poster

Standard Retrieval-Augmented Generation (RAG) relies on chunk-based retrieval, whereas GraphRAG advances this approach by graph-based knowledge representation. However, existing graph-based RAG approaches are constrained by binary relations, as each edge in an ordinary graph connects only two entiti…

Cited by 0SourceScholar
2025

INFER: A Neural-symbolic Model For Extrapolation Reasoning on Temporal Knowledge Graph

ICLR 2025poster

Temporal Knowledge Graph(TKG) serves as an efficacious way to store dynamic facts in real-world. Extrapolation reasoning on TKGs, which aims at predicting possible future events, has attracted consistent research interest. Recently, some rule-based methods have been proposed, which are considered mo…

Cited by 0SourcePDFScholar
2025

KBQA-o1: Agentic Knowledge Base Question Answering with Monte Carlo Tree Search

ICML 2025poster

Knowledge Base Question Answering (KBQA) aims to answer natural language questions with a large-scale structured knowledge base (KB). Despite advancements with large language models (LLMs), KBQA still faces challenges in weak KB awareness, imbalance between effectiveness and efficiency, and high rel…

2025

LS-TGNN: Long and Short-Term Temporal Graph Neural Network for Session-Based Recommendation

AAAI 2025technical

Session-Based Recommendation (SBR) based on Graph Neural Networks (GNN) has become a new paradigm for recommender systems, and plays a fundamental role in e-commerce and other relevant domains. Existing graph aggregation methods primarily form node representations by capturing basic relationships be…

Cited by 0SourcePDFScholar
2025

MASTER: Multi-Agent Security Through Exploration of Roles and Topological Structures - A Comprehensive Framework

EMNLP 2025

Large Language Models (LLMs)-based Multi-Agent Systems (MAS) exhibit remarkable problem-solving and task planning capabilities across diverse domains due to their specialized agentic roles and collaborative interactions. However, this also amplifies the severity of security risks under MAS attacks.

Cited by 0SourcePDFScholar
2025

One-Shot Real-to-Sim via End-to-End Differentiable Simulation and Rendering

RA-L 2025

Identifying predictive world models for robots from sparse online observations is essential for robot task planning and execution in novel environments. However, existing methods that leverage differentiable programming to identify world models are incapable of jointly optimizing the geometry, appea

Cited by 5SourcecodeScholar
2025

PowerMLP: An Efficient Version of KAN

AAAI 2025technical

The Kolmogorov-Arnold Network (KAN) is a new network architecture known for its high accuracy in several tasks such as function fitting and PDE solving. The superior expressive capability of KAN arises from the Kolmogorov-Arnold representation theorem and learnable spline functions. However, the com…

2025

Provable Robust Overfitting Mitigation in Wasserstein Distributionally Robust Optimization

ICLR 2025poster

Wasserstein distributionally robust optimization (WDRO) optimizes against worst-case distributional shifts within a specified uncertainty set, leading to enhanced generalization on unseen adversarial examples, compared to standard adversarial training which focuses on pointwise adversarial perturbat…

2025

Speech Is Not Enough: Interpreting Nonverbal Indicators of Common Knowledge and Engagement

AAAI 2025technical

Our goal is to develop an AI Partner that can provide support for group problem solving and social dynamics. In multi-party working group environments, multimodal analytics is crucial for identifying non-verbal interactions of group members. In conjunction with their verbal participation, this creat…

Cited by 1SourcePDFScholar
2025

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making

EMNLP 2025

Using effective generalization capabilities of vision language models (VLMs) in context-specific dynamic tasks for embodied artificial intelligence remains a significant challenge. Although supervised fine-tuned models can better align with the real physical world, they still exhibit sluggish respon

Cited by 0SourcePDFScholar
2025

TRACE: Real-Time Multimodal Common Ground Tracking in Situated Collaborative Dialogues

NAACL 2025system demonstrations

We present TRACE, a novel system for live *common ground* tracking in situated collaborative tasks. With a focus on fast, real-time performance, TRACE tracks the speech, actions, gestures, and visual attention of participants, uses these multimodal inputs to determine the set of task-relevant propos…

Cited by 0SourcePDFScholar
2025

TSVC: Tripartite Learning with Semantic Variation Consistency for Robust Image-Text Retrieval

AAAI 2025technical

Cross-modal retrieval maps data under different modalities via semantic relevance. Existing approaches implicitly assume that data pairs are well-aligned and ignore the widely existing annotation noise, i.e., noisy correspondence (NC). Consequently, it inevitably causes performance degradation. Desp…

Cited by 0SourcePDFScholar
2025

Towards Recognizing Spatial-temporal Collaboration of EEG Phase Brain Networks for Emotion Understanding

IJCAI 2025

Emotion recognition from EEG signals is crucial for understanding complex brain dynamics. Existing methods typically rely on static frequency bands and graph convolutional networks (GCNs) to model brain connectivity. However, EEG signals are inherently non-stationary and exhibit substantial individu

2025

Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents

EMNLP 2025

Role-playing agents (RPAs) have attracted growing interest for their ability to simulate immersive and interactive characters. However, existing approaches primarily focus on static role profiles, overlooking the dynamic perceptual abilities inherent to humans. To bridge this gap, we introduce the c

Cited by 0SourcePDFScholar
2024

ChatKBQA: A Generate-then-Retrieve Framework for Knowledge Base Question Answering with Fine-tuned Large Language Models

ACL 2024findings

Knowledge Base Question Answering (KBQA) aims to answer natural language questions over large-scale knowledge bases (KBs), which can be summarized into two crucial steps: knowledge retrieval and semantic parsing. However, three core challenges remain: inefficient knowledge retrieval, mistakes of ret…

2024

Common Ground Tracking in Multimodal Dialogue

COLING 2024main

Within Dialogue Modeling research in AI and NLP, considerable attention has been spent on “dialogue state tracking” (DST), which is the ability to update the representations of the speaker’s needs at each turn in the dialogue by taking into account the past dialogue moves and history. Less studied b…

2024

Efficient Availability Attacks against Supervised and Contrastive Learning Simultaneously

NeurIPS 2024poster

Availability attacks provide a tool to prevent the unauthorized use of private data and commercial datasets by generating imperceptible noise and crafting unlearnable examples before release. Ideally, the obtained unlearnability can prevent algorithms from training usable models. When supervised l…

2024

Structured Bayesian Meta-Learning for Data-Efficient Visual-Tactile Model Estimation

CoRL 2024poster

Estimating visual-tactile models of deformable objects is challenging because vision suffers from occlusion, while touch data is sparse and noisy. We propose a novel data-efficient method for dense heterogeneous model estimation by leveraging experience from diverse training objects. The method is…

Cited by 1SourceScholar
2024

T2VSafetyBench: Evaluating the Safety of Text-to-Video Generative Models

NeurIPS 2024poster

The recent development of Sora leads to a new era in text-to-video (T2V) generation. Along with this comes the rising concern about its safety risks. The generated videos may contain illegal or unethical content, and there is a lack of comprehensive quantitative understanding of their safety, posing…

2024

Text2NKG: Fine-Grained N-ary Relation Extraction for N-ary relational Knowledge Graph Construction

NeurIPS 2024poster

Beyond traditional binary relational facts, n-ary relational knowledge graphs (NKGs) are comprised of n-ary relational facts containing more than two entities, which are closer to real-world facts with broader applications. However, the construction of NKGs remains at a coarse-grained level, which i…

2023

Consistent Depth Prediction for Transparent Object Reconstruction from RGB-D Camera

ICCV 2023poster

Transparent objects are commonly seen in indoor scenes but are hard to estimate. Currently, commercial depth cameras face difficulties in estimating the depth of transparent objects due to the light reflection and refraction on their surface. As a result, they tend to make a noisy and incorrect dept…

Cited by 6PDFScholar
2023

Few-shot Adaptation for Manipulating Granular Materials Under Domain Shift

RSS 2023poster

Autonomous lander missions on extraterrestrial bodies will need to sample granular material while coping with domain shift, no matter how well a sampling strategy is tuned on Earth. This paper proposes an adaptive scooping strategy that uses deep Gaussian process method trained with meta-learning to…

Cited by 8SourcePDFScholar
2023

Looking Through the Glass: Neural Surface Reconstruction Against High Specular Reflections

CVPR 2023poster

Neural implicit methods have achieved high-quality 3D object surfaces under slight specular highlights. However, high specular reflections (HSR) often appear in front of target objects when we capture them through glasses. The complex ambiguity in these scenes violates the multi-view consistency, th…

2022

Excavation of Fragmented Rocks with Multi-modal Model-based Reinforcement Learning

IROS 2022poster

This paper presents a multi-modal model-based reinforcement learning (MBRL) approach to the excavation of fragmented rocks, which are very challenging to model due to their highly variable sizes and geometries, and visual occlusions. A multi-modal recurrent neural network (RNN) learns the dynamics o…

Cited by 9SourceScholar
2021

A Novel Robotic System for Ultrasound-guided Peripheral Vascular Localization

ICRA 2021poster

In this paper, we present an autonomous RGB-D and 2D ultrasound-guided robotic system for collecting 3D localized volumes of peripheral vessels. This compact design, with available commercial components, lends itself to platform utility throughout the human body. The fully integrated system works wi…

Cited by 18SourceScholar
2021

Contact-Implicit Trajectory Optimization With Learned Deformable Contacts Using Bilevel Optimization

ICRA 2021poster

We present a bilevel, contact-implicit trajectory optimization (TO) formulation that searches for robot trajectories with learned soft contact models. On the lower-level, contact forces are solved via a quadratic program (QP) with the maximum dissipation principle (MDP), based on which the dynamics…

Cited by 13SourceScholar
2020

Semi-Empirical Simulation of Learned Force Response Models for Heterogeneous Elastic Objects

ICRA 2020poster

This paper presents a semi-empirical method for simulating contact with elastically deformable objects whose force response is learned using entirely data-driven models. A point-based surface representation and an inhomogeneous, nonlinear force response model are learned from a robotic arm acquiring…

Cited by 3SourceScholar
2019

A Data-driven Approach for Fast Simulation of Robot Locomotion on Granular Media

ICRA 2019poster

In this paper, we propose a semi-empirical approach for simulating robot locomotion on granular media. We first develop a contact model based on the stick-slip behavior between rigid objects and granular grains, which is then learned through running extensive experiments. The contact model represent…

Cited by 24SourceScholar