← Search

Yi Zhao

45 accepted papers

2026

CMoE: Contrastive Mixture of Experts for Motion Control and Terrain Adaptation of Humanoid Robots

ICRA 2026poster

For effective deployment in real-world environments, humanoid robots must autonomously navigate a diverse range of complex terrains with abrupt transitions. While the Vanilla mixture of experts (MoE) framework is theoretically capable of modeling diverse terrain features, in practice, the gating net…

2026

Efficient Reinforcement Learning by Guiding World Models with Non-Curated Data

ICLR 2026poster

Leveraging offline data is a promising way to improve the sample efficiency of online reinforcement learning (RL). This paper expands the pool of usable data for offline-to-online RL by leveraging abundant non-curated data that is reward-free, of mixed quality, and collected across multiple embodime…

Cited by 0SourcecodeScholar
2026

Geometry-Preserving Unsupervised Alignment for Heterogeneous Foundation Models

ICML 2026poster

Foundation models have driven rapid progress in computer vision, yet the two dominant paradigm, vision-language foundation models (VLMs) and vision-only foundation models (VFMs), remain only partially compatible. VLMs offer language-grounded semantic alignment but are often visually coarse, while VF…

Cited by 0SourceScholar
2026

HIVE-3D: Hierarchical Voxel Enhancement for High-Quality 3D Scene Generation

ICML 2026poster

Recently, a line of works can generate impressive 3D objects from a single image, but they are limited by restricted representation resolution, making them unsuitable for 3D scene generation. In this work, we introduce \name, a novel method for high-quality 3D scene generation based on hierarchical …

Cited by 0SourceScholar
2026

Influence without Confounding: Causal Discovery from Temporal Data with Long-term Carry-over Effects

ICLR 2026poster

Learning causal structures from temporal data is fundamental to many practical tasks, such as physical laws discovery and root causes localization. Real-world systems often exhibit long-term carry-over effects, where the value of a variable at the current time can be influenced by distant past va…

Cited by 0SourceScholar
2026

LaF-GRPO: In-Situ Navigation Instruction Generation for the Visually Impaired via GRPO with LLM-as-Follower Reward

AAAI 2026technical

Navigation instruction generation for visually impaired (VI) individuals (NIG-VI) is critical yet relatively underexplored. This study focuses on generating precise, in-situ, step-by-step navigation instructions that are practically usable for VI users. Specifically, we propose LaF-GRPO (LLM-as-Foll

Cited by 0SourcePDFScholar
2026

MedLesionVQA: A Multimodal Benchmark Emulating Clinical Visual Diagnosis for Body Surface Health

ICLR 2026poster

Body-surface health conditions, spanning diverse clinical departments, represent some of the most frequent diagnostic scenarios and a primary target for medical multimodal large language models (MLLMs). Yet existing medical benchmarks are either built from publicly available sources with limited ex…

Cited by 0SourceScholar
2026

[CLS] is Not Enough: Multi-Label Recognition via Patch-Level Inference and Adaptive Aggregation

ICML 2026poster

Vision-Language Models such as CLIP exhibit strong zero-shot recognition capability by aligning images with textual concepts, yet they often underperform on multi-label recognition where multiple objects co-exist. A key bottleneck is that the CLS token, as a single global visual representation, is i…

Cited by 0SourceScholar
2025

DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression

ACL 2025long

Task-agnostic prompt compression leverages the redundancy in natural language to reduce computational overhead and enhance information density within prompts, especially in long-context scenarios. Existing methods predominantly rely on information entropy as the metric to compress lexical units, aim…

2025

Discrete Codebook World Models for Continuous Control

ICLR 2025poster

In reinforcement learning (RL), world models serve as internal simulators, enabling agents to predict environment dynamics and future outcomes in order to make informed decisions. While previous approaches leveraging discrete latent spaces, such as DreamerV3, have demonstrated strong performance in…

2025

Extracting Visual Plans from Unlabeled Videos via Symbolic Guidance

CoRL 2025poster

Visual planning, by offering a sequence of intermediate visual subgoals to a goal-conditioned low-level policy, achieves promising performance on long-horizon manipulation tasks. To obtain the subgoals, existing methods typically resort to video generation models but suffer from model hallucination…

Cited by 0SourceScholar
2025

Generative Adversarial Network with Adaptive Synthesis for Brain-Computer Interfaces in Motor Imagery Classification

ICASSP 2025accepted

Motor Imagery (MI) is essential in Brain-Computer Interfaces (BCIs), highlighting the central role of electroencephalography (EEG) in this technology. However, the amount of raw EEG data is often limited. Raw EEG data contains significant noise caused by individual and task-specific differences. The…

Cited by 0SourceScholar
2025

HEROS-GAN: Honed-Energy Regularized and Optimal Supervised GAN for Enhancing Accuracy and Range of Low-Cost Accelerometers

AAAI 2025technical

Low-cost accelerometers play a crucial role in modern society due to their advantages of small size, ease of integration, wearability, and mass production, making them widely applicable in automotive systems, aerospace, and wearable technology. However, this widely used sensor suffers from severe ac…

Cited by 0SourcePDFScholar
2025

HGS-Planner: Hierarchical Planning Framework for Active Scene Reconstruction Using 3D Gaussian Splatting

ICRA 2025

In complex missions such as search and rescue, robots must make intelligent decisions in unknown environments, relying on their ability to perceive and understand their surroundings. High-quality and real-time reconstruction enhances situational awareness and is crucial for intelligent robotics. Tra

Cited by 19SourceScholar
2025

MDRNet: Multi-Branch with Different Feature Representations Network for Motor Imagery Classification

ICASSP 2025accepted

A brain-computer interface (BCI) offers an innovative solution for facilitating communication and control in individuals with paralysis. BCI reflects brain activity by decoding electroencephalogram (EEG) signals. Despite numerous techniques for classifying motor imagery (MI) EEG signals, challenges…

Cited by 0SourceScholar
2025

SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference

NeurIPS 2025spotlight

KV cache eviction has emerged as an effective solution to alleviate resource constraints faced by LLMs in long-context scenarios. However, existing token-level eviction methods often overlook two critical aspects: (1) their irreversible eviction strategy fails to adapt to dynamic attention patterns…

Cited by 0SourceScholar
2025

Topology-Driven Trajectory Optimization for Modelling Controllable Interactions Within Multi-Vehicle Scenario

IROS 2025

Trajectory optimization in multi-vehicle scenarios faces challenges due to its non-linear, non-convex properties and sensitivity to initial values, making interactions between vehicles difficult to control. In this paper, inspired by topological planning, we propose a differentiable local homotopy i

Cited by 0SourceScholar
2025

Towards Unified and Lossless Latent Space for 3D Molecular Latent Diffusion Modeling

NeurIPS 2025poster

3D molecule generation is crucial for drug discovery and material science, requiring models to process complex multi-modalities, including atom types, chemical bonds, and 3D coordinates. A key challenge is integrating these modalities of different shapes while maintaining SE(3) equivariance for 3D c…

Cited by 0SourcecodeScholar
2024

Adversarial Robust Safeguard for Evading Deep Facial Manipulation

AAAI 2024technical

The non-consensual exploitation of facial manipulation has emerged as a pressing societal concern. In tandem with the identification of such fake content, recent research endeavors have advocated countering manipulation techniques through proactive interventions, specifically the incorporation of ad…

Cited by 3SourcePDFScholar
2024

Enhancing Realism in 3D Facial Animation Using Conformer-Based Generation and Automated Post-Processing

ICASSP 2024accepted

Recent progress has propelled the development of realistic talking-face videos for avatars. Yet, animating 3D cartoon avatars remains intricate due to the imprecise nature of facial-driven data. This often manifests as inconsistent mouth configurations and rigid facial expressions, curbing the anima…

Cited by 0SourceScholar
2024

FanLoRA: Fantastic LoRAs and Where to Find Them in Large Language Model Fine-tuning

EMNLP 2024industry

Full-parameter fine-tuning is computationally prohibitive for large language models (LLMs), making parameter-efficient fine-tuning (PEFT) methods like low-rank adaptation (LoRA) increasingly popular. However, LoRA and its existing variants introduce significant latency in multi-tenant settings, hind…

2024

GFMAE: Self-Supervised GNN-Free Masked Autoencoders

ICASSP 2024accepted

Generative self-supervised learning, represented by graph autoencoders (GAEs), has begun to exhibit significant potential in addressing graph tasks. However, GAEs often rely on Graph Neural Networks (GNNs) for encoding and decoding, this can pose a computation challenge due to the inherent complexit…

Cited by 0SourceScholar
2024

High-Fidelity Speech Synthesis with Minimal Supervision: All Using Diffusion Models

ICASSP 2024accepted

Text-to-speech (TTS) methods have shown promising results in voice cloning, but they require a large number of labeled text-speech pairs. Minimally-supervised speech synthesis decouples TTS by combining two types of discrete speech representations(semantic & acoustic) and using two sequence-to-seque…

Cited by 0SourceScholar
2024

MOS-FAD: Improving Fake Audio Detection Via Automatic Mean Opinion Score Prediction

ICASSP 2024accepted

IEEE Automatic Mean Opinion Score (MOS) prediction is employed to evaluate the quality of synthetic speech. This study extends the application of predicted MOS to the task of Fake Audio Detection (FAD) as we expect that MOS can be used to assess how close synthesized speech is to the natural human v…

Cited by 0SourceScholar
2024

MiLoRA: Efficient Mixture of Low-Rank Adaptation for Large Language Models Fine-tuning

EMNLP 2024finding

Low-rank adaptation (LoRA) and its mixture-of-experts (MOE) variants are highly effective parameter-efficient fine-tuning (PEFT) methods. However, they introduce significant latency in multi-tenant settings due to the LoRA modules and MOE routers added to multiple linear modules in the Transformer l…

2024

PARA: Parameter-Efficient Fine-tuning with Prompt-Aware Representation Adjustment

EMNLP 2024industry

In the realm of parameter-efficient fine-tuning (PEFT) methods, while options like LoRA are available, there is a persistent demand in the industry for a PEFT approach that excels in both efficiency and performance within the context of single-backbone multi-tenant applications. This paper introduce…

2024

RP1M: A Large-Scale Motion Dataset for Piano Playing with Bi-Manual Dexterous Robot Hands

CoRL 2024poster

Endowing robot hands with human-level dexterity is a long-lasting research objective. Bi-manual robot piano playing constitutes a task that combines challenges from dynamic tasks, such as generating fast while precise motions, with slower but contact-rich manipulation problems. Although reinforcemen…

Cited by 2SourceScholar
2024

RU22Fact: Optimizing Evidence for Multilingual Explainable Fact-Checking on Russia-Ukraine Conflict

COLING 2024main

Fact-checking is the task of verifying the factuality of a given claim by examining the available evidence. High-quality evidence plays a vital role in enhancing fact-checking systems and facilitating the generation of explanations that are understandable to humans. However, the provision of both su…

2023

Simplified Temporal Consistency Reinforcement Learning

ICML 2023poster

Reinforcement learning (RL) is able to solve complex sequential decision-making tasks but is currently limited by sample efficiency and required computation. To improve sample efficiency, recent work focuses on model-based RL which interleaves model learning with planning. Recent methods further uti…

2022

Melons: Generating Melody With Long-Term Structure Using Transformers And Structure Graph

ICASSP 2022accepted

The creation of long melody sequences requires effective expression of coherent musical structure. However, there is no clear representation of musical structure. Recent works on music generation have suggested various approaches to deal with the structural information of music, but generating a ful…

Cited by 0SourceScholar
2022

UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition

EMNLP 2022main

Multimodal sentiment analysis (MSA) and emotion recognition in conversation (ERC) are key research topics for computers to understand human behaviors. From a psychological perspective, emotions are the expression of affect or feelings during a short period, while sentiments are formed and held for a…

2021

Bidirectional Hierarchical Attention Networks based on Document-level Context for Emotion Cause Extraction

EMNLP 2021finding

Emotion cause extraction (ECE) aims to extract the causes behind the certain emotion in text. Some works related to the ECE task have been published and attracted lots of attention in recent years. However, these methods neglect two major issues: 1) pay few attentions to the effect of document-level…

2021

EM-POSE: 3D Human Pose Estimation From Sparse Electromagnetic Trackers

ICCV 2021poster

Fully immersive experiences in AR/VR depend on reconstructing the full body pose of the user without restricting their motion. In this paper we study the use of body-worn electromagnetic (EM) field-based sensing for the task of 3D human pose reconstruction. To this end, we present a method to estima…

Cited by 40PDFcodeScholar
2021

Learning Disentangled Phone and Speaker Representations in a Semi-Supervised VQ-VAE Paradigm

ICASSP 2021accepted

We present a new approach to disentangle speaker voice and phone content by introducing new components to the VQ-VAE architecture for speech synthesis. The original VQ-VAE does not generalize well to unseen speakers or content. To alleviate this problem, we have incorporated a speaker encoder and sp…

Cited by 0SourceScholar
2020

Hierarchical Scene Coordinate Classification and Regression for Visual Localization

CVPR 2020poster

Visual localization is critical to many applications in computer vision and robotics. To address single-image RGB localization, state-of-the-art feature-based methods match local descriptors between a query image and a pre-built 3D model. Recently, deep neural networks have been exploited to regress…

Cited by 152PDFScholar
2020

Transferring Neural Speech Waveform Synthesizers to Musical Instrument Sounds Generation

ICASSP 2020accepted

Recent neural waveform synthesizers such as WaveNet, WaveG-low, and the neural-source-filter (NSF) model have shown good performance in speech synthesis despite their different methods of waveform generation. The similarity between speech and music audio synthesis techniques suggests interesting ave…

Cited by 0SourceScholar