← Search

Yu Qi

24 accepted papers

2026

Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis

ICML 2026poster

Understanding the capability bottlenecks of embodied multimodal large language models (MLLMs) is crucial for improvement. However, existing embodied benchmarks fail to provide actionable insights because they focus on task-level evaluation rather than discovering capability bottlenecks. To address t…

Cited by 0SourceScholar
2026

EquAct: An SE(3)-Equivariant Multi-Task Transformer for 3D Robotic Manipulation

ICLR 2026poster

Multi-task manipulation policy often builds on transformer's ability to jointly process language instructions and 3D observations in a shared embedding space. However, real-world tasks frequently require robots to generalize to novel 3D object poses. Policies based on shared embedding break geometri…

Cited by 0SourceScholar
2026

Generalizable Hierarchical Skill Learning via Object-Centric Representation

RA-L 2026

We present Generalizable Hierarchical Skill Learning (GSL), a novel framework for hierarchical policy learning that improves policy generalization and sample efficiency in robot manipulation. One core idea of GSL is to use object-centric skills as an interface that bridges the high-level vision-lang

Cited by 3SourceScholar
2026

Parse, Search, and Confirmation: Training-Free Aerial Vision-and-Dialog Navigation with Chain-of-Thought Reasoning and Structured Spatial Memory

CVPR 2026

In this paper, we tackle the Aerial Vision-and-Dialog Navigation (AVDN) task in the training-free setting for resource-efficient high-altitude UAV navigation.Naively applying MLLMs leads to unreliable navigation due to weak directional grounding and the lack of explicit spatial memory.To address the

Cited by 0SourcecodeScholar
2026

Reducing Semantic Mismatch in Brain-to-Text Decoding Through Personalized Multimodal Masking

ICLR 2026poster

The rapid progress of large vision-language models (VLMs), such as CLIP, has spurred the development of a wide range of neural decoding frameworks. Nevertheless, most existing approaches still suffer from semantic mismatches during representational alignment. This challenge may stem from the fact th…

Cited by 0SourceScholar
2025

Bridging the Gap between Brain and Machine in Interpreting Visual Semantics: Towards Self-adaptive Brain-to-Text Decoding

ICCV 2025poster

Neural decoding has recently made significant progress in reconstructing images and text from brain activity, yet seeking biologically valid semantic alignment between artificial models and the brain remains challenging. Large pre-trained foundation models such as CLIP excel at capturing rich semant…

2025

CRRL: Learning Channel-invariant Neural Representations for High-performance Cross-day Decoding

NeurIPS 2025poster

Brain-computer interfaces have shown great potential in motor and speech rehabilitation, but still suffer from low performance stability across days, mostly due to the instabilities in neural signals. These instabilities, partially caused by neuron deaths and electrode shifts, leading to channel-lev…

Cited by 0SourceScholar
2025

Cauchy Diffusion: A Heavy-tailed Denoising Diffusion Probabilistic Model for Speech Synthesis

AAAI 2025technical

Denoising diffusion probabilistic models (DDPMs) have gained popularity in devising neural vocoders and obtained outstanding performance. However, existing DDPM-based neural vocoders struggle to handle the prosody diversities due to their susceptibility to mode-collapse issues confronted with imbala…

Cited by 0SourcePDFScholar
2025

DeCorrNet: Enhancing Neural Decoding Performance by Eliminating Correlations in Noise

AAAI 2025technical

Neural decoding, which transforms neural signals into motor commands, plays a key role in brain-computer interfaces (BCIs). Existing neural decoding approaches mainly rely on the assumption of independent noises, which could perform poorly in case the assumption is invalid. However, correlations in…

Cited by 0SourcePDFScholar
2025

Flow Matching for Few-Trial Neural Adaptation with Stable Latent Dynamics

ICML 2025poster

The primary goal of brain-computer interfaces (BCIs) is to establish a direct linkage between neural activities and behavioral actions via neural decoders. Due to the nonstationary property of neural signals, BCIs trained on one day usually obtain degraded performance on other days, hindering the us…

Cited by 0SourcePDFScholar
2025

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

ICML 2025poster

Answering questions with Chain-of-Thought (CoT) has significantly enhanced the reasoning capabilities of Large Language Models (LLMs), yet its impact on Large Multimodal Models (LMMs) still lacks a systematic assessment and in-depth investigation. In this paper, we introduce **MME-CoT**, a specializ…

Cited by 0SourcePDFScholar
2025

Two by Two: Learning Multi-Task Pairwise Objects Assembly for Generalizable Robot Manipulation

CVPR 2025poster

3D assembly tasks, such as furniture assembly and component fitting, play a crucial role in daily life and represent essential capabilities for future home robots. Existing benchmarks and datasets predominantly focus on assembling geometric fragments or factory parts, which fall short in addressing…

2024

Bridging the Semantic Latent Space between Brain and Machine: Similarity Is All You Need

AAAI 2024technical

How our brain encodes complex concepts has been a longstanding mystery in neuroscience. The answer to this problem can lead to new understandings about how the brain retrieves information in large-scale data with high efficiency and robustness. Neuroscience studies suggest the brain represents conce…

Cited by 5SourcePDFScholar
2024

ThinkGrasp: A Vision-Language System for Strategic Part Grasping in Clutter

CoRL 2024poster

Robotic grasping in cluttered environments remains a significant challenge due to occlusions and complex object arrangements. We have developed ThinkGrasp, a plug-and-play vision-language grasping system that makes use of GPT-4o's advanced contextual reasoning for grasping strategies. ThinkGrasp can…

Cited by 14SourcecodeScholar
2023

Exploring Stochastic Autoregressive Image Modeling for Visual Representation

AAAI 2023technical

Autoregressive language modeling (ALM) has been successfully used in self-supervised pre-training in Natural language processing (NLP). However, this paradigm has not achieved comparable results with other self-supervised approaches in computer vision (e.g., contrastive learning, masked image modeli…

2023

Extracting Semantic-Dynamic Features for Long-Term Stable Brain Computer Interface

AAAI 2023technical

Brain-computer Interface (BCI) builds a neural signal to the motor command pathway, which is a prerequisite for the realization of neural prosthetics. However, a long-term stable BCI suffers from the neural data drift across days while retraining the BCI decoder is expensive and restricts its applic…

Cited by 4SourcePDFScholar
2023

Rethinking Visual Reconstruction: Experience-Based Content Completion Guided by Visual Cues

ICML 2023poster

Decoding seen images from brain activities has been an absorbing field. However, the reconstructed images still suffer from low quality with existing studies. This can be because our visual system is not like a camera that ''remembers'' every pixel. Instead, only part of the information can be perce…

Cited by 7SourcePDFScholar
2022

Combined Fast Control of Drifting State and Trajectory Tracking for Autonomous Vehicles Based on MPC Controller

ICRA 2022poster

Slipping may cause a vehicle out of control with serious accident potential. However, a kind of car slipping named “drifting” can be seen in professional contests. So, it is reasonable to apply drift maneuvers in autonomous driving. This article proposes a controller for the particular driving skill…

Cited by 16SourceScholar
2022

Tracking Functional Changes in Nonstationary Signals with Evolutionary Ensemble Bayesian Model for Robust Neural Decoding

NeurIPS 2022accept

Neural signals are typical nonstationary data where the functional mapping between neural activities and the intentions (such as the velocity of movements) can occasionally change. Existing studies mostly use a fixed neural decoder, thus suffering from an unstable performance given neural functional…

Cited by 3SourcePDFScholar
2019

Dynamic Ensemble Modeling Approach to Nonstationary Neural Decoding in Brain-Computer Interfaces

NeurIPS 2019poster

Brain-computer interfaces (BCIs) have enabled prosthetic device control by decoding motor movements from neural activities. Neural signals recorded from cortex exhibit nonstationary property due to abrupt noises and neuroplastic changes in brain activities during motor control. Current state-of-the-…

Cited by 26SourcePDFScholar
2018

Epileptic State Segmentation with Temporal-Constrained Clustering

ICASSP 2018accepted

Automatic seizure identification plays an important role in epilepsy evaluation. Most existing methods regard seizure identification as a classification problem and rely on labelled training set. However, labelling seizure onset is very expensive and seizure data for each individual is especially li…

Cited by 0SourceScholar