← Search

Di Hu

44 accepted papers

2026

AnyTouch 2: General Optical Tactile Representation Learning For Dynamic Tactile Perception

ICLR 2026poster

Real-world contact-rich manipulation demands robots to perceive temporal tactile feedback, capture subtle surface deformations, and reason about object properties and force dynamics. Although optical tactile sensors are uniquely capable of providing such rich information, existing tactile datasets a…

Cited by 0SourcecodeScholar
2026

GeoISF: Instance Semantic Forest Inspired Large-Scale Cross-View Geo-Localization Via Ground LiDAR-To-Satellite Image

ICRA 2026poster

The problem of localization on a large-scale satellite image given a frame of query ground view point clouds remains challenging. Existing LiDAR-to-image cross-view localization methods struggle in large-scale scenarios due to limited semantic alignment and the modality gap between point clouds and …

2026

Imbalanced View Contribution Evaluation and Refinement for Deep Incomplete Multi-View Clustering

CVPR 2026

In real-world applications, multi-view data often suffer from missing situations due to privacy protection and sensor failures. Such incomplete scenarios not only reduce information availability but also cause significant imbalance among views: certain "strong views" dominate the fusion process, whi

Cited by 0SourcecodeScholar
2026

Information-Theoretic Decomposition for Multimodal Interaction Learning

CVPR 2026

Multimodal learning hinges on capturing redundant, unique, and synergistic information across modalities, which collectively constitute multimodal interactions. A critical yet underexplored challenge is that these implicit interactions vary dynamically across samples. In this work, we present the fi

Cited by 0SourcecodeScholar
2026

When would Vision-Proprioception Policies Fail in Robotic Manipulation?

ICLR 2026poster

Proprioceptive information is critical for precise servo control by providing real-time robotic states. Its collaboration with vision is highly expected to enhance performances of the manipulation policy in complex tasks. However, recent studies have reported inconsistent observations on the general…

Cited by 0SourcecodeScholar
2025

Adaptive Unimodal Regulation for Balanced Multimodal Information Acquisition

CVPR 2025poster

Sensory training during the early ages is vital for human development. Inspired by this cognitive phenomenon, we observe that the early training stage is also important for the multimodal learning process, where dataset information is rapidly acquired. We refer to this stage as the prime learning wi…

2025

AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-tactile Sensors

ICLR 2025poster

Visuo-tactile sensors aim to emulate human tactile perception, enabling robots to precisely understand and manipulate objects. Over time, numerous meticulously designed visuo-tactile sensors have been integrated into robotic systems, aiding in completing various tasks. However, the distinct data cha…

2025

Crab: A Unified Audio-Visual Scene Understanding Model with Explicit Cooperation

CVPR 2025poster

In recent years, numerous tasks have been proposed to encourage model to develop specified capability in understanding audio-visual scene, primarily categorized into temporal localization, spatial localization, spatio-temporal reasoning, and pixel-level understanding. Instead, human possesses a unif…

2025

Human-assisted Robotic Policy Refinement via Action Preference Optimization

NeurIPS 2025poster

Establishing a reliable and iteratively refined robotic system is essential for deploying real-world applications. While Vision-Language-Action (VLA) models are widely recognized as the foundation model for such robotic deployment, their reliance on offline expert demonstrations critically limi…

Cited by 0SourcecodeScholar
2025

Patch Matters: Training-free Fine-grained Image Caption Enhancement via Local Perception

CVPR 2025poster

High-quality image captions play a crucial role in improving the performance of cross-modal applications such as text-to-image generation, text-to-video generation, and text-image retrieval. To generate long-form, high-quality captions, many recent studies have employed multimodal large language mod…

Cited by 1SourcePDFScholar
2025

Phoenix: A Motion-based Self-Reflection Framework for Fine-grained Robotic Action Correction

CVPR 2025poster

Building a generalizable self-correction system is crucial for robots to recover from failures. Despite advancements in Multimodal Large Language Models (MLLMs) that empower robots with semantic reflection ability for failure, translating semantic reflection into how to correct fine-grained robotic…

2025

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer

ICML 2025poster

Multimodal learning faces challenges in effectively fusing information from diverse modalities, especially when modality quality varies across samples. Dynamic fusion strategies, such as attention mechanism in Transformers, aim to address such challenge by adaptively emphasizing modalities based on…

2025

Towards Effective and Efficient Continual Pre-training of Large Language Models

ACL 2025long

Continual pre-training (CPT) has been an important approach for adapting language models to specific domains or tasks. In this paper, we comprehensively study its key designs to balance the new abilities while retaining the original abilities, and present an effective CPT method that can greatly imp…

2024

Depth Helps: Improving Pre-trained RGB-based Policy with Depth Information Injection

IROS 2024poster

3D perception ability is crucial for generalizable robotic manipulation. While recent foundation models have made significant strides in perception and decision-making with RGB-based input, their lack of 3D perception limits their effectiveness in fine-grained robotic manipulation tasks. To address…

Cited by 2SourcecodeScholar
2024

Enhancing Multimodal Cooperation via Sample-level Modality Valuation

CVPR 2024poster

One primary topic of multimodal learning is to jointly incorporate heterogeneous information from different modalities. However most models often suffer from unsatisfactory multimodal cooperation which cannot jointly utilize all modalities well. Some methods are proposed to identify and enhance the…

2024

KOI: Accelerating Online Imitation Learning via Hybrid Key-state Guidance

CoRL 2024poster

Online Imitation Learning methods struggle with the gap between extensive online exploration space and limited expert trajectories, which hinder efficient exploration due to inaccurate task-aware reward estimation. Inspired by the findings from cognitive neuroscience that task decomposition coul…

Cited by 0SourceScholar
2024

Kinematic-aware Prompting for Generalizable Articulated Object Manipulation with LLMs

ICRA 2024poster

Generalizable articulated object manipulation is essential for home-assistant robots. Recent efforts focus on imitation learning from demonstrations or reinforcement learning in simulation, however, due to the prohibitive costs of real-world data collection and precise object simulation, it still re…

Cited by 23SourcecodeScholar
2024

Learning Manipulation by Predicting Interaction

RSS 2024poster

Representation learning approaches for robotic manipulation have boomed in recent years. Due to the scarcity of in-domain robot data, prevailing methodologies tend to leverage large-scale human video datasets to extract generalizable features for visuomotor policy learning. Despite the progress achi…

2024

Play to the Score: Stage-Guided Dynamic Multi-Sensory Fusion for Robotic Manipulation

CoRL 2024poster

Humans possess a remarkable talent for flexibly alternating to different senses when interacting with the environment. Picture a chef skillfully gauging the timing of ingredient additions and controlling the heat according to the colors, sounds, and aromas, seamlessly navigating through every stage…

Cited by 7SourceScholar
2024

Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer

AAAI 2024technical

Never having seen an object and heard its sound simultaneously, can the model still accurately localize its visual position from the input audio? In this work, we concentrate on the Audio-Visual Localization and Segmentation tasks but under the demanding zero-shot and few-shot scenarios. To achieve…

2024

Quantifying and Enhancing Multi-modal Robustness with Modality Preference

ICLR 2024poster

Multi-modal models have shown a promising capability to effectively integrate information from various sources, yet meanwhile, they are found vulnerable to pervasive perturbations, such as uni-modal attacks and missing conditions. To counter these perturbations, robust multi-modal representations ar…

2024

SphereDiffusion: Spherical Geometry-Aware Distortion Resilient Diffusion Model

AAAI 2024technical

Controllable spherical panoramic image generation holds substantial applicative potential across a variety of domains. However, it remains a challenging task due to the inherent spherical distortion and geometry characteristics, resulting in low-quality content generation. In this paper, we introduc…

Cited by 8SourcePDFScholar
2023

MMCosine: Multi-Modal Cosine Loss Towards Balanced Audio-Visual Fine-Grained Learning

ICASSP 2023accepted

Audio-visual learning helps to comprehensively under-stand the world by fusing practical information from multiple modalities. However, recent studies show that the imbalanced optimization of uni-modal encoders in a joint-learning model is a bottleneck to enhancing the model’s performance. We furthe…

Cited by 0SourceScholar
2023

Towards Inadequately Pre-trained Models in Transfer Learning

ICCV 2023poster

Transfer learning has been a popular learning paradigm in the deep learning era, especially in annotation-insufficient scenarios. Better ImageNet pre-trained models have been demonstrated, from the perspective of architecture, by previous research to have better transferability to downstream tasks.…

Cited by 11PDFScholar
2022

Balanced Multimodal Learning via On-the-Fly Gradient Modulation

CVPR 2022oral

Audio-visual learning helps to comprehensively understand the world, by integrating different senses. Accordingly, multiple input modalities are expected to boost model performance, but we actually find that they are not fully exploited even when the multi-modal model outperforms its uni-modal count…

Cited by 247PDFcodeScholar
2022

Learning To Answer Questions in Dynamic Audio-Visual Scenarios

CVPR 2022oral

In this paper, we focus on the Audio-Visual Question Answering (AVQA) task, which aims to answer questions regarding different visual objects, sounds, and their associations in videos. The problem requires comprehensive multimodal understanding and spatio-temporal reasoning over audio-visual scenes.…

Cited by 157PDFcodeScholar
2022

SepFusion: Finding Optimal Fusion Structures for Visual Sound Separation

AAAI 2022technical

Multiple modalities can provide rich semantic information; and exploiting such information will normally lead to better performance compared with the single-modality counterpart. However, it is not easy to devise an effective cross-modal fusion structure due to the variations of feature dimensions…

Cited by 15SourcePDFScholar
2022

Visual Sound Localization in the Wild by Cross-Modal Interference Erasing

AAAI 2022technical

The task of audiovisual sound source localization has been well studied under constrained scenes, where the audio recordings are clean. However, in real world scenarios, audios are usually contaminated by off screen sound and background noise. They will interfere with the procedure of identifying de…

2021

Temporal Relational Modeling with Self-Supervision for Action Segmentation

AAAI 2021technical

Temporal relational modeling in video is essential for human action understanding, such as action recognition and action segmentation. Although Graph Convolution Networks (GCNs) have shown promising advantages in relation reasoning on many tasks, it is still a challenge to apply graph convolution ne…

2021

Unsupervised Multi-Source Domain Adaptation for Person Re-Identification

CVPR 2021poster

Unsupervised domain adaptation (UDA) methods for person re-identification (re-ID) aim at transferring re-ID knowledge from labeled source data to unlabeled target data. Among these methods, the pseudo-label-based branch has achieved great success, whereas most of them only use limited data from a si…

Cited by 112PDFScholar
2020

Cross-Task Transfer for Geotagged Audiovisual Aerial Scene Recognition

ECCV 2020poster

Aerial scene recognition is a fundamental task in remote sensing and has recently received increased interest. While the visual information from overhead images with powerful models and efficient algorithms yields considerable performance on scene recognition, it still suffers from the variation of…

2020

Discriminative Sounding Objects Localization via Self-supervised Audiovisual Matching

NeurIPS 2020poster

Discriminatively localizing sounding objects in cocktail-party, i.e., mixed sound scenes, is commonplace for humans, but still challenging for machines. In this paper, we propose a two-stage learning framework to perform self-supervised class-aware sounding object localization. First, we propose to…

2020

Multiple Sound Sources Localization from Coarse to Fine

ECCV 2020poster

How to visually localize multiple sound sources in unconstrained videos is a formidable problem, especially when lack of the pairwise sound-object annotations. To solve this problem, we develop a two-stage audiovisual learning framework that disentangles audio and visual representations of different…