← Search

Yunhao Li

18 accepted papers

2026

DiffVL: Diffusion-Based Visual Localization on 2D Maps Via BEV-Conditioned GPS Denoising

ICRA 2026poster

Accurate visual localization is crucial for autonomous driving, yet existing methods face a fundamental dilemma: While high-definition (HD) maps provide high-precision localization references, their costly construction and maintenance hinder scalability, which drives research toward standard-definit…

2026

MotionMaster: Generalizable Text-Driven Motion Generation and Editing

CVPR 2026

Synthesizing realistic human motion from natural language holds transformative potential for animation, robotics, and virtual reality. Recent methods handle single-action sequences and simple textual instructions, yet multi-action compositions and precise editing remain elusive due to limited data d

Cited by 0SourcecodeScholar
2026

ODI-Bench: Can MLLMs Understand Immersive Omnidirectional Environments?

ICLR 2026poster

Omnidirectional images (ODIs) provide full 360$^{\circ} \times$ 180$^{\circ}$ view which are widely adopted in VR, AR and embodied intelligence applications. While multi-modal large language models (MLLMs) have demonstrated remarkable performance on conventional 2D image and video understanding benc…

Cited by 0SourceScholar
2026

TI-3DGS: 3D Thermal Reconstruction Via Thermal Imaging-Guided 3D Gaussian Splatting

ICRA 2026poster

Thermal imaging, with its all-weather capabilities and strong penetration, enables 3D reconstruction in low- light and adverse conditions. In this paper, we investigate RGB-independent pure 3D thermal reconstruction, aiming to overcome the challenges of 3D reconstruction in extreme environments wher…

Cited by 0Scholar
2025

Attention to Trajectory: Trajectory-Aware Open-Vocabulary Tracking

ICCV 2025poster

Open-Vocabulary Multi-Object Tracking (OV-MOT) aims to enable approaches to track objects without being limited to a predefined set of categories. Current OV-MOT methods typically rely primarily on instance-level detection and association, often overlooking trajectory information that is unique and…

2025

DFMA: Adaptive Dual Fusion for Multimodal Relation Extraction with Mutual Attention

ICASSP 2025accepted

Multimodal relation extraction (MRE) is an emerging research field that combines techniques from natural language processing, computer vision, and machine learning, helping us better understand and interpret data. However, current methods are faced with two main issues. The first issue is that the a…

Cited by 0SourceScholar
2025

GSOT3D: Towards Generic 3D Single Object Tracking in the Wild

ICCV 2025poster

In this paper, we present a novel benchmark, GSOT3D, that aims at facilitating development of generic 3D single object tracking (SOT) in the wild. Specifically, GSOT3D offers 620 sequences with 123K frames, and covers a wide selection of 54 object categories. Each sequence is offered with multiple m…

2025

SCI-Gaussian: Optimizing 3D Gaussian Radiance Fields from a Snapshot Compressive Image

ICASSP 2025accepted

Snapshot compressive imaging (SCI) is a compressed sensing (CS)-based high-speed imaging modality. Recent efforts have explored the underlying 3D representation from only an SCI image using neural radiance fields (NeRF), yet the training time, rendering computation cost, and reconstruction quality l…

Cited by 0SourceScholar
2025

WMAJL: Watcher-Mediated Attention Joint Learning Model for Multimodal Relation Extraction

ICASSP 2025accepted

In the domain of Multimodal Relation Extraction (MRE), we present the $\color{Red}{\text{W}}$atcher-$\color{Red}{\text{M}}$ediated $\color{Red}{\text{A}}$ttention $\color{Red}{\text{J}}$oint $\color{Red}{\text{L}}$earning Model ($\color{Red}{\text{WMAJL}}$), a novel approach addressing the challenge…

Cited by 0SourceScholar
2024

Beyond MOT: Semantic Multi-Object Tracking

ECCV 2024poster

"Current multi-object tracking (MOT) aims to predict trajectories of targets (, “where”) in videos. Yet, knowing merely “where” is insufficient in many crucial applications. In comparison, semantic understanding such as fine-grained behaviors, interactions, and overall summarized captions (, “what”)…

2024

DiffStega: Towards Universal Training-Free Coverless Image Steganography with Diffusion Models

IJCAI 2024poster

Traditional image steganography focuses on concealing one image within another, aiming to avoid steganalysis by unauthorized entities. Coverless image steganography (CIS) enhances imperceptibility by not using any cover image. Recent works have utilized text prompts as keys in CIS through diffusion…

2024

SCINeRF: Neural Radiance Fields from a Snapshot Compressive Image

CVPR 2024highlight

In this paper we explore the potential of Snapshot Com- pressive Imaging (SCI) technique for recovering the under- lying 3D scene representation from a single temporal com- pressed image. SCI is a cost-effective method that enables the recording of high-dimensional data such as hyperspec- tral or te…

2023

Creating a Dynamic Quadrupedal Robotic Goalkeeper with Reinforcement Learning

IROS 2023poster

We present a reinforcement learning (RL) framework that enables quadrupedal robots to perform soccer goalkeeping tasks in the real world. Soccer goalkeeping with quadrupeds is a challenging problem, that combines highly dynamic locomotion with precise and fast non-prehensile object (ball) manipulati…

Cited by 50SourceScholar
2023

GANHead: Towards Generative Animatable Neural Head Avatars

CVPR 2023poster

To bring digital avatars into people's lives, it is highly demanded to efficiently generate complete, realistic, and animatable head avatars. This task is challenging, and it is difficult for existing methods to satisfy all the requirements at once. To achieve these goals, we propose GANHead (Genera…

Cited by 20SourcePDFScholar
2021

Looking Here or There? Gaze Following in 360-Degree Images

ICCV 2021poster

Gaze following, i.e., detecting the gaze target of a human subject, in 2D images has become an active topic in computer vision. However, it usually suffers from the out of frame issue due to the limited field-of-view (FoV) of 2D images. In this paper, we introduce a novel task, gaze following in 360…

Cited by 23PDFScholar