← Search

Yanhao Zhang

32 accepted papers

2026

Aligning Cross-View Visual Geometries in LVLMs Through Human-Like Reasoning Learning

AAAI 2026technical

Spatial understanding is a critical capability for LVLMs (Large Vision-Language Models) to advance embodied AI applications. Existing works primarily focus on enhancing spatial understanding within a single frame, i.e., injecting 3D spatial concepts into LVLMs under single coordinate system. However

Cited by 0SourcePDFScholar
2026

DMTrack: Spatio-Temporal Multimodal Tracking Via Dual-Adapter

ICRA 2026poster

In this paper, we explore adapter tuning and introduce a novel dual-adapter architecture for spatio-temporal multimodal tracking, dubbed DMTrack. The key of our DMTrack lies in two simple yet effective modules, including a spatio-temporal modality adapter (STMA) and a progressive modality complement…

2026

Non-Rigid Structure-From-Motion Via Differential Geometry with Recoverable Conformal Scale

ICRA 2026poster

Non-rigid structure-from-motion (NRSfM), a promising technique for addressing the mapping challenges in monocular visual deformable simultaneous localization and mapping (SLAM), has attracted growing attention. We introduce a novel method, called Con-NRSfM, for NRSfM under conformal deformations, en…

2026

OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning

CVPR 2026

Streaming video reasoning requires models to operate in a setting where history grows without bound while meaningful evidence remains scarce. In such a landscape, relevant signal is like an oasis -- small, critical, and easily lost in a desert of redundancy. Enlarging memory only widens the desert;

Cited by 0SourcecodeScholar
2026

OwlCap: Harmonizing Motion-Detail for Video Captioning via HMD-270K and Caption Set Equivalence Reward

AAAI 2026technical

Video captioning aims to generate comprehensive and coherent descriptions of the video content, contributing to the advancement of both video understanding and generation. However, existing methods often suffer from motion-detail imbalance, as models tend to overemphasize one aspect while neglecting

Cited by 0SourcePDFScholar
2026

StreamRAG: Enhancing Real-Time Video Understanding with Retrieval Augmentation

CVPR 2026

The transition of Retrieval-Augmented Generation (RAG) from offline video analysis to online, streaming scenarios presents a set of critical, unexplored challenges. These include the need for on-the-fly semantic segmentation of continuous video, the inherent tension between low-latency processing an

Cited by 0SourceScholar
2025

Free-MoRef: Instantly Multiplexing Context Perception Capabilities of Video-MLLMs within Single Inference

ICCV 2025poster

Video Multimodal Large Language Models (Video-MLLM) have achieved remarkable advancements in video understanding tasks. However, constrained by the context length limitation in the underlying LLMs, existing Video-MLLMs typically exhibit suboptimal performance on long video scenarios. To understand e…

2025

HEIE: MLLM-Based Hierarchical Explainable AIGC Image Implausibility Evaluator

CVPR 2025poster

AIGC images are prevalent across various fields, yet they frequently suffer from quality issues like artifacts and unnatural textures. Specialized models aim to predict defect region heatmaps but face two primary challenges: (1) lack of explainability, failing to provide reasons and analyses for sub…

2025

InstructHOI: Context-Aware Instruction for Multi-Modal Reasoning in Human-Object Interaction Detection

NeurIPS 2025spotlight

Recently, Large Foundation Models (LFMs), e.g., CLIP and GPT, have significantly advanced the Human-Object Interaction (HOI) detection, due to their superior generalization and transferability. Prior HOI detectors typically employ single- or multi-modal prompts to generate discriminative representat…

Cited by 0SourceScholar
2025

Self-supervised 3D Reconstruction of Tibia and Fibula from Biplanar X-rays

IROS 2025

With the growing number of patients experiencing knee-related conditions, total knee arthroplasty (TKA) has become a common procedure, where a 3D visualisation of the patient’s tibia and fibula is essential for preoperative planning. Traditional imaging techniques, such as computed tomography (CT),

Cited by 0SourcecodeScholar
2024

BARET: Balanced Attention Based Real Image Editing Driven by Target-Text Inversion

AAAI 2024technical

Image editing approaches with diffusion models have been rapidly developed, yet their applicability are subject to requirements such as specific editing types (e.g., foreground or background object editing, style transfer), multiple conditions (e.g., mask, sketch, caption), and time consuming fine-t…

Cited by 5SourcePDFScholar
2024

Increasing SLAM Pose Accuracy by Ground-to-Satellite Image Registration

ICRA 2024poster

Vision-based localization for autonomous driving has been of great interest among researchers. When a pre-built 3D map is not available, the techniques of visual simultaneous localization and mapping (SLAM) are typically adopted. Due to error accumulation, visual SLAM (vSLAM) usually suffers from lo…

Cited by 6SourcecodeScholar
2024

Lightweight High-Resolution Subject Matting in the Real World

ICASSP 2024accepted

Existing saliency object detection (SOD) methods struggle to satisfy fast inference and accurate results simultaneously in high resolution scenes. They are limited by the quality of public datasets and efficient network modules for high-resolution images. To alleviate these issues, we propose to con…

Cited by 0SourceScholar
2024

R$^2$-Gaussian: Rectifying Radiative Gaussian Splatting for Tomographic Reconstruction

NeurIPS 2024poster

3D Gaussian splatting (3DGS) has shown promising results in image rendering and surface reconstruction. However, its potential in volumetric reconstruction tasks, such as X-ray computed tomography, remains under-explored. This paper introduces R$^2$-Gaussian, the first 3DGS-based framework for spars…

2024

View From Above: Orthogonal-View aware Cross-view Localization

CVPR 2024poster

This paper presents a novel aerial-to-ground feature aggregation strategy tailored for the task of cross-view image-based geo-localization. Conventional vision-based methods heavily rely on matching ground-view image features with a pre-recorded image database often through establishing planar homog…

Cited by 5SourcePDFScholar
2024

Zero-shot High-fidelity and Pose-controllable Character Animation

IJCAI 2024poster

Image-to-video (I2V) generation aims to create a video sequence from a single image, which requires high temporal coherence and visual fidelity. However, existing approaches suffer from inconsistency of character appearances and poor preservation of fine details. Moreover, they require a large amoun…

Cited by 4SourcePDFScholar
2023

3D Reconstruction of Tibia and Fibula using One General Model and Two X-ray Images

ICRA 2023poster

The 3D reconstruction of patient specific bone models plays a crucial role in orthopaedic surgery for clinical evaluation, surgical planning and precise implant design or selection. This paper considers the problem of reconstructing a patient-specific 3D tibia and fibula model from only two 2D X-ray…

Cited by 2SourceScholar
2023

GAM: Gradient Attention Module of Optimization for Point Clouds Analysis

AAAI 2023technical

In the point cloud analysis task, the existing local feature aggregation descriptors (LFAD) do not fully utilize the neighborhood information of center points. Previous methods only use the distance information to constrain the local aggregation process, which is easy to be affected by abnormal poin…

2023

Homography Guided Temporal Fusion for Road Line and Marking Segmentation

ICCV 2023poster

Reliable segmentation of road lines and markings is critical to autonomous driving. Our work is motivated by the observations that road lines and markings are (1) frequently occluded in the presence of moving vehicles, shadow, and glare and (2) highly structured with low intra-class shape variance a…

Cited by 5PDFcodeScholar
2023

Learning Audio-Visual Source Localization via False Negative Aware Contrastive Learning

CVPR 2023poster

Self-supervised audio-visual source localization aims to locate sound-source objects in video frames without extra annotations. Recent methods often approach this goal with the help of contrastive learning, which assumes only the audio and visual contents from the same video are positive samples for…

2023

Matting Moments: A Unified Data-Driven Matting Engine for Mobile AIGC in Photo Gallery

IJCAI 2023poster

Image matting is a fundamental technique in visual understanding and has become one of the most significant capabilities in mobile phones. Despite the development of mobile storage and computing power, achieving diverse mobile Artificial Intelligence Generated Content (AIGC) applications remains a g…

Cited by 3SourcePDFScholar
2023

Satellite Image Based Cross-view Localization for Autonomous Vehicle

ICRA 2023poster

Existing spatial localization techniques for au-tonomous vehicles mostly use a pre-built 3D-HD map, often constructed using a survey-grade 3D mapping vehicle, which is not only expensive but also laborious. This paper shows that by using an off-the-shelf high-definition satellite image as a ready-to…

Cited by 25SourceScholar
2023

TALL: Thumbnail Layout for Deepfake Video Detection

ICCV 2023poster

The growing threats of deepfakes to society and cybersecurity have raised enormous public concerns, and increasing efforts have been devoted to this critical topic of deepfake video detection. Existing video methods achieve good performance but are computationally intensive. This paper introduces a…

Cited by 68PDFcodeScholar
2023

View Consistent Purification for Accurate Cross-View Localization

ICCV 2023poster

This paper proposes a fine-grained self-localization method for outdoor robotics that utilizes a flexible number of onboard cameras and readily accessible satellite images. The proposed method addresses limitations in existing cross-view localization methods that struggle to handle noise sources suc…

Cited by 5PDFScholar
2021

Exploring Visual-Audio Composition Alignment Network for Quality Fashion Retrieval in Video

ICASSP 2021accepted

Fashion retrieval in video suffers from the issues of imperfect visual representation and low quality of search results under the E-commercial circumstance. Previous works generally focus on searching the identical images from visual perspective only, but lack of leveraging multi-modal information f…

Cited by 0SourceScholar
2021

Some Research Questions for SLAM in Deformable Environments

IROS 2021poster

SLAM in deformable environments is a very challenging research topic. Some research works have been presented by different research groups in the past few years. However, there are still some challenging research questions remaining unanswered. This paper discusses some of these research questions f…

Cited by 6SourcecodeScholar
2020

Aortic 3D Deformation Reconstruction using 2D X-ray Fluoroscopy and 3D Pre-operative Data for Endovascular Interventions

ICRA 2020poster

Current clinical endovascular interventions rely on 2D guidance for catheter manipulation. Although an aortic 3D surface is available from the pre-operative CT/MRI imaging, it cannot be used directly as a 3D intra-operative guidance since the vessel will deform during the procedure. This paper aims…

Cited by 7SourceScholar
2020

Dense Isometric Non-Rigid Shape-From-Motion Based on Graph Optimization and Edge Selection

RA-L 2020

In this letter, we propose a novel framework for dense isometric non-rigid shape-from-motion (Iso-NRSfM) based on graph topology and edge selection. A weighted undirected graph, of which nodes, edges, and weighted values are respectively the images, the image warps, and the number of the common feat

Cited by 3SourceScholar