← Search

Yujiao Shi

19 accepted papers

2026

AffordGrasp: Cross-Modal Diffusion for Affordance-Aware Grasp Synthesis

CVPR 2026

Generating human grasping poses that accurately reflect both object geometry and user-specified interaction semantics is essential for natural hand-object interactions in AR/VR and embodied AI. However, existing semantic grasping approaches struggle with the large modality gap between 3D object repr

Cited by 0SourceScholar
2026

CylinderSplat: 3D Gaussian Splatting with Cylindrical Triplanes for Panoramic Novel View Synthesis

ICLR 2026poster

Feed-forward 3D Gaussian Splatting (3DGS) has shown great promise for real-time novel view synthesis, but its application to panoramic imagery remains challenging. Existing methods often rely on multi-view cost volumes for geometric refinement, which struggle to resolve occlusions in sparse-view sce…

Cited by 0SourcecodeScholar
2026

DynOPETs: A Versatile Benchmark for Dynamic Object Pose Estimation and Tracking in Moving Camera Scenarios

ICRA 2026poster

In the realm of object pose estimation, scenarios involving both dynamic objects and moving cameras are prevalent. However, the scarcity of corresponding real-world datasets significantly hinders the development and evaluation of robust pose estimation models. This is largely attributed to the inher…

2026

PanoVGGT: Feed-Forward 3D Reconstruction from Panoramic Imagery

CVPR 2026

Panoramic imagery offers a full 360^\circ field of view and is increasingly common in consumer devices. However, it introduces non-pinhole distortions that challenge joint pose estimation and 3D reconstruction. Existing feed-forward models, built for perspective cameras, generalize poorly to this se

Cited by 0SourcecodeScholar
2026

SatDreamer360: Multiview-Consistent Generation of Ground-Level Scenes from Satellite Imagery

ICLR 2026poster

Generating multiview-consistent $360^\circ$ ground-level scenes from satellite imagery is a challenging task with broad applications in simulation, autonomous navigation, and digital twin cities. Existing approaches primarily focus on synthesizing individual ground-view panoramas, often relying on a…

Cited by 0SourceScholar
2025

BevSplat: Resolving Height Ambiguity via Feature-Based Gaussian Primitives for Weakly-Supervised Cross-View Localization

NeurIPS 2025spotlight

This paper addresses the problem of weakly supervised cross-view localization, where the goal is to estimate the pose of a ground camera relative to a satellite image with noisy ground truth annotations. A common approach to bridge the cross-view domain gap for pose estimation is Bird’s-Eye View (BE…

Cited by 0SourceScholar
2025

Controllable Satellite-to-Street-View Synthesis with Precise Pose Alignment and Zero-Shot Environmental Control

ICLR 2025poster

Generating street-view images from satellite imagery is a challenging task, particularly in maintaining accurate pose alignment and incorporating diverse environmental conditions. While diffusion models have shown promise in generative tasks, their ability to maintain strict pose alignment throughou…

Cited by 0SourcePDFScholar
2025

DynOPETs: A Versatile Benchmark for Dynamic Object Pose Estimation and Tracking in Moving Camera Scenarios

RA-L 2025

In the realm of object pose estimation, scenarios involving both dynamic objects and moving cameras are prevalent. However, the scarcity of corresponding real-world datasets significantly hinders the development and evaluation of robust pose estimation models. This is largely attributed to the inher

Cited by 0SourceScholar
2025

GeoDistill: Geometry-Guided Self-Distillation for Weakly Supervised Cross-View Localization

ICCV 2025poster

Cross-view localization, the task of estimating a camera's 3-degrees-of-freedom (3-DoF) pose by aligning ground-level images with aerial images, is crucial for large-scale outdoor applications like autonomous navigation and augmented reality. Existing methods often rely on fully supervised learning,…

2024

Adapting Fine-Grained Cross-View Localization to Areas without Fine Ground Truth

ECCV 2024poster

"Given a ground-level query image and a geo-referenced aerial image that covers the query’s local surroundings, fine-grained cross-view localization aims to estimate the location of the ground camera inside the aerial image. Recent works have focused on developing advanced networks trained with accu…

2024

Increasing SLAM Pose Accuracy by Ground-to-Satellite Image Registration

ICRA 2024poster

Vision-based localization for autonomous driving has been of great interest among researchers. When a pre-built 3D map is not available, the techniques of visual simultaneous localization and mapping (SLAM) are typically adopted. Due to error accumulation, visual SLAM (vSLAM) usually suffers from lo…

Cited by 6SourcecodeScholar
2023

Boosting 3-DoF Ground-to-Satellite Camera Localization Accuracy via Geometry-Guided Cross-View Transformer

ICCV 2023poster

Image retrieval-based cross-view localization methods often lead to very coarse camera pose estimation, due to the limited sampling density of the database satellite images. In this paper, we propose a method to increase the accuracy of a ground camera's location and orientation by estimating the re…

Cited by 35PDFcodeScholar
2023

Learning Dense Flow Field for Highly-accurate Cross-view Camera Localization

NeurIPS 2023poster

This paper addresses the problem of estimating the 3-DoF camera pose for a ground-level image with respect to a satellite image that encompasses the local surroundings. We propose a novel end-to-end approach that leverages the learning of dense pixel-wise flow fields in pairs of ground and satellite…

Cited by 9SourcePDFScholar
2022

Beyond Cross-View Image Retrieval: Highly Accurate Vehicle Localization Using Satellite Image

CVPR 2022poster

This paper addresses the problem of vehicle-mounted camera localization by matching a ground-level image with an overhead-view satellite map. Existing methods often treat this problem as cross-view image retrieval, and use learned deep features to match the ground-level query image to a partition (e…

Cited by 96PDFcodeScholar
2020

Where Am I Looking At? Joint Location and Orientation Estimation by Cross-View Matching

CVPR 2020poster

Cross-view geo-localization is the problem of estimating the position and orientation (latitude, longitude and azimuth angle) of a camera at ground level given a large-scale database of geo-tagged aerial (eg., satellite) images. Existing approaches treat the task as a pure location estimation proble…

Cited by 213PDFcodeScholar
2019

Spatial-Aware Feature Aggregation for Image based Cross-View Geo-Localization

NeurIPS 2019poster

In this paper, we develop a new deep network to explicitly address these inherent differences between ground and aerial views. We observe there exist some approximate domain correspondences between ground and aerial images. Specifically, pixels lying on the same azimuth direction in an aerial image…