← Search

Junda Cheng

12 accepted papers

2026

GemDepth: Geometry-Embedded Features for 3D-Consistent Video Depth

ICML 2026poster

Video depth estimation extends monocular prediction into the temporal domain to ensure coherence. However, existing methods often suffer from spatial blurring in fine-detail regions and temporal inconsistencies. We argue that current approaches, which primarily rely on temporal smoothing via Transfo…

Cited by 0SourceScholar
2026

PromptStereo: Zero-Shot Stereo Matching via Structure and Motion Prompts

CVPR 2026

Modern stereo matching methods have leveraged monocular depth foundation models to achieve superior zero-shot generalization performance. However, most existing methods primarily focus on extracting robust features for cost volume construction or disparity initialization. At the same time, the itera

Cited by 0SourcecodeScholar
2025

ACP-MVS: Efficient Multi-View Stereo with Attention-based Context Perception

IROS 2025

The core of Multi-View Stereo (MVS) is to find corresponding pixels in neighboring images. However, due to challenging regions in input images such as untextured areas, repetitive patterns, or reflective surfaces, existing methods struggle to find precise pixel correspondence therein, resulting in i

Cited by 0SourcecodeScholar
2025

BANet: Bilateral Aggregation Network for Mobile Stereo Matching

ICCV 2025poster

State-of-the-art stereo matching methods typically use costly 3D convolutions to aggregate a full cost volume, but their computational demands make mobile deployment challenging. Directly applying 2D convolutions for cost aggregation often results in edge blurring, detail loss, and mismatches in tex…

2025

Leveraging Consistent Spatio-Temporal Correspondence for Robust Visual Odometry

AAAI 2025technical

Recent approaches to VO have significantly improved performance by using deep networks to predict optical flow between video frames. However, existing methods still suffer from noisy and inconsistent flow matching, making it difficult to handle challenging scenarios and long-sequence estimation.To o…

2025

LiDAR-Inertial Odometry in Dynamic Driving Scenarios using Label Consistency Detection

IROS 2025

In this paper, a LiDAR-inertial odometry (LIO) method that eliminates the influence of moving objects in dynamic driving scenarios is proposed. This method constructs binarized labels for 3D points of current sweep, and utilizes the label difference between each point and its surrounding points in g

Cited by 2SourceScholar
2025

MonSter: Marry Monodepth to Stereo Unleashes Power

CVPR 2025highlight

Stereo matching recovers depth from image correspondences. Existing methods struggle to handle ill-posed regions with limited matching cues, such as occlusions and textureless areas. To address this, we propose MonSter, a novel method that leverages the complementary strengths of monocular depth est…

2025

PriOr-Flow: Enhancing Primitive Panoramic Optical Flow with Orthogonal View

ICCV 2025poster

Panoramic optical flow enables a comprehensive understanding of temporal dynamics across wide fields of view. However, severe distortions caused by sphere-to-plane projections, such as the equirectangular projection (ERP), significantly degrade the performance of conventional perspective-based optic…

2024

Adaptive Fusion of Single-View and Multi-View Depth for Autonomous Driving

CVPR 2024poster

Multi-view depth estimation has achieved impressive performance over various benchmarks. However almost all current multi-view systems rely on given ideal camera poses which are unavailable in many real-world scenarios such as autonomous driving. In this work we propose a new robustness benchmark to…

2024

L-MAGIC: Language Model Assisted Generation of Images with Coherence

CVPR 2024poster

In the current era of generative AI breakthroughs generating panoramic scenes from a single input image remains a key challenge. Most existing methods use diffusion-based iterative or simultaneous multi-view inpainting. However the lack of global scene layout priors leads to subpar outputs with dupl…

2022

Attention Concatenation Volume for Accurate and Efficient Stereo Matching

CVPR 2022poster

Stereo matching is a fundamental building block for many vision and robotics applications. An informative and concise cost volume representation is vital for stereo matching of high accuracy and efficiency. In this paper, we present a novel cost volume construction method which generates attention w…

Cited by 283PDFcodeScholar