← Search

Jian Shi

10 accepted papers

2026

Any Resolution Any Geometry: From Multi-View To Multi-Patch

CVPR 2026

Joint estimation of surface normals and depth is essential for holistic 3D scene understanding, yet high-resolution prediction remains difficult due to the trade-off between preserving fine local detail and maintaining global consistency. To address this challenge, we propose the Ultra Resolution Ge

Cited by 0SourcecodeScholar
2026

Exploring Reliable Spatiotemporal Dependencies for Efficient Visual Tracking

AAAI 2026technical

Recent advances in transformer-based lightweight object tracking have established new standards across benchmarks, leveraging the global receptive field and powerful feature extraction capabilities of attention mechanisms. Despite these achievements, existing methods universally employ sparse sampli

Cited by 0SourcePDFScholar
2025

Amodal Depth Anything: Amodal Depth Estimation in the Wild

ICCV 2025poster

Amodal depth estimation aims to predict the depth of occluded (invisible) parts of objects in a scene. This task addresses the question of whether models can effectively perceive the geometry of occluded regions based on visible cues. Prior methods primarily rely on synthetic datasets and focus on m…

Cited by 0SourcePDFScholar
2025

HOGSA: Bimanual Hand-Object Interaction Understanding with 3D Gaussian Splatting Based Data Augmentation

AAAI 2025technical

Understanding of bimanual hand-object interaction plays an important role in robotics and virtual reality. However, due to significant occlusions between hands and object as well as the high degree-of-freedom motions, it is challenging to collect and annotate a high-quality, large-scale dataset, whi…

Cited by 0SourcePDFScholar
2025

VoxelKP: A Voxel-based Network Architecture for Human Keypoint Estimation in LiDAR Data

ICCV 2025poster

We present VoxelKP, a novel fully sparse network architecture tailored for human keypoint estimation in LiDAR data. The key challenge is that objects are distributed sparsely in 3D space, while human keypoint detection requires detailed local information wherever humans are present. First, we introd…