← Search

Tze Ho Elden Tse

15 accepted papers

2026

HumanBA: Human-Aware Bundle Adjustment via Global Human-Camera Decoupling

CVPR 2026

Recovering global human and camera motion from monocular video is essential for world-coordinate human reconstruction but remains challenging due to entangled motions in image space. Traditional SLAM methods estimate monocular camera motion but fail in scenes dominated by foreground objects such as

Cited by 0SourcecodeScholar
2026

Learning Scene Coordinate Reconstruction from Unposed Images via Pose Graph Optimization

CVPR 2026

Learning-based structure-from-motion methods such as ACE-Zero have demonstrated strong performance in estimating camera poses and scene coordinates from unordered image collections without requiring ground truth supervision. However, the lack of global and multi-view consistency constraints in ACE-Z

Cited by 0SourceScholar
2026

TIGeR: Text-Instructed Generation and Refinement for Template-Free Hand-Object Interaction

ICRA 2026poster

Pre-defined 3D object templates are widely used in 3D reconstruction of hand-object interactions. However, they often require substantial manual efforts to capture or source, and inherently restrict the adaptability of models to unconstrained interaction scenarios, e.g., heavily-occluded objects. To…

2025

A Constrained Optimization Approach for Gaussian Splatting from Coarsely-posed Images and Noisy Lidar Point Clouds

ICCV 2025poster

3D Gaussian Splatting (3DGS) is a powerful reconstruction technique; however, it requires initialization from accurate camera poses and high-fidelity point clouds. Typically, the initialization is taken from Structure-from-Motion (SfM) algorithms; however, SfM is time-consuming and restricts the app…

Cited by 0SourcePDFScholar
2025

Collaborative Learning for 3D Hand-Object Reconstruction and Compositional Action Recognition from Egocentric RGB Videos Using Superquadrics

AAAI 2025technical

With the availability of egocentric 3D hand-object interaction datasets, there is increasing interest in developing unified models for hand-object pose estimation and action recognition. However, existing methods still struggle to recognise seen actions on unseen objects due to the limitations in re…

Cited by 0SourcePDFScholar
2025

High-Resolution Spatiotemporal Modeling with Global-Local State Space Models for Video-Based Human Pose Estimation

ICCV 2025poster

Modeling high-resolution spatiotemporal representations, including both global dynamic contexts (e.g., holistic human motion tendencies) and local motion details (e.g., high-frequency changes of keypoints), is essential for video-based human pose estimation (VHPE). Current state-of-the-art methods t…

Cited by 0SourcePDFScholar
2025

Humans as Checkerboards: Calibrating Camera Motion Scale for World-Coordinate Human Mesh Recovery

ICCV 2025poster

Accurate camera motion estimation is essential for recovering global human motion in world coordinates from RGB video inputs. While SLAM is widely used for estimating camera trajectory and point cloud, monocular SLAM does so only up to an unknown scale factor. Previous works estimate the scale facto…

2025

Visual Intention Grounding for Egocentric Assistants

ICCV 2025poster

Visual grounding associates textual descriptions with objects in an image. Conventional methods target third-person image inputs and named object queries. In applications such as AI assistants, the perspective shifts -- inputs are egocentric, and objects may be referred to implicitly through needs a…

2024

GeoReF: Geometric Alignment Across Shape Variation for Category-level Object Pose Refinement

CVPR 2024poster

Object pose refinement is essential for robust object pose estimation. Previous work has made significant progress towards instance-level object pose refinement. Yet category-level pose refinement is a more challenging problem due to large shape variations within a category and the discrepancies bet…

Cited by 4SourcePDFScholar
2023

DiffPose: SpatioTemporal Diffusion Model for Video-Based Human Pose Estimation

ICCV 2023poster

Denoising diffusion probabilistic models that were initially proposed for realistic image generation have recently shown success in various perception tasks (e.g., object detection and image segmentation) and are increasingly gaining attention in computer vision. However, extending such models to mu…

Cited by 43PDFScholar
2023

Mutual Information-Based Temporal Difference Learning for Human Pose Estimation in Video

CVPR 2023poster

Temporal modeling is crucial for multi-frame human pose estimation. Most existing methods directly employ optical flow or deformable convolution to predict full-spectrum motion fields, which might incur numerous irrelevant cues, such as a nearby person or background. Without further efforts to excav…

Cited by 24SourcePDFScholar
2023

Spectral Graphormer: Spectral Graph-Based Transformer for Egocentric Two-Hand Reconstruction using Multi-View Color Images

ICCV 2023poster

We propose a novel transformer-based framework that reconstructs two high fidelity hands from multi-view RGB images. Unlike existing hand pose estimation methods, where one typically trains a deep network to regress hand model parameters from single RGB image, we consider a more challenging problem…

Cited by 4PDFScholar
2022

Collaborative Learning for Hand and Object Reconstruction With Attention-Guided Graph Convolution

CVPR 2022poster

Estimating the pose and shape of hands and objects under interaction finds numerous applications including augmented and virtual reality. Existing approaches for hand and object reconstruction require explicitly defined physical constraints and known objects, which limits its application domains. Ou…

Cited by 46PDFScholar
2022

S2Contact: Graph-Based Network for 3D Hand-Object Contact Estimation with Semi-Supervised Learning

ECCV 2022poster

"Being able to reason about the physical contacts between hands and objects is crucial in understanding hand-object manipulation. However, despite the efforts in accurate 3D annotations in hand and object datasets, there still exist gaps in 3D hand and object reconstructions. Recent works leverage c…

Cited by 21SourcePDFScholar
2022

TP-AE: Temporally Primed 6D Object Pose Tracking with Auto-Encoders

ICRA 2022poster

Fast and accurate tracking of an object's motion is one of the key functionalities of a robotic system for achieving reliable interaction with the environment. This paper focuses on the instance-level six-dimensional (6D) pose tracking problem with a symmetric and textureless object under occlusion.…

Cited by 9SourceScholar