← Search

Sitong Zhuang

2 accepted papers

2026

Bootstrap Dynamic-Aware 3D Visual Representation for Scalable Robot Learning

CVPR 2026

Despite strong results on recognition and segmentation, current 3D visual pre-training methods often underperform on robotic manipulation. We attribute this gap to two factors: the lack of state-action-state dynamics modeling and the unnecessary redundancy of explicit geometric reconstruction. We in

Cited by 0SourcecodeScholar
2025

Temporal Action Detection Model Compression by Progressive Block Drop

CVPR 2025poster

Temporal action detection (TAD) aims to identify and localize action instances in untrimmed videos, which is essential for various video understanding tasks. However, recent improvements in model performance, driven by larger feature extractors and datasets, have led to increased computational deman…

Cited by 0SourcePDFScholar