← Search

Chongyang Xu

8 accepted papers

2026

Action-Geometry Prediction with 3D Geometric Prior for Bimanual Manipulation

CVPR 2026

Bimanual manipulation requires policies that can reason about 3D geometry, anticipate how it evolves under action, and generate smooth, coordinated motions. However, existing methods typically rely on 2D features with limited spatial awareness, or require explicit point clouds that are difficult to

Cited by 0SourcecodeScholar
2026

Causality-inspired Federated Learning for Dynamic Spatio-Temporal Graphs

AAAI 2026technical

Federated Graph Learning (FGL) has emerged as a powerful paradigm for decentralized training of graph neural networks while preserving data privacy. However, existing FGL methods are predominantly designed for static graphs and rely on parameter averaging or distribution alignment, which implicitly

Cited by 0SourcePDFScholar
2026

HeRO: Hierarchical 3D Semantic Representation for Pose-Aware Object Manipulation

ICRA 2026poster

Imitation learning for robotic manipulation has progressed from 2D image policies to 3D representations that explicitly encode geometry. Yet purely geometric policies often lack explicit part-level semantics, which are critical for pose-aware manipulation (e.g., distinguishing a shoe's toe from heel…

2026

Occluded Human Body Capture with Frequency Domain Denoising Prior

CVPR 2026

Monocular human motion capture in occlusion scenarios presents significant challenges. Although a few works have explicitly considered the occlusion problem, image-based methods are unreliable due to the lack of temporal constraints while video-based approaches cannot gain sufficient knowledge from

Cited by 0SourcecodeScholar
2026

Stability-Driven Motion Generation for Object-Guided Human-Human Co-Manipulation

CVPR 2026

Co-manipulation requires multiple humans to synchronize their motions with a shared object while ensuring reasonable interactions, maintaining natural poses, and preserving stable states. However, most existing motion generation approaches are designed for single-character scenarios or fail to accou

Cited by 0SourceScholar
2025

Cache Saver: A Modular Framework for Efficient, Affordable, and Reproducible LLM Inference

EMNLP 2025

Inference constitutes the majority of costs throughout the lifecycle of a large language model (LLM). While numerous LLM inference engines focusing primarily on low-level optimizations have been developed, there is a scarcity of non-intrusive client-side frameworks that perform high-level optimizati

2025

Reconstructing Close Human Interaction with Appearance and Proxemics Reasoning

CVPR 2025poster

Due to visual ambiguities and inter-person occlusions, existing human pose estimation methods cannot recover plausible close interactions from in-the-wild videos. Even state-of-the-art large foundation models (e.g., SAM) cannot accurately distinguish human semantics in such challenging scenarios. In…

Cited by 0SourcePDFScholar
2024

Closely Interactive Human Reconstruction with Proxemics and Physics-Guided Adaption

CVPR 2024poster

Existing multi-person human reconstruction approaches mainly focus on recovering accurate poses or avoiding penetration but overlook the modeling of close interactions. In this work we tackle the task of reconstructing closely interactive humans from a monocular video. The main challenge of this tas…