← Search

Shaowei Liu

12 accepted papers

2025

MoBA: Mixture of Block Attention for Long-Context LLMs

NeurIPS 2025spotlight

Scaling the effective context length is essential for advancing large language models (LLMs) toward artificial general intelligence (AGI). However, the quadratic increase in computational complexity inherent in traditional attention mechanisms presents a prohibitive overhead. Existing approaches eit…

Cited by 0SourcecodeScholar
2025

PhysGen3D: Crafting a Miniature Interactive World from a Single Image

CVPR 2025poster

Envisioning physically plausible outcomes from a single image requires a deep understanding of the world's dynamics. To address this, we introduce MiniTwin, a novel framework that transforms a single image into an amodal, camera-centric, interactive 3D scene. By combining advanced image-based geomet…

Cited by 3SourcePDFScholar
2025

Ponimator: Unfolding Interactive Pose for Versatile Human-human Interaction Animation

ICCV 2025poster

Close-proximity human-human interactive poses convey rich contextual information about interaction dynamics. Given such poses, humans can intuitively infer the context and anticipate possible past and future dynamics, drawing on strong priors of human behavior. Inspired by this observation, we propo…

2025

Visual Sync: Multi‑Camera Synchronization via Cross‑View Object Motion

NeurIPS 2025poster

Today, people can easily record memorable moments, ranging from concerts, sports events, lectures, family gatherings, and birthday parties with multiple consumer cameras. However, synchronizing these cross‑camera streams remains challenging. Existing methods assume controlled settings, specific targ…

Cited by 0SourceScholar
2024

PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation

ECCV 2024poster

"We present PhysGen, a novel image-to-video generation method that converts a single image and an input condition (, force and torque applied to an object in the image) to produce a realistic, physically plausible, and temporally consistent video. Our key insight is to integrate model-based physical…

2023

Building Rearticulable Models for Arbitrary 3D Objects From 4D Point Clouds

CVPR 2023poster

We build rearticulable models for arbitrary everyday man-made objects containing an arbitrary number of parts that are connected together in arbitrary ways via 1-degree-of-freedom joints. Given point cloud videos of such everyday objects, our method identifies the distinct object parts, what parts a…

2023

ContactGen: Generative Contact Modeling for Grasp Generation

ICCV 2023poster

This paper presents a novel object-centric contact representation ContactGen for hand-object interaction. The ContactGen comprises 3 components: a contact map indicates the contact location, a part map represents the contact hand part, and a direction map tells the contact direction within each part…

Cited by 30PDFcodeScholar
2022

CASA: Category-agnostic Skeletal Animal Reconstruction

NeurIPS 2022accept

Recovering a skeletal shape from a monocular video is a longstanding challenge. Prevailing nonrigid animal reconstruction methods often adopt a control-point driven animation model and optimize bone transforms individually without considering skeletal topology, yielding unsatisfactory shape and arti…

Cited by 33SourcePDFScholar
2022

DexMV: Imitation Learning for Dexterous Manipulation from Human Videos

ECCV 2022poster

"While in computer vision we have made significant progress on understanding hand-object interactions, it is still very challenging for robots to perform complex dexterous manipulation. In this paper, we propose a new platform and pipeline, DexMV (Dexterous Manipulation from Videos), for imitation l…

2022

Joint Hand Motion and Interaction Hotspots Prediction From Egocentric Videos

CVPR 2022poster

We propose to forecast future hand-object interactions given an egocentric video. Instead of predicting action labels or pixels, we directly predict the hand motion trajectory and the future contact points on the next active object (i.e., interaction hotspots). This relatively low-dimensional repres…

Cited by 105PDFcodeScholar
2021

Hand-Object Contact Consistency Reasoning for Human Grasps Generation

ICCV 2021poster

While predicting robot grasps with parallel jaw grippers have been well studied and widely applied in robot manipulation tasks, the study on natural human grasp generation with a multi-finger hand remains a very challenging problem. In this paper, we propose to generate human grasps given a 3D objec…

Cited by 189PDFcodeScholar
2021

Semi-Supervised 3D Hand-Object Poses Estimation With Interactions in Time

CVPR 2021poster

Estimating 3D hand and object pose from a single image is an extremely challenging problem: hands and objects are often self-occluded during interactions, and the 3D annotations are scarce as even humans cannot directly label the ground-truths from a single image perfectly. To tackle these challenge…

Cited by 195PDFcodeScholar