← Search

Liang-Yan Gui

18 accepted papers

2026

HandX: Scaling Bimanual Motion and Interaction Generation

CVPR 2026

Synthesizing human motion has advanced rapidly, yet realistic hand motion and bimanual interaction remain underexplored. Whole-body models often miss the fine-grained cues that drive dexterous behavior, finger articulation, contact timing, and inter-hand coordination, and existing resources lack hig

Cited by 0SourcecodeScholar
2026

InterPrior: Scaling Generative Control for Physics-Based Human-Object Interactions

CVPR 2026

Humans rarely plan whole-body interactions with objects at the level of explicit whole-body movements. High-level intentions, such as affordance, define the goal, while coordinated balance, contact, and manipulation can emerge naturally from underlying physical and motor priors. Scaling such priors

Cited by 0SourceScholar
2026

LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight

CVPR 2026

To act in the world, a model must name what it sees and know where it is in 3D. Today's vision-language models excel at open-ended 2D description and grounding, yet multi-object 3D detection remains largely missing from the VLM toolbox. We present LocateAnything3D, a VLM-native recipe that casts 3D

Cited by 0SourcecodeScholar
2026

Think in Latent, Explain in Language: Self-Explainable Latent Reasoning

ICML 2026poster

Latent reasoning has emerged as a powerful alternative to text-based Chain-of-Thought (CoT), offering significant gains in computational efficiency by compressing verbose reasoning into compact embeddings. However, compressing reasoning into the latent space renders the thinking opaque, hindering it…

Cited by 0SourceScholar
2025

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought

CVPR 2025poster

Recent advances in multimodal large language models (MLLMs) have demonstrated remarkable capabilities in vision-language tasks, yet they often struggle with vision-centric scenarios where precise visual focus is needed for accurate reasoning. In this paper, we introduce Argus to address these limita…

Cited by 0SourcePDFScholar
2025

Floating No More: Object-Ground Reconstruction from a Single Image

CVPR 2025poster

Recent advancements in 3D object reconstruction from single images have primarily focused on improving the accuracy of object shapes. Yet, these techniques often fail to accurately capture the inter-relation between the object, ground, and camera. As a result, the reconstructed objects often appear…

Cited by 3SourcePDFScholar
2025

InterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction Generation

CVPR 2025poster

While large-scale human motion capture datasets have advanced human motion generation, modeling and generating dynamic 3D human-object interactions (HOIs) remain challenging due to dataset limitations. Existing datasets often lack extensive, high-quality motion and annotation and exhibit artifacts s…

Cited by 2SourcePDFScholar
2025

InterMimic: Towards Universal Whole-Body Control for Physics-Based Human-Object Interactions

CVPR 2025highlight

Achieving realistic simulations of humans interacting with a wide range of objects has long been a fundamental goal. Extending physics-based motion imitation to complex human-object interactions (HOIs) is challenging due to intricate human-object coupling, variability in object geometries, and artif…

2025

Refer to Any Segmentation Mask Group With Vision-Language Prompts

ICCV 2025poster

Recent image segmentation models have advanced to segment images into high-quality masks for visual entities, and yet they cannot provide comprehensive semantic understanding for complex queries based on both language and vision. This limitation reduces their effectiveness in applications that requi…

2023

Contrastive Mean Teacher for Domain Adaptive Object Detectors

CVPR 2023poster

Object detectors often suffer from the domain gap between training (source domain) and real-world applications (target domain). Mean-teacher self-training is a powerful paradigm in unsupervised domain adaptation for object detection, but it struggles with low-quality pseudo-labels. In this work, we…

2023

InterDiff: Generating 3D Human-Object Interactions with Physics-Informed Diffusion

ICCV 2023poster

This paper addresses a novel task of anticipating 3D human-object interactions (HOIs). Most existing research on HOI synthesis lacks comprehensive whole-body interactions with dynamic objects, e.g., often limited to manipulating small or static objects. Our task is significantly more challenging, as…

Cited by 116PDFcodeScholar
2023

SDFusion: Multimodal 3D Shape Completion, Reconstruction, and Generation

CVPR 2023poster

In this work, we present a novel framework built to simplify 3D asset generation for amateur users. To enable interactive generation, our method supports a variety of input modalities that can be easily provided by a human, including images, texts, partially observed shapes and combinations of these…

2022

Diverse Human Motion Prediction Guided by Multi-level Spatial-Temporal Anchors

ECCV 2022poster

"Predicting diverse human motions given a sequence of historical poses has received increasing attention. Despite rapid progress, existing work captures the multi-modal nature of human motions primarily through likelihood-based sampling, where the mode collapse has been widely observed. In this pape…

2018

Adversarial Geometry-Aware Human Motion Prediction

ECCV 2018poster

We explore an approach to forecasting human motion in a few milliseconds given an input 3D skeleton sequence based on a recurrent encoder-decoder framework. Current approaches suffer from the problem of prediction discontinuities and may fail to predict human-like motion in longer time horizons due…

Cited by 317SourcePDFScholar
2018

Few-Shot Human Motion Prediction via Meta-Learning

ECCV 2018poster

Human motion prediction, forecasting human motion in a few milliseconds conditioning on a historical 3D skeleton sequence, is a long-standing problem in computer vision and robotic vision. Existing forecasting algorithms rely on extensive annotated motion capture data and are brittle to novel action…

Cited by 155SourcePDFScholar
2018

Teaching Robots to Predict Human Motion

IROS 2018poster

Teaching a robot to predict and mimic how a human moves or acts in the near future by observing a series of historical human movements is a crucial first step in human-robot interaction and collaboration. In this paper, we instrument a robot with such a prediction ability by leveraging recent deep l…

Cited by 137SourceScholar