← Search

Liqian Ma

17 accepted papers

2025

Bootstraping Clustering of Gaussians for View-consistent 3D Scene Understanding

AAAI 2025technical

Injecting semantics into 3D Gaussian Splatting (3DGS) has recently garnered significant attention. While current approaches typically distill 3D semantic features from 2D foundational models (e.g., CLIP and SAM) to facilitate novel view segmentation and semantic understanding, their heavy reliance o…

2025

CCL-LGS: Contrastive Codebook Learning for 3D Language Gaussian Splatting

ICCV 2025poster

Recent advances in 3D reconstruction techniques and vision-language models have fueled significant progress in 3D semantic understanding, a capability critical to robotics, autonomous driving, and virtual/augmented reality. However, methods that rely on 2D priors are prone to a critical challenge: c…

2025

Generalizable Domain Adaptation for Sim-and-Real Policy Co-Training

NeurIPS 2025poster

Behavior cloning has shown promise for robot manipulation, but real-world demonstrations are costly to acquire at scale. While simulated data offers a scalable alternative, particularly with advances in automated demonstration generation, transferring policies to the real world is hampered by variou…

Cited by 0SourceScholar
2025

VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical Prior

ICCV 2025accepted

Video diffusion models (VDMs) have advanced significantly in recent years, enabling the generation of highly realistic videos and drawing the attention of the community in their potential as world simulators. However, despite their capabilities, VDMs often fail to produce physically plausible videos…

2024

Meta-Control: Automatic Model-based Control Synthesis for Heterogeneous Robot Skills

CoRL 2024poster

The requirements for real-world manipulation tasks are diverse and often conflicting; some tasks require precise motion while others require force compliance; some tasks require avoidance of certain regions while others require convergence to certain states. Satisfying these varied requirements with…

Cited by 4SourceScholar
2024

PTUS: Photo-Realistic Talking Upper-Body Synthesis via 3D-Aware Motion Decomposition Warping

AAAI 2024technical

Talking upper-body synthesis is a promising task due to its versatile potential for video creation and consists of animating the body and face from a source image with the motion from a given driving video. However, prior synthesis approaches fall short in addressing this task and have been either l…

2023

Cloth2Body: Generating 3D Human Body Mesh from 2D Clothing

ICCV 2023poster

In this paper, we define and study a new Cloth2Body problem which has a goal of generating 3d human body meshes from a 2D clothing image. Unlike the existing human mesh recovery problem, Cloth2Body needs to address new and emerging challenges raised by the partial observation of the input and the hi…

Cited by 4PDFcodeScholar
2023

GM-NeRF: Learning Generalizable Model-Based Neural Radiance Fields From Multi-View Images

CVPR 2023poster

In this work, we focus on synthesizing high-fidelity novel view images for arbitrary human performers, given a set of sparse multi-view images. It is a challenging task due to the large variation among articulated body poses and heavy self-occlusions. To alleviate this, we introduce an effective gen…

2023

Sim2Real2: Actively Building Explicit Physics Model for Precise Articulated Object Manipulation

ICRA 2023poster

Accurately manipulating articulated objects is a challenging yet important task for real robot applications. In this paper, we present a novel framework called Sim2Real2 to enable the robot to manipulate an unseen articulated object to the desired state precisely in the real world with no human demo…

Cited by 14SourcecodeScholar
2022

UNIF: United Neural Implicit Functions for Clothed Human Reconstruction and Animation

ECCV 2022poster

"We propose united implicit functions (UNIF), a part-based method for clothed human reconstruction and animation with raw scans and skeletons as the input. Previous part-based methods for human reconstruction rely on ground-truth part labels from SMPL and thus are limited to minimal-clothed humans.…

2021

FoV-Net: Field-of-View Extrapolation Using Self-Attention and Uncertainty

RA-L 2021

The ability to make educated predictions about their surroundings, and associate them with certain confidence, is important for intelligent systems, like autonomous vehicles and robots. It allows them to plan early and decide accordingly. Motivated by this observation, in this letter we utilize info

Cited by 7SourcecodeScholar
2020

Unselfie: Translating Selfies to Neutral-pose Portraits in the Wild

ECCV 2020poster

Due to the ubiquity of smartphones, it is popular to take photos of one's self, or ""selfies."" Such photos are convenient to take, because they do not require specialized equipment or a third-party photographer. However, in selfies, constraints such as human arm length often make the body pose look…

Cited by 14SourcePDFScholar
2019

Exemplar Guided Unsupervised Image-to-Image Translation with Semantic Consistency

ICLR 2019poster

Image-to-image translation has recently received significant attention due to advances in deep learning. Most works focus on learning either a one-to-one mapping in an unsupervised way or a many-to-many mapping in a supervised way. However, a more practical setting is many-to-many mapping in an unsu…

Cited by 165SourcePDFScholar
2018

Disentangled Person Image Generation

CVPR 2018poster

Generating novel, yet realistic, images of persons is a challenging task due to the complex interplay between the different image factors, such as the foreground, background and pose information. In this work, we aim at generating such images based on a novel, two-stage reconstruction pipeline that…

Cited by 541SourcePDFScholar
2018

Natural and Effective Obfuscation by Head Inpainting

CVPR 2018poster

As more and more personal photos are shared online, being able to obfuscate identities in such photos is becoming a necessity for privacy protection. People have largely resorted to blacking out or blurring head regions, but they result in poor user experience while being surprisingly ineffective ag…

Cited by 266SourcePDFScholar
2017

Pose Guided Person Image Generation

NeurIPS 2017poster

This paper proposes the novel Pose Guided Person Generation Network (PG$^2$) that allows to synthesize person images in arbitrary poses, based on an image of that person and a novel pose. Our generation framework PG$^2$ utilizes the pose information explicitly and consists of two key stages: pose in…

Cited by 1080SourcePDFScholar