← Search

Hai Ci

19 accepted papers

2026

CyC3D: Fine-grained Controllable 3D Generation via Cycle Consistency Regularization

AAAI 2026technical

Despite the remarkable progress of 3D generation, achieving controllability, i.e., ensuring consistency between generated 3D content and input conditions like edge and depth, remains a significant challenge. Existing methods often struggle to maintain accurate alignment, leading to noticeable discre

Cited by 0SourcePDFScholar
2026

LumosX: Relate Any Identities with Their Attributes for Personalized Video Generation

ICLR 2026poster

Recent advances in diffusion models have significantly improved text-to-video generation, enabling personalized content creation with fine-grained control over both foreground and background elements. However, precise face–attribute alignment across subjects remains challenging, as existing methods…

Cited by 0SourcecodeScholar
2026

OptMark: Robust Multi-bit Diffusion Watermarking via Inference Time Optimization

AAAI 2026technical

Watermarking diffusion-generated images is crucial for copyright protection and user tracking. However, current diffusion watermarking methods face significant limitations: zero-bit watermarking systems lack the capacity for large-scale user tracking, while multi-bit methods are highly sensitive to

Cited by 0SourcePDFScholar
2026

RobotSeg: A Model and Dataset for Segmenting Robots in Image and Video

CVPR 2026

Accurate robot segmentation is a fundamental capability for robotic perception. It enables precise visual servoing for VLA systems, scalable robot-centric data augmentation, accurate real-to-sim transfer, and reliable safety monitoring in dynamic human-robot environments. Despite the strong capabili

Cited by 0SourcecodeScholar
2025

FreeCloth: Free-form Generation Enhances Challenging Clothed Human Modeling

CVPR 2025highlight

Achieving realistic animated human avatars requires accurate modeling of pose-dependent clothing deformations. Existing learning-based methods heavily rely on the Linear Blend Skinning (LBS) of minimally-clothed human models like SMPL to model deformation. However, they struggle to handle loose clot…

Cited by 0SourcePDFScholar
2025

IDProtector: An Adversarial Noise Encoder to Protect Against ID-Preserving Image Generation

CVPR 2025poster

Recently, zero-shot methods like InstantID have revolutionized identity-preserving generation. Unlike multi-image finetuning approaches such as DreamBooth, these zero-shot methods leverage powerful facial encoders to extract identity information from a single portrait photo, enabling efficient ident…

2025

Image Watermarks are Removable using Controllable Regeneration from Clean Noise

ICLR 2025poster

Image watermark techniques provide an effective way to assert ownership, deter misuse, and trace content sources, which has become increasingly essential in the era of large generative models. A critical attribute of watermark techniques is their robustness against various manipulations. In this pap…

2025

Impossible Videos

ICML 2025poster

Synthetic videos nowadays is widely used to complement data scarcity and diversity of real-world videos. Current synthetic datasets primarily replicate real-world scenarios, leaving impossible, counterfactual and anti-reality video concepts underexplored. This work aims to answer two questions: 1) C…

Cited by 1SourcePDFScholar
2025

UnrealZoo: Enriching Photo-realistic Virtual Worlds for Embodied AI

ICCV 2025poster

We introduce UnrealZoo, a collection of over 100 photo-realistic 3D virtual worlds built on Unreal Engine, designed to reflect the complexity and variability of open-world environments. We also provide a rich variety of playable entities, including humans, animals, robots, and vehicles for embodied…

2025

WMAdapter: Adding WaterMark Control to Latent Diffusion Models

ICML 2025poster

Watermarking is essential for protecting the copyright of AI-generated images. We propose WMAdapter, a diffusion model watermark plugin that embeds user-specified watermark information seamlessly during the diffusion generation process. Unlike previous methods that modify diffusion modules to incorp…

Cited by 14SourcePDFScholar
2024

Empowering Embodied Visual Tracking with Visual Foundation Models and Offline RL

ECCV 2024poster

"Embodied visual tracking is to follow a target object in dynamic 3D environments using an agent’s egocentric vision. This is a vital and challenging skill for embodied agents. However, existing methods suffer from inefficient training and poor generalization. In this paper, we propose a novel frame…

Cited by 5SourcePDFScholar
2023

GFPose: Learning 3D Human Pose Prior With Gradient Fields

CVPR 2023poster

Learning 3D human pose prior is essential to human-centered AI. Here, we present GFPose, a versatile framework to model plausible 3D human poses for various applications. At the core of GFPose is a time-dependent score network, which estimates the gradient on each body joint and progressively denois…

2023

Proactive Multi-Camera Collaboration for 3D Human Pose Estimation

ICLR 2023poster

This paper presents a multi-agent reinforcement learning (MARL) scheme for proactive Multi-Camera Collaboration in 3D Human Pose Estimation in dynamic human crowds. Traditional fixed-viewpoint multi-camera solutions for human motion capture (MoCap) are limited in capture space and susceptible to dyn…

Cited by 17SourcePDFScholar
2023

Social Motion Prediction with Cognitive Hierarchies

NeurIPS 2023poster

Humans exhibit a remarkable capacity for anticipating the actions of others and planning their own actions accordingly. In this study, we strive to replicate this ability by addressing the social motion prediction problem. We introduce a new benchmark, a novel formulation, and a cognition-inspired f…

Cited by 9SourcePDFScholar
2021

Context Modeling in 3D Human Pose Estimation: A Unified Perspective

CVPR 2021poster

Estimating 3D human pose from a single image suffers from severe ambiguity since multiple 3D joint configurations may have the same 2D projection. The state-of-the-art methods often rely on context modeling methods such as pictorial structure model (PSM) or graph neural network (GNN) to reduce ambig…

Cited by 97PDFScholar