← Search

Li Cheng

40 accepted papers

2026

Online Multi-Relational Clustering with Dominant View Mining

AAAI 2026technical

Multi-relational graph clustering aims to uncover complex node interactions by leveraging multiple relational views, yet existing methods often suffer from two key limitations: they assume equal importance across views and decouple representation learning from clustering, both of which hinder overal

Cited by 0SourcePDFScholar
2026

Reliable Clustering Number Estimation for Contrastive Multi-View Clustering

CVPR 2026

In recent years, contrastive multi-view clustering has achieved remarkable performance improvements. However, existing methods still face two key challenges: (1) reliance on a predefined number of clusters k, which is often unknown in real-world scenarios; and (2) contrastive learning might cause re

Cited by 0SourceScholar
2026

VPHO: Joint Visual-Physical Cue Learning and Aggregation for Hand-Object Pose Estimation

AAAI 2026technical

Estimating the 3D poses of hands and objects from a single RGB image is a fundamental yet challenging problem, with broad applications in augmented reality and human-computer interaction. Existing methods largely rely on visual cues alone, often producing results that violate physical constraints su

Cited by 0SourcePDFScholar
2025

Act to See, See to Act: Diffusion-Driven Perception-Action Interplay for Adaptive Policies

NeurIPS 2025poster

Existing imitation learning methods decouple perception and action, which overlooks the causal reciprocity between sensory representations and action execution that humans naturally leverage for adaptive behaviors. To bridge this gap, we introduce Action-Guided Diffusion Policy (DP-AG), a unified re…

Cited by 0SourcecodeScholar
2025

BOOTPLACE: Bootstrapped Object Placement with Detection Transformers

CVPR 2025poster

In this paper, we tackle the copy-paste image-to-image composition problem with a focus on object placement learning. Prior methods have leveraged generative models to reduce the reliance for dense supervision. However, this often limits their capacity to model complex data distributions. Alternativ…

2025

CASAGPT: Cuboid Arrangement and Scene Assembly for Interior Design

CVPR 2025highlight

We present a novel approach for indoor scene synthesis, which learns to arrange decomposed cuboid primitives to represent 3D objects within a scene. Unlike conventional methods that use bounding boxes to determine the placement and scale of 3D objects, our approach leverages cuboids as a straightfor…

Cited by 1SourcePDFScholar
2025

HiPoser: 3D Human Pose Estimation with Hierarchical Shared Learning at Parts-Level Using Inertial Measurement Units

AAAI 2025technical

This paper considers the challenging problem of 3D Human Pose Estimation (HPE) from a sparse set of Inertial Measurement Units (IMUs). Existing efforts typically reconstruct a pose sequence by either directly tackling whole-body motions or focusing on distinctive spatio-temporal features of local bo…

Cited by 0SourcePDFScholar
2025

InterMask: 3D Human Interaction Generation via Collaborative Masked Modeling

ICLR 2025poster

Generating realistic 3D human-human interactions from textual descriptions remains a challenging task. Existing approaches, typically based on diffusion models, often produce results lacking realism and fidelity. In this work, we introduce *InterMask*, a novel framework for generating human interact…

Cited by 5SourcePDFScholar
2025

MotionDreamer: One-to-Many Motion Synthesis with Localized Generative Masked Transformer

ICLR 2025poster

Generative masked transformer have demonstrated remarkable success across various content generation tasks, primarily due to their ability to effectively model large-scale dataset distributions with high consistency. However, in the animation domain, large datasets are not always available. Applying…

Cited by 0SourcePDFScholar
2025

MotionScript: Natural Language Descriptions for Expressive 3D Human Motions

IROS 2025

We introduce MotionScript, a novel framework for generating highly detailed, natural language descriptions of 3D human motions. Unlike existing motion datasets that rely on broad action labels or generic captions, MotionScript provides fine-grained, structured descriptions that capture the full comp

Cited by 29SourcecodeScholar
2024

GSD: View-Guided Gaussian Splatting Diffusion for 3D Reconstruction

ECCV 2024poster

"We present GSD, a diffusion model approach based on Gaussian Splatting (GS) representation for 3D object reconstruction from a single view. Prior works suffer from inconsistent 3D geometry or mediocre rendering quality due to improper representations. We take a step towards resolving these shortcom…

Cited by 7SourcePDFScholar
2024

Generative Human Motion Stylization in Latent Space

ICLR 2024poster

Human motion stylization aims to revise the style of an input motion while keeping its content unaltered. Unlike existing works that operate directly in pose space, we leverage the \textit{latent space} of pretrained autoencoders as a more expressive and robust representation for motion extraction a…

Cited by 13SourcePDFScholar
2024

MoMask: Generative Masked Modeling of 3D Human Motions

CVPR 2024poster

We introduce MoMask a novel masked modeling framework for text-driven 3D human motion generation. In MoMask a hierarchical quantization scheme is employed to represent human motion as multi-layer discrete motion tokens with high-fidelity details. Starting at the base layer with a sequence of motion…

2024

TexGen: Text-Guided 3D Texture Generation with Multi-view Sampling and Resampling

ECCV 2024poster

"Given a 3D mesh, we aim to synthesize 3D textures that correspond to arbitrary textual descriptions. Current methods for generating and assembling textures from sampled views often result in prominent seams or excessive smoothing. To tackle these issues, we present TexGen, a novel multi-view sampli…

2024

Unleashing Multispectral Video's Potential in Semantic Segmentation: A Semi-supervised Viewpoint and New UAV-View Benchmark

NeurIPS 2024poster

Thanks to the rapid progress in RGB & thermal imaging, also known as multispectral imaging, the task of multispectral video semantic segmentation, or MVSS in short, has recently drawn significant attentions. Noticeably, it offers new opportunities in improving segmentation performance under unfavora…

Cited by 0SourcePDFScholar
2023

Decorate3D: Text-Driven High-Quality Texture Generation for Mesh Decoration in the Wild

NeurIPS 2023poster

This paper presents Decorate3D, a versatile and user-friendly method for the creation and editing of 3D objects using images. Decorate3D models a real-world object of interest by neural radiance field (NeRF) and decomposes the NeRF representation into an explicit mesh representation, a view-dependen…

2023

Multispectral Video Semantic Segmentation: A Benchmark Dataset and Baseline

CVPR 2023poster

Robust and reliable semantic segmentation in complex scenes is crucial for many real-life applications such as autonomous safe driving and nighttime rescue. In most approaches, it is typical to make use of RGB images as input. They however work well only in preferred weather conditions; when facing…

2022

Contrastive Learning for Unsupervised Video Highlight Detection

CVPR 2022poster

Video highlight detection can greatly simplify video browsing, potentially paving the way for a wide range of applications. Existing efforts are mostly fully-supervised, requiring humans to manually identify and label the interesting moments (called highlights) in a video. Recent weakly supervised m…

Cited by 53PDFcodeScholar
2022

Exploring Denoised Cross-Video Contrast for Weakly-Supervised Temporal Action Localization

CVPR 2022poster

Weakly-supervised temporal action localization aims to localize actions in untrimmed videos with only video-level labels. Most existing methods address this problem with a "localization-by-classification" pipeline that localizes action regions based on snippet-wise classification sequences. Snippet-…

Cited by 75PDFcodeScholar
2022

Generating Diverse and Natural 3D Human Motions From Text

CVPR 2022poster

Automated generation of 3D human motions from text is a challenging problem. The generated motions are expected to be sufficiently diverse to explore the text-grounded motion space, and more importantly, accurately depicting the content in prescribed text descriptions. Here we tackle this problem wi…

Cited by 615PDFcodeScholar
2022

Object Wake-Up: 3D Object Rigging from a Single Image

ECCV 2022poster

"Given a single chair image, could we wake it up by reconstructing its 3D shape and skeleton, as well as animating its plausible articulations and motions, similar to that of human modeling? It is a new problem that not only goes beyond image-based object reconstruction but also involves articulated…

Cited by 7SourcePDFScholar
2022

Promoting Saliency From Depth: Deep Unsupervised RGB-D Saliency Detection

ICLR 2022poster

Growing interests in RGB-D salient object detection (RGB-D SOD) have been witnessed in recent years, owing partly to the popularity of depth sensors and the rapid progress of deep learning techniques. Unfortunately, existing RGB-D SOD methods typically demand large quantity of training images being…

2022

TM2T: Stochastic and Tokenized Modeling for the Reciprocal Generation of 3D Human Motions and Texts

ECCV 2022poster

"Inspired by the strong ties between vision and language, the two intimate human sensing and communication modalities, our paper aims to explore the generation of 3D human full-body motions from texts, as well as its reciprocal task, shorthanded for text2motion and motion2text, respectively. To tack…

2021

Automated Generation of Accurate & Fluent Medical X-ray Reports

EMNLP 2021main

Our paper aims to automate the generation of medical reports from chest X-ray image inputs, a critical yet time-consuming task for radiologists. Existing medical report generation efforts emphasize producing human-readable reports, yet the generated text may not be well aligned to the clinical facts…

2021

EventHPE: Event-Based 3D Human Pose and Shape Estimation

ICCV 2021poster

Event camera is an emerging imaging sensor for capturing dynamics of moving objects as events, which motivates our work in estimating 3D human pose and shape from the event signals. Events, on the other hand, have their unique challenges: rather than capturing static body postures, the event signals…

Cited by 58PDFcodeScholar
2021

Fden: Mining Effective Information of Features in Detecting Network Anomalies

ICASSP 2021accepted

Network anomaly detection is important for detecting and reacting to the presence of network attacks. In this paper, we propose a novel method to effectively leverage the features in detecting network anomalies, named FDEn, consisting of flow-based Feature Derivation (FD) and prior knowledge incorpo…

Cited by 0SourceScholar
2021

Joint Semantic Mining for Weakly Supervised RGB-D Salient Object Detection

NeurIPS 2021poster

Training saliency detection models with weak supervisions, e.g., image-level tags or captions, is appealing as it removes the costly demand of per-pixel annotations. Despite the rapid progress of RGB-D saliency detection in fully-supervised setting, it however remains an unexplored territory when on…

2021

Learning Calibrated Medical Image Segmentation via Multi-Rater Agreement Modeling

CVPR 2021poster

In medical image analysis, it is typical to collect multiple annotations, each from a different clinical expert or rater, in the expectation that possible diagnostic errors could be mitigated. Meanwhile, from the computer vision practitioner viewpoint, it has been a common practice to adopt the grou…

Cited by 188PDFcodeScholar
2021

Neighborhood Consensus Networks for Unsupervised Multi-view Outlier Detection

AAAI 2021technical

Multi-view outlier detection recently attracted rapidly growing attention with the development of multi-view learning. Although promising performance demonstrated, we observe that identifying outliers in multi-view data is still a challenging task due to the complicated characteristics of multi-view…

Cited by 16SourcePDFScholar
2020

3D Human Shape Reconstruction from a Polarization Image

ECCV 2020poster

This paper tackles the problem of estimating 3D body shape of clothed humans from single polarized 2D images, i.e. polarization images. Polarization images are known to be able to capture polarized reflected lights that preserve rich geometric cues of an object, which has motivated its recent applic…

Cited by 56SourcePDFScholar
2019

Automated Cell Patterning System with a Microchip using Dielectrophoresis

ICRA 2019poster

The ability to patterning cells is an important technique to facilitate cell-based assay and characterization. In this paper, an automated cell patterning system was developed for the fabrication of large-scale cell patterns. To resolve the challenge of the limited printable area, the cell-printing…

Cited by 7SourceScholar
2019

Graph-based RGB-D Image Segmentation Using Color-directional-region Merging

ICASSP 2019accepted

Color and depth information provided simultaneously in RGB-D images can be used to segment scenes into disjoint regions. In this paper, a graph-based segmentation method for RGB-D image is proposed, in which an adaptive data-driven combination of color- and normal-variation is presented to construct…

Cited by 0SourceScholar
2019

Towards Natural and Accurate Future Motion Prediction of Humans and Animals

CVPR 2019poster

Anticipating the future motions of 3D articulate objects is challenging due to its non-linear and highly stochastic nature. Current approaches typically represent the skeleton of an articulate object as a set of 3D joints, which unfortunately ignores the relationship between joints, and fails to enc…

Cited by 157PDFScholar