← Search

Lincheng Li

26 accepted papers

2026

DiffPBR: Point-Based Rendering via Spatial-Aware Residual Diffusion

ICLR 2026poster

Neural radiance fields and 3D Gaussian splatting (3DGS) have significantly advanced 3D reconstruction and novel view synthesis (NVS). Yet, achieving high-fidelity and view-consistent renderings directly from point clouds---without costly per-scene optimization---remains a core challenge. In this wor…

Cited by 0SourceScholar
2025

Dynamic Derivation and Elimination: Audio Visual Segmentation with Enhanced Audio Semantics

CVPR 2025poster

Sound-guided object segmentation has drawn considerable attention for its potential to enhance multimodal perception. Previous methods primarily focus on developing advanced architectures to facilitate effective audio-visual interactions, without fully addressing the inherent challenges posed by aud…

2025

EasyCraft: A Robust and Efficient Framework for Automatic Avatar Crafting

CVPR 2025poster

Character customization, or 'face crafting,' is a vital feature in role-playing games (RPGs), enhancing player engagement by enabling the creation of personalized avatars. Existing automated methods often struggle with generalizability across diverse game engines due to their reliance on the interme…

Cited by 0SourcePDFScholar
2025

LDPose: Towards Inclusive Human Pose Estimation for Limb-Deficient Individuals in the Wild

ICCV 2025poster

Human pose estimation aims to predict the location of body keypoints and enable various practical applications. However, existing research focuses solely on individuals with full physical bodies and overlooks those with limb deficiencies. As a result, current pose estimation methods cannot be genera…

Cited by 0SourcePDFScholar
2025

Multimodal Latent Diffusion Model for Complex Sewing Pattern Generation

ICCV 2025poster

Generating sewing patterns in garment design is receiving increasing attention due to its CG-friendly and flexible-editing nature. Previous sewing pattern generation methods have been able to produce exquisite clothing, but struggle to design complex garments with detailed control. To address these…

Cited by 0SourcePDFScholar
2025

Robust Audio-Visual Segmentation via Audio-Guided Visual Convergent Alignment

CVPR 2025poster

Accurately localizing audible objects based on audio-visual cues is the core objective of audio-visual segmentation. Most previous methods emphasize spatial or temporal multi-modal modeling, yet overlook challenges from ambiguous audio-visual correspondences--such as nearby visually similar but acou…

Cited by 0SourcePDFScholar
2024

Benchmarking Audio Visual Segmentation for Long-Untrimmed Videos

CVPR 2024poster

Existing audio-visual segmentation datasets typically focus on short-trimmed videos with only one pixel-map annotation for a per-second video clip. In contrast for untrimmed videos the sound duration start- and end-sounding time positions and visual deformation of audible objects vary significantly.…

2024

EfficientDreamer: High-Fidelity and Robust 3D Creation via Orthogonal-view Diffusion Priors

CVPR 2024poster

While image diffusion models have made significant progress in text-driven 3D content creation they often fail to accurately capture the intended meaning of text prompts especially for view information. This limitation leads to the Janus problem where multi-faced 3D models are generated under the gu…

2024

HIMO: A New Benchmark for Full-Body Human Interacting with Multiple Objects

ECCV 2024poster

"Generating human-object interactions (HOIs) is critical with the tremendous advances of digital avatars. Existing datasets are typically limited to humans interacting with a single object while neglecting the ubiquitous manipulation of multiple objects. Thus, we propose HIMO, a large-scale MoCap da…

Cited by 3SourcePDFScholar
2024

Text-Guided 3D Face Synthesis - From Generation to Editing

CVPR 2024poster

Text-guided 3D face synthesis has achieved remarkable results by leveraging text-to-image (T2I) diffusion models. However most existing works focus solely on the direct generation ignoring the editing restricting them from synthesizing customized 3D faces through iterative adjustments. In this paper…

2023

Diverse 3D Hand Gesture Prediction From Body Dynamics by Bilateral Hand Disentanglement

CVPR 2023poster

Predicting natural and diverse 3D hand gestures from the upper body dynamics is a practical yet challenging task in virtual avatar creation. Previous works usually overlook the asymmetric motions between two hands and generate two hands in a holistic manner, leading to unnatural results. In this wor…

2023

DyGait: Exploiting Dynamic Representations for High-performance Gait Recognition

ICCV 2023poster

Gait recognition is a biometric technology that recognizes the identity of humans through their walking patterns. Compared with other biometric technologies, gait recognition is more difficult to disguise and can be applied to the condition of long-distance without the cooperation of subjects. Thus,…

Cited by 48PDFScholar
2023

Efficient View Path Planning for Autonomous Implicit Reconstruction

ICRA 2023poster

Implicit neural representations have shown promising potential for 3D scene reconstruction. Recent work applies it to autonomous 3D reconstruction by learning information gain for view path planning. Effective as it is, the computation of the information gain is expensive, and compared with that usi…

Cited by 20SourceScholar
2023

FlowFace: Semantic Flow-Guided Shape-Aware Face Swapping

AAAI 2023technical

In this work, we propose a semantic flow-guided two-stage framework for shape-aware face swapping, namely FlowFace. Unlike most previous methods that focus on transferring the source inner facial features but neglect facial contours, our FlowFace can transfer both of them to a target face, thus lead…

2023

Implementation and Optimization of Grasping Learning with Dual-modal Soft Gripper

ICRA 2023poster

Robust and efficient grasping of different objects is still an open problem due to the difficulty of integrating multidisciplinary knowledge such as gripper ontology design, perception, control, and learning. In recent years, learning-based methods have achieved excellent results in grasping various…

Cited by 4SourceScholar
2023

NeFII: Inverse Rendering for Reflectance Decomposition With Near-Field Indirect Illumination

CVPR 2023poster

Inverse rendering methods aim to estimate geometry, materials and illumination from multi-view RGB images. In order to achieve better decomposition, recent approaches attempt to model indirect illuminations reflected from different materials via Spherical Gaussians (SG), which, however, tends to blu…

2023

NeurAR: Neural Uncertainty for Autonomous 3D Reconstruction With Implicit Neural Representations

RA-L 2023

Implicit neural representations have shown compelling results in offline 3D reconstruction and also recently demonstrated the potential for online SLAM systems. However, applying them to autonomous 3D reconstruction, where a robot is required to explore a scene and plan a view path for the reconstru

Cited by 91SourceScholar
2023

Object-Goal Visual Navigation via Effective Exploration of Relations Among Historical Navigation States

CVPR 2023poster

Object-goal visual navigation aims at steering an agent toward an object via a series of moving steps. Previous works mainly focus on learning informative visual representations for navigation, but overlook the impacts of navigation states on the effectiveness and efficiency of navigation. We observ…

Cited by 27SourcePDFScholar
2023

Towards Unbiased Volume Rendering of Neural Implicit Surfaces With Geometry Priors

CVPR 2023poster

Learning surface by neural implicit rendering has been a promising way for multi-view reconstruction in recent years. Existing neural surface reconstruction methods, such as NeuS and VolSDF, can produce reliable meshes from multi-view posed images. Although they build a bridge between volume renderi…

2023

Zero-Shot Text-to-Parameter Translation for Game Character Auto-Creation

CVPR 2023poster

Recent popular Role-Playing Games (RPGs) saw the great success of character auto-creation systems. The bone-driven face model controlled by continuous parameters (like the position of bones) and discrete parameters (like the hairstyles) makes it possible for users to personalize and customize in-gam…

2022

Learning Implicit Body Representations from Double Diffusion Based Neural Radiance Fields

IJCAI 2022poster

In this paper, we present a novel double diffusion based neural radiance field, dubbed DD-NeRF, to reconstruct human body geometry and render the human body appearance in novel views from a sparse set of images. We first propose a double diffusion mechanism to achieve expressive representations of i…

Cited by 10SourcePDFScholar
2022

One-Shot Talking Face Generation from Single-Speaker Audio-Visual Correlation Learning

AAAI 2022technical

Audio-driven one-shot talking face generation methods are usually trained on video resources of various persons. However, their created videos often suffer unnatural mouth shapes and asynchronous lips because those methods struggle to learn a consistent speech style from different speakers. We obser…

2021

Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

IJCAI 2021poster

We propose an audio-driven talking-head method to generate photo-realistic talking-head videos from a single reference image. In this work, we tackle two key challenges: (i) producing natural head motions that match speech prosody, and (ii)} maintaining the appearance of a speaker in a large head mo…

2021

Flow-Guided One-Shot Talking Face Generation With a High-Resolution Audio-Visual Dataset

CVPR 2021poster

One-shot talking face generation should synthesize high visual quality facial videos with reasonable animations of expression and head pose, and just utilize arbitrary driving audio and arbitrary single face image as the source. Current works fail to generate over 256 x 256 resolution realistic-look…

Cited by 368PDFcodeScholar
2021

Write-a-speaker: Text-based Emotional and Rhythmic Talking-head Generation

AAAI 2021technical

In this paper, we propose a novel text-based talking-head video generation framework that synthesizes high-fidelity facial expressions and head motions in accordance with contextual sentiments as well as speech rhythm and pauses. To be specific, our framework consists of a speaker-independent stage…