← Search

Liefeng Bo

44 accepted papers

2026

Think Fast and Slow: Step-Level Cognitive Depth Adaptation for LLM Agents

ICML 2026poster

Large language models (LLMs) are increasingly deployed as autonomous agents for multi-turn decision-making tasks. However, current agents typically rely on fixed cognitive patterns: non-thinking models generate immediate responses, while thinking models engage in deep reasoning uniformly. This rigid…

Cited by 0SourceScholar
2026

Twins: Learn to Predict Unified Representations with Focal Loss

ICML 2026poster

Unified multimodal models seek a shared visual token space that supports both multimodal understanding and image generation. Discrete methods unify the interface via a shared codebook, whereas continuous pipelines often rely on two disparate representations—semantic features (e.g., ViT) for understa…

Cited by 0SourceScholar
2025

AniGS: Animatable Gaussian Avatar from a Single Image with Inconsistent Gaussian Reconstruction

CVPR 2025poster

Generating animatable human avatars from a single image is essential for various digital human modeling applications. Existing 3D reconstruction methods often struggle to capture fine details in animatable models, while generative approaches for controllable animation, though avoiding explicit 3D mo…

2025

Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance

ICCV 2025poster

Recent character image animation methods based on diffusion models, such as Animate Anyone, have made significant progress in generating consistent and generalizable character animations. However, these approaches fail to produce reasonable associations between characters and their environments. To…

Cited by 0SourcePDFScholar
2025

Controllable and Expressive One-Shot Video Head Swapping

ICCV 2025poster

In this paper, we propose a novel diffusion-based multi-condition controllable framework for video head swapping, which seamlessly transplant a human head from a static image into a dynamic video, while preserving the original body and background of target video, and further allowing to tweak head e…

Cited by 0SourcePDFScholar
2025

Exploring Timeline Control for Facial Motion Generation

CVPR 2025poster

This paper introduces a new control signal for facial motion generation: timeline control. Compared to audio and text signals, timelines provide more fine-grained control, such as generating specific facial motions with precise timing. Users can specify a multi-track timeline of facial actions arran…

Cited by 0SourcePDFScholar
2025

ExtPose: Robust and Coherent Pose Estimation by Extending ViTs

ICML 2025poster

Vision Transformers (ViT) are remarkable at 3D pose estimation, yet they still encounter certain challenges. One issue is that the popular ViT architecture for pose estimation is limited to images and lacks temporal information. Another challenge is that the prediction often fails to maintain pixel…

Cited by 0SourcePDFScholar
2025

GaussianIP: Identity-Preserving Realistic 3D Human Generation via Human-Centric Diffusion Prior

CVPR 2025poster

Text-guided 3D human generation has advanced with the development of efficient 3D representations and 2D-lifting methods like score distillation sampling (SDS). However, current methods suffer from prolonged training times and often produce results that lack fine facial and garment details. In this…

2025

LHM: Large Animatable Human Reconstruction Model for Single Image to 3D in Seconds

ICCV 2025poster

Animatable 3D human reconstruction from a single image is a challenging problem due to the ambiguity in decoupling geometry, appearance, and deformation. Recent advances in 3D human reconstruction mainly focus on static human modeling, and the reliance of using synthetic 3D scans for training limits…

2025

MIMO: Controllable Character Video Synthesis with Spatial Decomposed Modeling

CVPR 2025poster

Character video synthesis aims to produce realistic videos of animatable characters within lifelike scenes. As a fundamental problem in the computer vision and graphics community, 3D works typically require multi-view captures for per-case training, which severely limits their applicability of model…

Cited by 18SourcePDFScholar
2025

Perception-as-Control: Fine-grained Controllable Image Animation with 3D-aware Motion Representation

ICCV 2025poster

Motion-controllable image animation is a fundamental task with a wide range of potential applications. Recent works have made progress in controlling camera or object motion via various motion representations, while they still struggle to support collaborative camera and object motion control with a…

Cited by 0SourcePDFScholar
2025

SemTalk: Holistic Co-speech Motion Generation with Frame-level Semantic Emphasis

ICCV 2025poster

A good co-speech motion generation cannot be achieved without a careful integration of common rhythmic motion and rare yet essential semantic motion. In this work, we propose SemTalk for holistic co-speech motion generation with frame-level semantic emphasis. Our key insight is to separately learn b…

Cited by 0SourcePDFScholar
2025

Synchronized Video-to-Audio Generation via Mel Quantization-Continuum Decomposition

CVPR 2025poster

Video-to-audio generation is essential for synthesizing realistic audio tracks that synchronize effectively with silent videos.Following the perspective of extracting essential signals from videos that can precisely control the mature text-to-audio generative diffusion models, this paper presents ho…

Cited by 0SourcePDFScholar
2025

Towards Fine-grained Interactive Segmentation in Images and Videos

ICCV 2025poster

The recent Segment Anything Models (SAMs) have emerged as foundational visual models for general interactive segmentation. Despite demonstrating robust generalization abilities, they still suffer from performance degradations in scenarios that demand accurate masks. Existing methods for high-precisi…

Cited by 0SourcePDFScholar
2025

UniPortrait: A Unified Framework for Identity-Preserving Single- and Multi-Human Image Personalization

ICCV 2025poster

This paper presents UniPortrait, an innovative human image personalization framework that unifies single- and multi-ID customization with high face fidelity, extensive facial editability, free-form input description, and diverse layout generation. UniPortrait consists of only two plug-and-play modul…

2024

An Optimization Framework to Enforce Multi-View Consistency for Texturing 3D Meshes

ECCV 2024poster

"A fundamental problem in the texturing of 3D meshes using pre-trained text-to-image models is to ensure multi-view consistency. State-of-the-art approaches typically use diffusion models to aggregate multi-view inputs, where common issues are the blurriness caused by the averaging operation in the…

2024

Evaluate Geometry of Radiance Fields with Low-Frequency Color Prior

AAAI 2024technical

A radiance field is an effective representation of 3D scenes, which has been widely adopted in novel-view synthesis and 3D reconstruction. It is still an open and challenging problem to evaluate the geometry, i.e., the density field, as the ground-truth is almost impossible to obtain. One alternativ…

2024

GIC: Gaussian-Informed Continuum for Physical Property Identification and Simulation

NeurIPS 2024oral

This paper studies the problem of estimating physical properties (system identification) through visual observations. To facilitate geometry-aware guidance in physical property estimation, we introduce a novel hybrid framework that leverages 3D Gaussian representation to not only capture explicit sh…

2024

GPLD3D: Latent Diffusion of 3D Shape Generative Models by Enforcing Geometric and Physical Priors

CVPR 2024poster

State-of-the-art man-made shape generative models usually adopt established generative models under a suitable implicit shape representation. A common theme is to perform distribution alignment which does not explicitly model important shape priors. As a result many synthetic shapes are not connecte…

Cited by 8SourcePDFScholar
2024

High-Fidelity 3D Textured Shapes Generation by Sparse Encoding and Adversarial Decoding

ECCV 2024poster

"3D vision is inherently characterized by sparse spatial structures, which propels the necessity for an efficient paradigm tailored to 3D generation. Another discrepancy is the amount of training data, which undeniably affects generalization if we only use limited 3D data. To solve these, we design…

2024

High-Fidelity and Transferable NeRF Editing by Frequency Decomposition

ECCV 2024poster

"This paper enables high-fidelity, transferable NeRF editing by frequency decomposition. Recent NeRF editing pipelines lift 2D stylization results to 3D scenes while suffering from blurry results, and fail to capture detailed structures caused by the inconsistency between 2D editings. Our critical i…

2024

IPoD: Implicit Field Learning with Point Diffusion for Generalizable 3D Object Reconstruction from Single RGB-D Images

CVPR 2024highlight

Generalizable 3D object reconstruction from single-view RGB-D images remains a challenging task particularly with real-world data. Current state-of-the-art methods develop Transformer-based implicit field learning necessitating an intensive learning paradigm that requires dense query-supervision uni…

2024

MoGenTS: Motion Generation based on Spatial-Temporal Joint Modeling

NeurIPS 2024poster

Motion generation from discrete quantization offers many advantages over continuous regression, but at the cost of inevitable approximation errors. Previous methods usually quantize the entire body pose into one code, which not only faces the difficulty in encoding all joints within one vector but a…

Cited by 2SourcePDFScholar
2024

Open-Vocabulary Category-Level Object Pose and Size Estimation

RA-L 2024

This letter studies a new open-set problem, the open-vocabulary category-level object pose and size estimation. Given human text descriptions of arbitrary novel object categories, the robot agent seeks to predict the position, orientation, and size of the target object in the observed scene image. T

Cited by 11SourceScholar
2024

PanoContext-Former: Panoramic Total Scene Understanding with a Transformer

CVPR 2024poster

Panoramic images enable deeper understanding and more holistic perception of 360 surrounding environment which can naturally encode enriched scene context information compared to standard perspective image. Previous work has made lots of effort to solve the scene understanding task in a hybrid solut…

Cited by 11SourcePDFScholar
2024

RichDreamer: A Generalizable Normal-Depth Diffusion Model for Detail Richness in Text-to-3D

CVPR 2024highlight

Lifting 2D diffusion for 3D generation is a challenging problem due to the lack of geometric prior and the complex entanglement of materials and lighting in natural images. Existing methods have shown promise by first creating the geometry through score-distillation sampling (SDS) applied to rendere…

2023

Compressing Volumetric Radiance Fields to 1 MB

CVPR 2023poster

Approximating radiance fields with discretized volumetric grids is one of promising directions for improving NeRFs, represented by methods like DVGO, Plenoxels and TensoRF, which achieve super-fast training convergence and real-time rendering. However, these methods typically require a tremendous st…

2023

DG3D: Generating High Quality 3D Textured Shapes by Learning to Discriminate Multi-Modal Diffusion-Renderings

ICCV 2023poster

Many virtual reality applications require massive 3D content, which impels the need for low-cost and efficient modeling tools in terms of quality and quantity. In this paper, we present a Diffusion-augmented Generative model to generate high-fidelity 3D textured meshes that can be directly used in m…

Cited by 3PDFcodeScholar
2023

One-Shot High-Fidelity Talking-Head Synthesis With Deformable Neural Radiance Field

CVPR 2023poster

Talking head generation aims to generate faces that maintain the identity information of the source image and imitate the motion of the driving image. Most pioneering methods rely primarily on 2D representations and thus will inevitably suffer from face distortion when large head rotations are encou…

Cited by 54SourcePDFScholar
2023

Reducing Shape-Radiance Ambiguity in Radiance Fields with a Closed-Form Color Estimation Method

NeurIPS 2023poster

A neural radiance field (NeRF) enables the synthesis of cutting-edge realistic novel view images of a 3D scene. It includes density and color fields to model the shape and radiance of a scene, respectively. Supervised by the photometric loss in an end-to-end training manner, NeRF inherently suffers…

2023

RenderIH: A Large-Scale Synthetic Dataset for 3D Interacting Hand Pose Estimation

ICCV 2023poster

The current interacting hand (IH) datasets are relatively simplistic in terms of background and texture, with hand joints being annotated by a machine annotator, which may result in inaccuracies, and the diversity of pose distribution is limited. However, the variability of background, pose distribu…

Cited by 19PDFcodeScholar
2023

Towards Stable Human Pose Estimation via Cross-View Fusion and Foot Stabilization

CVPR 2023poster

Towards stable human pose estimation from monocular images, there remain two main dilemmas. On the one hand, the different perspectives, i.e., front view, side view, and top view, appear the inconsistent performances due to the depth ambiguity. On the other hand, foot posture plays a significant rol…

Cited by 5SourcePDFScholar
2021

Detection, Tracking, and Counting Meets Drones in Crowds: A Benchmark

CVPR 2021poster

To promote the developments of object detection, tracking and counting algorithms in drone-captured videos, we construct a benchmark with a new drone-captured large-scale dataset, named as DroneCrowd, formed by 112 video clips with 33,600 HD frames in various scenarios. Notably, we annotate 20,800 p…

Cited by 133PDFcodeScholar
2021

Graph-Enhanced Multi-Task Learning of Multi-Level Transition Dynamics for Session-based Recommendation

AAAI 2021technical

Session-based recommendation plays a central role in a wide spectrum of online applications, ranging from e-commerce to online advertising services. However, the majority of existing session-based recommendation techniques (e.g., attention-based recurrent network or graph neural network) are not wel…

2021

Knowledge-Enhanced Hierarchical Graph Transformer Network for Multi-Behavior Recommendation

AAAI 2021technical

Accurate user and item embedding learning is crucial for modern recommender systems. However, most existing recommendation techniques have thus far focused on modeling users' preferences over singular type of user-item interactions. Many practical recommendation scenarios involve multi-typed user in…

2021

Knowledge-aware Coupled Graph Neural Network for Social Recommendation

AAAI 2021technical

Social recommendation task aims to predict users' preferences over items with the incorporation of social connections among users, so as to alleviate the sparse issue of collaborative filtering. While many recent efforts show the effectiveness of neural network-based social recommender systems, seve…

2021

Spatial-Temporal Sequential Hypergraph Network for Crime Prediction with Dynamic Multiplex Relation Learning

IJCAI 2021poster

Crime prediction is crucial for public safety and resource optimization, yet is very challenging due to two aspects: i) the dynamics of criminal patterns across time and space, crime events are distributed unevenly on both spatial and temporal domains; ii) time-evolving dependencies between differen…

2020

Cross-Interaction Hierarchical Attention Networks for Urban Anomaly Prediction

IJCAI 2020poster

Predicting anomalies (e.g., blocked driveway and vehicle collisions) in urban space plays an important role in assisting governments and communities for building smart city applications, ranging from intelligent transportation to public safety. However, predicting urban anomalies is not trivial due…

Cited by 0SourcePDFScholar
2020

Efficient Pig Counting in Crowds with Keypoints Tracking and Spatial-aware Temporal Response Filtering

ICRA 2020poster

Pig counting is a crucial task for large-scale pig farming. Pigs are usually visually counted by human. But this process is very time-consuming and error-prone. Few studies in literature developed automated pig counting method. The existing works only focused on pig counting using single image, and…

Cited by 32SourceScholar
2019

ScratchDet: Training Single-Shot Object Detectors From Scratch

CVPR 2019oral

Current state-of-the-art object objectors are fine-tuned from the off-the-shelf networks pretrained on large-scale classification dataset ImageNet, which incurs some additional problems: 1) The classification and detection have different degrees of sensitivity to translation, resulting in the learni…

Cited by 188PDFcodeScholar