← Search

Yu-Kun Lai

57 accepted papers

2026

CG-Floor: Centroid-Guided Diffusion for Large-Scale Floorplan Generation

CVPR 2026

Large-scale floorplan generation is critical for virtual space planning and architectural simulation. Although existing methods have shown success in generating small-scale floorplans with simple room shapes, they struggle to handle complex room connections and irregular room shapes that arise in la

Cited by 0SourceScholar
2026

Crowd4D: Scene-Aware Monocular 4D Crowd Reconstruction

ICML 2026poster

Recovering scene-consistent 4D crowd motion from monocular video in large-scale scenes remains challenging due to severe depth ambiguity and complex scene geometry. Existing monocular crowd reconstruction methods typically rely on single-plane assumptions, leading to unreliable metric scale and spat…

Cited by 0SourceScholar
2026

Diagram2Structure: Unlocking LLMs' Diagram Comprehension through DiagramDiff, a Framework for Structuring Offline Diagrams

CVPR 2026

Diagrams are widely used in daily life. However, offline diagrams typically exist in the form of images, lacking structured data representation, which significantly limits their reusability and editability. Current research mainly focuses on supporting basic query tasks for online diagrams and does

Cited by 0SourceScholar
2026

InterCoser: Interactive 3D Character Creation with Disentangled Fine-Grained Features

AAAI 2026technical

This paper aims to interactively generate and edit disentangled 3D characters based on precise user instructions. Existing methods generate and edit 3D characters via rough and simple editing guidance and entangled representations, making it difficult to achieve precise and comprehensive control ove

Cited by 0SourcePDFScholar
2026

One-Shot View Planning and Online Optimization-Based Replanning for Unknown Object Reconstruction

ICRA 2026poster

Robotic inspection tasks often require constructing high-quality 3D models of objects from a minimal number of views. Traditional next-best view planning (NBVP) approaches incrementally select view poses but fail to account for global optimality of the inspection trajectory, thus leading to ineffici…

Cited by 0Scholar
2026

Temporal Inconsistency Guidance for Super-resolution Video Quality Assessment

AAAI 2026technical

As super-resolution (SR) techniques introduce unique distortions that fundamentally differ from those caused by traditional degradation processes (e.g., compression), there is an increasing demand for specialized video quality assessment (VQA) methods tailored to SR-generated content. One critical f

Cited by 0SourcePDFScholar
2025

Celebi's Choice: Causality-Guided Skill Optimisation for Granular Manipulation via Differentiable Simulation

IROS 2025

Robotic soil manipulation is essential for automated farming, particularly in excavation and levelling tasks. However, the nonlinear dynamics of granular materials challenge traditional control methods, limiting stability and efficiency. We propose Celebi, a causality-enhanced optimisation method th

Cited by 0SourceScholar
2025

Chat-Driven 3D Human Pose and Shape Editing with Large Language Models

ICASSP 2025accepted

Generating and creating humanoid 3D models has received increasing attention recently due to its fundamental support for many high-level 3D applications. Although automatic 3D pose and shape reconstruction methods have achieved promising results, there are still some failure cases due to self-occlus…

Cited by 0SourceScholar
2025

LGA-Net: Learning Local and Global Affinities for Sparse Scribble based Image Colorization

ICCV 2025poster

Image colorization is a typical ill-posed problem. Among various colorization methods, scribble-based methods have a unique advantage that allows users to accurately resolve ambiguities and modify the colors of any objects to suit their specific tastes. However, due to the time-consuming scribble dr…

2025

RESCUE: Crowd Evacuation Simulation via Controlling SDM-United Characters

ICCV 2025poster

Crowd evacuation simulation is critical for enhancing public safety, and demanded for realistic virtual environments. Current mainstream evacuation models overlook the complex human behaviors that occur during evacuation, such as pedestrian collisions, interpersonal interactions, and variations in b…

Cited by 0SourcePDFScholar
2025

Skeleton-Guided Rolling-Contact Kinematics for Arbitrary Point Clouds via Locally Controllable Parameterized Curve Fitting

IROS 2025

Rolling contact kinematics plays a vital role in dexterous manipulation and rolling-based locomotion. Yet, in practical applications, the environments and objects involved are often captured as discrete point clouds, creating substantial difficulties for traditional motion control and planning frame

Cited by 0SourceScholar
2024

AHRNET: Attention and Heatmap-Based Regressor for Hand Pose Estimation and Mesh Recovery

ICASSP 2024accepted

Estimating 3D hand pose and recovering the full hand surface mesh from a single RGB image is a challenging task due to self-occlusions, viewpoint changes, and the complexity of hand articulations. In this paper, we propose a novel framework that combines an attention mechanism with heatmap regressio…

Cited by 0SourceScholar
2024

Efficient Precision and Recall Metrics for Assessing Generative Models using Hubness-aware Sampling

ICML 2024spotlight

Despite impressive results, deep generative models require massive datasets for training, and as dataset size increases, effective evaluation metrics like precision and recall (P&R) become computationally infeasible on commodity hardware. In this paper, we address this challenge by proposing efficie…

2024

GLSkeleton: A Geometric Laplacian-Based Skeletonisation Framework for Object Point Clouds

RA-L 2024

The curve skeleton is known to geometric modelling and computer graphics communities as one of the shape descriptors which intuitively indicates the topological properties of the objects. In recent years, studies have also suggested the potential of applying curve skeletons to assist robotic reasoni

Cited by 7SourceScholar
2024

Real-time 3D-aware Portrait Video Relighting

CVPR 2024highlight

Synthesizing realistic videos of talking faces under custom lighting conditions and viewing angles benefits various downstream applications like video conferencing. However most existing relighting methods are either time-consuming or unable to adjust the viewpoints. In this paper we present the fir…

2024

SAMVG: A Multi-Stage Image Vectorization Model with the Segment-Anything Model

ICASSP 2024accepted

Vector graphics are widely used in graphical designs and have received more and more attention. However, unlike raster images which can be easily obtained, acquiring high-quality vector graphics, typically through automatically converting from raster images, remains a significant challenge, especial…

Cited by 0SourceScholar
2024

SceneDiff: Generative Scene-Level Image Retrieval with Text and Sketch Using Diffusion Models

IJCAI 2024poster

Jointly using text and sketch for scene-level image retrieval utilizes the complementary between text and sketch to describe the fine-grained scene content and retrieve the target image, which plays a pivotal role in accurate image retrieval. Existing methods directly fuse the features of sketch and…

Cited by 0SourcePDFScholar
2024

Texture-GS: Disentangle the Geometry and Texture for 3D Gaussian Splatting Editing

ECCV 2024poster

"3D Gaussian splatting, emerging as a groundbreaking approach, has drawn increasing attention for its capabilities of high-fidelity reconstruction and real-time rendering. However, it couples the appearance and geometry of the scene within the Gaussian attributes, which hinders the flexibility of ed…

Cited by 16SourcePDFScholar
2023

BBDM: Image-to-Image Translation With Brownian Bridge Diffusion Models

CVPR 2023poster

Image-to-image translation is an important and challenging problem in computer vision and image processing. Diffusion models(DM) have shown great potentials for high-quality image synthesis, and have gained competitive performance on the task of image-to-image translation. However, most of the exist…

2023

CXTrack: Improving 3D Point Cloud Tracking With Contextual Information

CVPR 2023poster

3D single object tracking plays an essential role in many applications, such as autonomous driving. It remains a challenging problem due to the large appearance variation and the sparsity of points caused by occlusion and limited sensor capabilities. Therefore, contextual information across two cons…

Cited by 42SourcePDFScholar
2023

Crowd3D: Towards Hundreds of People Reconstruction From a Single Image

CVPR 2023poster

Image-based multi-person reconstruction in wide-field large scenes is critical for crowd analysis and security alert. However, existing methods cannot deal with large scenes containing hundreds of people, which encounter the challenges of large number of people, large variations in human scale, and…

Cited by 13SourcePDFScholar
2023

E3Sym: Leveraging E(3) Invariance for Unsupervised 3D Planar Reflective Symmetry Detection

ICCV 2023poster

Detecting symmetrical properties is a fundamental task in 3D shape analysis. In the case of a 3D model with planar symmetries, each point has a corresponding mirror point w.r.t. a symmetry plane, and the correspondences remain invariant under any arbitrary Euclidean transformation. Our proposed meth…

Cited by 11PDFcodeScholar
2023

FEditNet: Few-Shot Editing of Latent Semantics in GAN Spaces

AAAI 2023technical

Generative Adversarial networks (GANs) have demonstrated their powerful capability of synthesizing high-resolution images, and great efforts have been made to interpret the semantics in the latent spaces of GANs. However, existing works still have the following limitations: (1) the majority of works…

2023

Feature Proliferation -- the "Cancer" in StyleGAN and its Treatments

ICCV 2023poster

Despite the success of StyleGAN in image synthesis, the images it synthesizes are not always perfect and the well-known truncation trick has become a standard post-processing technique for StyleGAN to synthesize high-quality images. Although effective, it has long been noted that the truncation tric…

Cited by 0PDFcodeScholar
2023

Learning Semantic-Aware Disentangled Representation for Flexible 3D Human Body Editing

CVPR 2023poster

3D human body representation learning has received increasing attention in recent years. However, existing works cannot flexibly, controllably and accurately represent human bodies, limited by coarse semantics and unsatisfactory representation capability, particularly in the absence of supervised da…

Cited by 8SourcePDFScholar
2023

MBPTrack: Improving 3D Point Cloud Tracking with Memory Networks and Box Priors

ICCV 2023poster

3D single object tracking has been a crucial problem for decades with numerous applications such as autonomous driving. Despite its wide-ranging use, this task remains challenging due to the significant appearance variation caused by occlusion and size differences among tracked targets. To address t…

Cited by 29PDFScholar
2023

NeuralSlice: Neural 3D Triangle Mesh Reconstruction via Slicing 4D Tetrahedral Meshes

ICML 2023poster

Learning-based high-fidelity reconstruction of 3D shapes with varying topology is a fundamental problem in computer vision and computer graphics. Recent advances in learning 3D shapes using explicit and implicit representations have achieved impressive results in 3D modeling. However, the template-b…

2023

Towards Artistic Image Aesthetics Assessment: A Large-Scale Dataset and a New Method

CVPR 2023poster

Image aesthetics assessment (IAA) is a challenging task due to its highly subjective nature. Most of the current studies rely on large-scale datasets (e.g., AVA and AADB) to learn a general model for all kinds of photography images. However, little light has been shed on measuring the aesthetic qual…

2022

Exploring and Exploiting Hubness Priors for High-Quality GAN Latent Sampling

ICML 2022spotlight

Despite the extensive studies on Generative Adversarial Networks (GANs), how to reliably sample high-quality images from their latent spaces remains an under-explored topic. In this paper, we propose a novel GAN latent sampling method by exploring and exploiting the hubness priors of GAN latent dist…

2022

FOF: Learning Fourier Occupancy Field for Monocular Real-time Human Reconstruction

NeurIPS 2022accept

The advent of deep learning has led to significant progress in monocular human reconstruction. However, existing representations, such as parametric models, voxel grids, meshes and implicit neural representations, have difficulties achieving high-quality results and real-time speed at the same time.…

Cited by 39SourcePDFScholar
2022

High-Fidelity Human Avatars From a Single RGB Camera

CVPR 2022poster

In this paper, we propose a coarse-to-fine framework to reconstruct a personalized high-fidelity human avatar from a monocular video. To deal with the misalignment problem caused by the changed poses and shapes in different frames, we design a dynamic surface network to recover pose-dependent surfac…

Cited by 40PDFScholar
2022

NeRF-Editing: Geometry Editing of Neural Radiance Fields

CVPR 2022poster

Implicit neural rendering, especially Neural Radiance Field (NeRF), has shown great potential in novel view synthesis of a scene. However, current NeRF-based methods cannot enable users to perform user-controlled shape deformation in the scene. While existing works have proposed some approaches to m…

Cited by 285PDFScholar
2022

StylizedNeRF: Consistent 3D Scene Stylization As Stylized NeRF via 2D-3D Mutual Learning

CVPR 2022poster

3D scene stylization aims at generating stylized images of the scene from arbitrary novel views following a given set of style examples, while ensuring consistency when rendered from different views. Directly applying methods for image or video stylization to 3D scenes cannot achieve such consistenc…

Cited by 172PDFcodeScholar
2021

Hierarchical Layout-Aware Graph Convolutional Network for Unified Aesthetics Assessment

CVPR 2021poster

Learning computational models of image aesthetics can have a substantial impact on visual art and graphic design. Although automatic image aesthetics assessment is a challenging topic by its subjective nature, psychological studies have confirmed a strong correlation between image layouts and percei…

Cited by 98PDFcodeScholar
2021

MLVSNet: Multi-Level Voting Siamese Network for 3D Visual Tracking

ICCV 2021poster

Benefiting from the excellent performance of Siamese-based trackers, huge progress on 2D visual tracking has been achieved. However, 3D visual tracking is still under-explored. Inspired by the idea of Hough voting in 3D object detection, in this paper, we propose a Multi-level Voting Siamese Network…

Cited by 66PDFcodeScholar
2021

Manifold Alignment for Semantically Aligned Style Transfer

ICCV 2021poster

Most existing style transfer methods follow the assumption that styles can be represented with global statistics (e.g., Gram matrices or covariance matrices), and thus address the problem by forcing the output and style images to have similar global statistics. An alternative is the assumption of lo…

Cited by 60PDFcodeScholar
2021

Single Image 3D Shape Retrieval via Cross-Modal Instance and Category Contrastive Learning

ICCV 2021poster

In this work, we tackle the problem of single image-based 3D shape retrieval (IBSR), where we seek to find the most matched shape of a given single 2D image from a shape repository. Most of the existing works learn to embed 2D images and 3D shapes into a common feature space and perform metric learn…

Cited by 40PDFcodeScholar
2021

VENet: Voting Enhancement Network for 3D Object Detection

ICCV 2021poster

Hough voting, as has been demonstrated in VoteNet, is effective for 3D object detection, where voting is a key step. In this paper, we propose a novel VoteNet-based 3D detector with vote enhancement to improve the detection accuracy in cluttered indoor scenes. It addresses the limitations of current…

Cited by 61PDFScholar
2020

MLCVNet: Multi-Level Context VoteNet for 3D Object Detection

CVPR 2020poster

In this paper, we address the 3D object detection task by capturing multi-level contextual information with the self-attention mechanism and multi-scale feature fusion. Most existing 3D object detection methods recognize objects individually, without giving any consideration on contextual informatio…

Cited by 229PDFcodeScholar
2020

SceneSketcher: Fine-Grained Image Retrieval with Scene Sketches

ECCV 2020poster

Sketch-based image retrieval (SBIR) has been a popular research topic in recent years. Existing works concentrate on mapping the visual information of sketches and images to a semantic space at the object level. In this paper, for the first time, we study the fine-grained scene-level SBIR problem wh…

Cited by 45SourcePDFScholar
2019

APDrawingGAN: Generating Artistic Portrait Drawings From Face Photos With Hierarchical GANs

CVPR 2019oral

Significant progress has been made with image stylization using deep learning, especially with generative adversarial networks (GANs). However, existing methods fail to produce high quality artistic portrait drawings. Such drawings have a highly abstract style, containing a sparse set of continuous…

Cited by 200PDFScholar
2019

Attention-Aware Polarity Sensitive Embedding for Affective Image Retrieval

ICCV 2019poster

Images play a crucial role for people to express their opinions online due to the increasing popularity of social networks. While an affective image retrieval system is useful for obtaining visual contents with desired emotions from a massive repository, the abstract and subjective characteristics m…

Cited by 45PDFScholar
2019

ClusterSLAM: A SLAM Backend for Simultaneous Rigid Body Clustering and Motion Estimation

ICCV 2019poster

We present a practical backend for stereo visual SLAM which can simultaneously discover individual rigid bodies and compute their motions in dynamic environments. While recent factor graph based state optimization algorithms have shown their ability to robustly solve SLAM problems by treating dynami…

Cited by 97PDFScholar
2019

IP102: A Large-Scale Benchmark Dataset for Insect Pest Recognition

CVPR 2019oral

Insect pests are one of the main factors affecting agricultural product yield. Accurate recognition of insect pests facilitates timely preventive measures to avoid economic losses. However, the existing datasets for the visual classification task mainly focus on common objects, e.g., flowers and do…

Cited by 519PDFcodeScholar
2019

Joint Acne Image Grading and Counting via Label Distribution Learning

ICCV 2019accepted

Accurate grading of skin disease severity plays a crucial role in precise treatment for patients. Acne vulgaris, the most common skin disease in adolescence, can be graded by evidence-based lesion counting as well as experience-based global estimation in the medical field. However, due to the appear…

2019

Probabilistic Projective Association and Semantic Guided Relocalization for Dense Reconstruction

ICRA 2019poster

We present a real-time dense mapping system which uses the predicted 2D semantic labels for optimizing the geometric quality of reconstruction. With a combination of Convolutional Neural Networks (CNNs) for 2D labeling and a Simultaneous Localization and Mapping (SLAM) system for camera trajectory e…

Cited by 14SourceScholar
2019

SketchGAN: Joint Sketch Completion and Recognition With Generative Adversarial Network

CVPR 2019poster

Hand-drawn sketch recognition is a fundamental problem in computer vision, widely used in sketch-based image and video retrieval, editing, and reorganization. Previous methods often assume that a complete sketch is used as input; however, hand-drawn sketches in common application scenarios are often…

Cited by 71PDFScholar
2019

VV-Net: Voxel VAE Net With Group Convolutions for Point Cloud Segmentation

ICCV 2019poster

We present a novel algorithm for point cloud segmentation.Our approach transforms unstructured point clouds into regular voxel grids, and further uses a kernel-based interpolated variational autoencoder (VAE) architecture to encode the local geometry within each voxel.Traditionally, the voxel repres…

Cited by 348PDFcodeScholar
2018

Weakly Supervised Coupled Networks for Visual Sentiment Analysis

CVPR 2018poster

Automatic assessment of sentiment from visual content has gained considerable attention with the increasing tendency of expressing opinions on-line. In this paper, we solve the problem of visual sentiment analysis using the high-level abstraction in the recognition process. Existing methods based on…

Cited by 161SourcePDFScholar