← Search

Yulan Guo

65 accepted papers

2026

Design Your Ad: Personalized Advertising Image and Text Generation with Unified Autoregressive Models

CVPR 2026

Generating realistic and user-preferred advertisements is a key challenge in e-commerce. Existing approaches utilize multiple independent models driven by click-through-rate (CTR) to controllably create attractive image or text advertisements. However, their pipelines lack cross-modal perception and

Cited by 0SourcecodeScholar
2026

Edge-Centric Relational Reasoning for 3D Scene Graph Prediction

AAAI 2026technical

3D scene graph prediction aims to abstract complex 3D environments into structured graphs consisting of objects and their pairwise relationships. Existing approaches typically adopt object-centric graph neural networks, where relation edge features are iteratively updated by aggregating messages fro

Cited by 0SourcePDFScholar
2026

MangoBench: A Benchmark for Multi-Agent Goal-Conditioned Offline Reinforcement Learning

CVPR 2026

Offline Multi-Agent Reinforcement Learning (MARL) is critical for coordinating multiple agents in costly and unsafe environments, yet existing methods struggle with high sensitivity to reward functions and weak generalization to new goals, limiting its practical impact. Inspired by single-agent Offl

Cited by 0SourceScholar
2026

Seeing Motion, Generating Action: Explicit Motion-Aware Policy for Robotic Action Generation

ICRA 2026poster

Imitation learning (IL) offers a scalable framework for teaching robots complex manipulation skills from human demonstrations. However, conventional end-to-end visuomotor IL models often suffer from poor performance and robustness due to the significant modality mismatch between high-dimensional vis…

Cited by 0Scholar
2026

Viper: Verifiable Imitation Learning Policy for Efficient Robotic Manipulation

ICRA 2026poster

Imitation learning (IL) presents a promising paradigm for enabling embodied robots to efficiently acquire human-like manipulation skills. However, prevailing methods face a persistent trade-off between motion precision and computational tractability. To resolve this fundamental challenge, this paper…

Cited by 0Scholar
2025

$U2$ Frame: A Unified and Unsupervised Learning Framework for LiDAR-Based Loop Closing

ICRA 2025

Loop closing is critically important in Simultaneous Localization and Mapping (SLAM) due to its ability to correct accumulated localization errors. However, existing methods are hindered by the difficulty of acquiring pose labels and the unreliability of ground truth data. In this paper, we propose

Cited by 0SourcecodeScholar
2025

3D Whole-Body Pose Estimation Using Graph High-Resolution Network for Humanoid Robot Teleoperation

ICRA 2025

In the realm of robotics, teleoperation plays a pivotal role in performing high-risk or intricate tasks, and obtaining precise 3D whole-body pose is crucial for this purpose. Traditional two-stage methods have limitations in estimating different body parts, leading to complex systems and higher esti

Cited by 0SourcecodeScholar
2025

AIQViT: Architecture-Informed Post-Training Quantization for Vision Transformers

AAAI 2025technical

Post-training quantization (PTQ) has emerged as a promising solution for reducing the storage and computational cost of vision transformers (ViTs). Recent advances primarily target at crafting quantizers to deal with peculiar activations characterized by ViTs. However, most existing methods underest…

Cited by 0SourcePDFScholar
2025

CDP: Towards Robust Autoregressive Visuomotor Policy Learning via Causal Diffusion

CoRL 2025poster

Diffusion Policy (DP) enables robots to learn complex behaviors by imitating expert demonstrations through action diffusion. However, in practical applications, hardware limitations often degrade data quality, while real-time constraints restrict model inference to instantaneous state and scene obse…

Cited by 0SourceScholar
2025

DropoutGS: Dropping Out Gaussians for Better Sparse-view Rendering

CVPR 2025poster

Although 3D Gaussian Splatting (3DGS) has demonstrated promising results in novel view synthesis, its performance degrades dramatically with sparse inputs and generates undesirable artifacts. As the number of training views decreases, the novel view synthesis task degrades to a highly under-determin…

Cited by 0SourcePDFScholar
2025

Graph2Scene: Versatile 3D Indoor Scene Generation with Interaction-aware Scene Graph

IROS 2025

Embodied artificial intelligence requires a wide variety of large-scale simulated environments for development. Previous scene reconstruction approaches based on multiview images can produce high-fidelity 3D scenes but lack diversity. In contrast, existing prompt-based scene generation approaches ca

Cited by 0SourceScholar
2025

Multi-Modality Test-Time Adaptation for Semantic Segmentation in Robotic Perception

ICRA 2025

Test-Time Adaptation (TTA) adjusts pre-trained models in unlabeled unseen environments during the test phase, making it more practical for robotic applications. However, the constant changes of the physical world create significant domain gaps between the received data during robot deployment and th

Cited by 0SourceScholar
2025

OnlineAnySeg: Online Zero-Shot 3D Segmentation by Visual Foundation Model Guided 2D Mask Merging

CVPR 2025poster

Online 3D open-vocabulary segmentation of a progressively reconstructed scene is both a critical and challenging task for embodied applications. With the success of visual foundation models (VFMs) in the image domain, leveraging 2D priors to address 3D online segmentation has become a prominent rese…

Cited by 0SourcePDFScholar
2025

Progressive Correspondence Regenerator for Robust 3D Registration

CVPR 2025poster

Obtaining enough high-quality correspondences is crucial for robust registration. Existing correspondence refinement methods mostly follow the paradigm of outlier removal, which either fails to correctly identify the accurate correspondences under extreme outlier ratios, or select too few correct co…

2025

SaMam: Style-aware State Space Model for Arbitrary Image Style Transfer

CVPR 2025highlight

Global effective receptive field plays a crucial role for image style transfer (ST) to obtain high-quality stylized results. However, existing ST backbones (e.g., CNNs and Transformers) suffer huge computational complexity to achieve global receptive fields. Recently, the State Space Model (SSM), es…

2025

Self-Distilled Stereo Matching: Real-Time Domain Generalization for Robotic Depth Perception

IROS 2025

While human vision inherently achieves robust cross-domain depth estimation through binocular coordination, robotic systems employing stereo matching still confront significant challenges in maintaining robustness across domains when performing real-time environmental depth perception. Furthermore,

Cited by 0SourceScholar
2025

TopNet: Transformer-Efficient Occupancy Prediction Network for Octree-Structured Point Cloud Geometry Compression

CVPR 2025poster

Efficient Point Cloud Geometry Compression (PCGC) with a lower bits per point (BPP) and higher peak signal-to-noise ratio (PSNR) is essential for the transportation of large-scale 3D data. Although octree-based entropy models can reduce BPP without introducing geometry distortion, existing CNN-based…

2025

VideoDirector: Precise Video Editing via Text-to-Video Models

CVPR 2025poster

Despite the typical inversion-then-editing paradigm using text-to-image (T2I) models has demonstrated promising results, directly extending it to text-to-video (T2V) models still suffers severe artifacts such as color flickering and content distortion. Consequently, current video editing methods pri…

Cited by 0SourcePDFScholar
2024

ACRF: Compressing Explicit Neural Radiance Fields via Attribute Compression

ICLR 2024poster

In this work, we study the problem of explicit NeRF compression. Through analyzing recent explicit NeRF models, we reformulate the task of explicit NeRF compression as 3D data compression. We further introduce our NeRF compression framework, Attributed Compression of Radiance Field (ACRF), which foc…

Cited by 3SourcePDFScholar
2024

Density-guided Translator Boosts Synthetic-to-Real Unsupervised Domain Adaptive Segmentation of 3D Point Clouds

CVPR 2024poster

3D synthetic-to-real unsupervised domain adaptive segmentation is crucial to annotating new domains. Self-training is a competitive approach for this task but its performance is limited by different sensor sampling patterns (i.e. variations in point density) and incomplete training strategies. In th…

2024

L4D-Track: Language-to-4D Modeling Towards 6-DoF Tracking and Shape Reconstruction in 3D Point Cloud Stream

CVPR 2024poster

3D visual language multi-modal modeling plays an important role in actual human-computer interaction. However the inaccessibility of large-scale 3D-language pairs restricts their applicability in real-world scenarios. In this paper we aim to handle a real-time multi-task for 6-DoF pose tracking of u…

Cited by 0SourcePDFScholar
2024

Learning Coupled Dictionaries from Unpaired Data for Image Super-Resolution

CVPR 2024poster

The difficulty of acquiring high-resolution (HR) and low-resolution (LR) image pairs in real scenarios limits the performance of existing learning-based image super-resolution (SR) methods in the real world. To conduct training on real-world unpaired data current methods focus on synthesizing pseudo…

Cited by 3SourcePDFScholar
2024

LoS: Local Structure-Guided Stereo Matching

CVPR 2024poster

Estimating disparities in challenging areas is difficult and limits the performance of stereo matching models. In this paper we exploit local structure information (LSI) to enhance stereo matching. Specifically our LSI comprises a series of key elements including the slant plane (parameterised by di…

Cited by 14SourcePDFScholar
2023

2D3D-MATR: 2D-3D Matching Transformer for Detection-Free Registration Between Images and Point Clouds

ICCV 2023poster

The commonly adopted detect-then-match approach to registration finds difficulties in the cross-modality cases due to the incompatible keypoint detection and inconsistent feature description. We propose, 2D3D-MATR, a detection-free method for accurate and robust registration between images and point…

Cited by 18PDFcodeScholar
2023

3D Spatial Multimodal Knowledge Accumulation for Scene Graph Prediction in Point Cloud

CVPR 2023poster

In-depth understanding of a 3D scene not only involves locating/recognizing individual objects, but also requires to infer the relationships and interactions among them. However, since 3D scenes contain partially scanned objects with physical connections, dense placement, changing sizes, and a wide…

2023

BUFFER: Balancing Accuracy, Efficiency, and Generalizability in Point Cloud Registration

CVPR 2023poster

An ideal point cloud registration framework should have superior accuracy, acceptable efficiency, and strong generalizability. However, this is highly challenging since existing registration techniques are either not accurate enough, far from efficient, or generalized poorly. It remains an open ques…

2023

Context-Aware Alignment and Mutual Masking for 3D-Language Pre-Training

CVPR 2023highlight

3D visual language reasoning plays an important role in effective human-computer interaction. The current approaches for 3D visual reasoning are task-specific, and lack pre-training methods to learn generic representations that can transfer across various tasks. Despite the encouraging progress in v…

2023

HI-Net: Boosting Self-Supervised Indoor Depth Estimation via Pose Optimization

RA-L 2023

Pose estimation plays a critical role in self-supervised monocular depth estimation for indoor scenes, especially those involving complex ego-motion. In this letter, we leverage the two-view geometry constraints into pose estimation to boost the accuracy of pose estimation, which ultimately improves

Cited by 1SourceScholar
2023

Learning Non-Local Spatial-Angular Correlation for Light Field Image Super-Resolution

ICCV 2023poster

Exploiting spatial-angular correlation is crucial to light field (LF) image super-resolution (SR), but is highly challenging due to its non-local property caused by the disparities among LF images. Although many deep neural networks (DNNs) have been developed for LF image SR and achieved continuousl…

Cited by 69PDFcodeScholar
2023

Masked Spatio-Temporal Structure Prediction for Self-supervised Learning on Point Cloud Videos

ICCV 2023poster

Recently, the community has made tremendous progress in developing effective methods for point cloud video understanding that learn from massive amounts of labeled data. However, annotating point cloud videos is usually notoriously expensive. Moreover, training via one or only a few traditional task…

Cited by 18PDFcodeScholar
2023

Monte Carlo Linear Clustering with Single-Point Supervision is Enough for Infrared Small Target Detection

ICCV 2023poster

Single-frame infrared small target (SIRST) detection aims at separating small targets from clutter backgrounds on infrared images. Recently, deep learning based methods have achieved promising performance on SIRST detection, but at the cost of a large amount of training data with expensive pixel-lev…

Cited by 56PDFcodeScholar
2023

Point Contrastive Prediction with Semantic Clustering for Self-Supervised Learning on Point Cloud Videos

ICCV 2023poster

We propose a unified point cloud video self-supervised learning framework for object-centric and scene-centric data. Previous methods commonly conduct representation learning at the clip or frame level and cannot well capture fine-grained semantics. Instead of contrasting the representations of clip…

Cited by 24PDFScholar
2023

PointCMP: Contrastive Mask Prediction for Self-Supervised Learning on Point Cloud Videos

CVPR 2023poster

Self-supervised learning can extract representations of good quality from solely unlabeled data, which is appealing for point cloud videos due to their high labelling cost. In this paper, we propose a contrastive mask prediction (PointCMP) framework for self-supervised learning on point cloud videos…

2023

Robust Multiview Point Cloud Registration With Reliable Pose Graph Initialization and History Reweighting

CVPR 2023poster

In this paper, we present a new method for the multiview registration of point cloud. Previous multiview registration methods rely on exhaustive pairwise registration to construct a densely-connected pose graph and apply Iteratively Reweighted Least Square (IRLS) on the pose graph to compute the sca…

2023

Semi-Weakly Supervised Object Kinematic Motion Prediction

CVPR 2023poster

Given a 3D object, kinematic motion prediction aims to identify the mobile parts as well as the corresponding motion parameters. Due to the large variations in both topological structure and geometric details of 3D objects, this remains a challenging task and the lack of large scale labeled data als…

Cited by 11SourcePDFScholar
2023

VAPCNet: Viewpoint-Aware 3D Point Cloud Completion

ICCV 2023poster

Most existing learning-based 3D point cloud completion methods ignore the fact that the completion process is highly coupled with the viewpoint of a partial scan. However, the various viewpoints of incompletely scanned objects in real-world applications are normally unknown and directly estimating t…

Cited by 12PDFcodeScholar
2022

3DAC: Learning Attribute Compression for Point Clouds

CVPR 2022poster

We study the problem of attribute compression for large-scale unstructured 3D point clouds. Through an in-depth exploration of the relationships between different encoding steps and different attribute channels, we introduce a deep compression network, termed 3DAC, to explicitly compress the attribu…

Cited by 40PDFcodeScholar
2022

Decoupling Makes Weakly Supervised Local Feature Better

CVPR 2022poster

Weakly supervised learning can help local feature methods to overcome the obstacle of acquiring a large-scale dataset with densely labeled correspondences. However, since weak supervision cannot distinguish the losses caused by the detection and description steps, directly conducting weakly supervis…

Cited by 61PDFcodeScholar
2022

Depth Estimation by Combining Binocular Stereo and Monocular Structured-Light

CVPR 2022poster

It is well known that the passive stereo system cannot adapt well to weak texture objects, e.g., white walls. However, these weak texture targets are very common in indoor environments. In this paper, we present a novel stereo system, which consists of two cameras (an RGB camera and an IR camera) an…

Cited by 18PDFcodeScholar
2022

Geometric Transformer for Fast and Robust Point Cloud Registration

CVPR 2022oral

We study the problem of extracting accurate correspondences for point cloud registration. Recent keypoint-free methods bypass the detection of repeatable keypoints which is difficult in low-overlap scenarios, showing great potential in registration. They seek correspondences over downsampled superpo…

Cited by 454PDFcodeScholar
2022

Learnable Lookup Table for Neural Network Quantization

CVPR 2022poster

Neural network quantization aims at reducing bit-widths of weights and activations for memory and computational efficiency. Since a linear quantizer (i.e., round(*) function) cannot well fit the bell-shaped distributions of weights and activations, many existing methods use pre-defined functions (e.…

Cited by 65PDFScholar
2022

Not All Points Are Equal: Learning Highly Efficient Point-Based Detectors for 3D LiDAR Point Clouds

CVPR 2022oral

We study the problem of efficient object detection of 3D LiDAR point clouds. To reduce the memory and computational cost, existing point-based pipelines usually adopt task-agnostic random sampling or farthest point sampling to progressively downsample input point clouds, despite the fact that not al…

Cited by 389PDFcodeScholar
2022

Occlusion-Aware Cost Constructor for Light Field Depth Estimation

CVPR 2022poster

Matching cost construction is a key step in light field (LF) depth estimation, but was rarely studied in the deep learning era. Recent deep learning-based LF depth estimation methods construct matching cost by sequentially shifting each sub-aperture image (SAI) with a series of predefined offsets, w…

Cited by 106PDFcodeScholar
2022

RayMVSNet: Learning Ray-Based 1D Implicit Fields for Accurate Multi-View Stereo

CVPR 2022poster

Learning-based multi-view stereo (MVS) has by far centered around 3D convolution on cost volumes. Due to the high computation and memory consumption of 3D CNN, the resolution of output depth is often considerably limited. Different from most existing works dedicated to adaptive refinement of cost vo…

Cited by 34PDFScholar
2022

SLFNet: A Stereo and LiDAR Fusion Network for Depth Completion

RA-L 2022

Acquiring dense and precise depth information in real time is highly demanded for robotic perception and automatic driving. Motivated by the complementary nature of stereo images and LiDAR point clouds, we propose an efficient stereo-LiDAR fusion network (SLFNet) to predict a dense depth map of a sc

Cited by 14SourceScholar
2022

SQN: Weakly-Supervised Semantic Segmentation of Large-Scale 3D Point Clouds

ECCV 2022poster

"Labelling point clouds fully is highly time-consuming and costly. As larger point cloud datasets containing billions of points become more common, we ask whether the full annotation is even necessary, demonstrating that existing baselines designed under a fully annotated assumption only degrade sli…

2021

Cgan-Net: Class-Guided Asymmetric Non-Local Network for Real-Time Semantic Segmentation

ICASSP 2021accepted

By introducing various non-local blocks to capture the long-range dependencies, remarkable progress has been achieved in semantic segmentation recently. However, the improvement in segmentation accuracy usually comes at the price of significant reductions in network efficiency, as non-local block us…

Cited by 0SourceScholar
2021

Exploring Sparsity in Image Super-Resolution for Efficient Inference

CVPR 2021poster

Current CNN-based super-resolution (SR) methods process all locations equally with computational resources being uniformly assigned in space. However, since missing details in low-resolution (LR) images mainly exist in regions of edges and textures, less computational resources are required for thos…

Cited by 311PDFcodeScholar
2021

Learning a Single Network for Scale-Arbitrary Super-Resolution

ICCV 2021poster

Recently, the performance of single image super-resolution (SR) has been significantly improved with powerful networks. However, these networks are developed for image SR with specific integer scale factors (e.g., x2/3/4), and cannot handle non-integer and asymmetric SR. In this paper, we propose to…

Cited by 145PDFScholar
2021

Sparse-to-Dense Feature Matching: Intra and Inter Domain Cross-Modal Learning in Domain Adaptation for 3D Semantic Segmentation

ICCV 2021poster

Domain adaptation is critical for success when confronting with the lack of annotations in a new domain. As the huge time consumption of labeling process on 3D point cloud, domain adaptation for 3D semantic segmentation is of great expectation. With the rise of multi-modal datasets, large amount of…

Cited by 66PDFcodeScholar
2021

SpinNet: Learning a General Surface Descriptor for 3D Point Cloud Registration

CVPR 2021poster

Extracting robust and general 3D local features is key to downstream tasks such as point cloud registration and reconstruction. Existing learning-based local descriptors are either sensitive to rotation transformations, or rely on classical handcrafted features which are neither general nor represen…

Cited by 380PDFcodeScholar
2021

Unsupervised Degradation Representation Learning for Blind Super-Resolution

CVPR 2021poster

Most existing CNN-based super-resolution (SR) methods are developed based on an assumption that the degradation is fixed and known (e.g., bicubic downsampling). However, these methods suffer a severe performance drop when the real degradation is different from their assumption. To handle various unk…

Cited by 429PDFcodeScholar
2020

ARPDR: An Accurate and Robust Pedestrian Dead Reckoning System for Indoor Localization on Handheld Smartphones

IROS 2020poster

The proliferation of mobile computing has prompted Pedestrian Dead Reckoning (PDR) to be one of the most attractive and promising indoor localization techniques for ubiquitous applications. The existing PDR approaches either suffer position drifts caused by accumulative errors or are sensitive to va…

Cited by 9SourceScholar
2020

RandLA-Net: Efficient Semantic Segmentation of Large-Scale Point Clouds

CVPR 2020oral

We study the problem of efficient semantic segmentation for large-scale 3D point clouds. By relying on expensive sampling techniques or computationally heavy pre/post-processing steps, most existing approaches are only able to be trained and operate over small-scale point clouds. In this paper, we i…

Cited by 2143PDFcodeScholar
2020

Spatial-Angular Interaction for Light Field Image Super-Resolution

ECCV 2020poster

Light field (LF) cameras record both intensity and directions of light rays, and capture scenes from a number of viewpoints. Both information within each perspective (i.e., spatial information) and among different perspectives (i.e., angular information) is beneficial to image super-resolution (SR).…

2019

Ground-to-Aerial Image Geo-Localization With a Hard Exemplar Reweighting Triplet Loss

ICCV 2019poster

The task of ground-to-aerial image geo-localization can be achieved by matching a ground view query image to a reference database of aerial/satellite images. It is highly challenging due to the dramatic viewpoint changes and unknown orientations. In this paper, we propose a novel in-batch reweightin…

Cited by 151PDFScholar
2019

Learning Parallax Attention for Stereo Image Super-Resolution

CVPR 2019poster

Stereo image pairs can be used to improve the performance of super-resolution (SR) since additional information is provided from a second viewpoint. However, it is challenging to incorporate this information for SR since disparities between stereo images vary significantly. In this paper, we propose…

Cited by 327PDFcodeScholar
2018

Learning for Disparity Estimation Through Feature Constancy

CVPR 2018poster

Stereo matching algorithms usually consist of four steps, including matching cost calculation, matching cost aggregation, disparity calculation, and disparity refinement. Existing CNN-based methods only adopt CNN to solve parts of the four steps, or use different networks to deal with different step…