← Search

Liu Ren

41 accepted papers

2026

3DGEER: 3D Gaussian Rendering Made Exact and Efficient for Generic Cameras

ICLR 2026poster

3D Gaussian Splatting (3DGS) achieves an appealing balance between rendering quality and efficiency, but relies on approximating 3D Gaussians as 2D projections—an assumption that degrades accuracy, especially under generic large field-of-view (FoV) cameras. Despite recent extensions, no prior work…

Cited by 6SourcecodeScholar
2026

MTA: Multimodal Task Alignment for BEV Perception and Captioning

CVPR 2026

Bird's eye view (BEV)-based 3D perception plays a crucial role in autonomous driving applications. The rise of large language models has spurred interest in BEV-based captioning to understand object behavior in the surrounding environment. However, existing approaches treat perception and captioning

Cited by 6SourceScholar
2026

NexusFlow: Unifying Disparate Tasks under Partial Supervision via Invertible Flow Networks

CVPR 2026

Partially Supervised Multi-Task Learning (PS-MTL) aims to leverage knowledge across tasks when annotations are incomplete. Existing approaches, however, have largely focused on the simpler setting of homogeneous, dense prediction tasks, leaving the more realistic challenge of learning from structura

Cited by 0SourcecodeScholar
2026

No Calibration, No Depth, No Problem: Cross-Sensor View Synthesis with 3D Consistency

CVPR 2026

We present the first study of cross-sensor view synthesis across different modalities. We examine a practical, fundamental, yet widely overlooked problem: getting aligned RGB-X data, where most RGB-X prior work assumes such pairs exist and focuses on modality fusion, but it empirically requires huge

Cited by 0SourceScholar
2026

Pantheon360: Taming Digital Twin Generation via 3D-Aware 360deg Video Diffusion

CVPR 2026

Generating complete digital twins from videos requires precise camera control, global scene coverage, and strict spatial-temporal consistency--constraints that remain challenging for perspective video generators due to their limited field of view (FoV). Their narrow FoV forces long or multi-view tra

Cited by 0SourceScholar
2025

AdaWM: Adaptive World Model based Planning for Autonomous Driving

ICLR 2025poster

World model based reinforcement learning (RL) has emerged as a promising approach for autonomous driving, which learns a latent dynamics model and uses it to train a planning policy. To speed up the learning process, the pretrain-finetune paradigm is often used, where online RL is initialized by a…

Cited by 1SourcePDFScholar
2025

BEVDiffuser: Plug-and-Play Diffusion Model for BEV Denoising with Ground-Truth Guidance

CVPR 2025highlight

Bird's-eye-view (BEV) representations play a crucial role in autonomous driving tasks. Despite recent advancements in BEV generation, inherent noise, stemming from sensor limitations and the learning process, remains largely unaddressed, resulting in suboptimal BEV representations that adversely imp…

Cited by 2SourcePDFScholar
2025

CHARM3R: Towards Unseen Camera Height Robust Monocular 3D Detector

ICCV 2025poster

Monocular 3D object detectors, while effective on data from one ego camera height, struggle with unseen or out-of-distribution camera heights. Existing methods often rely on Plucker embeddings, image transformations or data augmentation. This paper takes a step towards this understudied problem by i…

2025

Depth Any Camera: Zero-Shot Metric Depth Estimation from Any Camera

CVPR 2025poster

While recent depth foundation models exhibit strong zero-shot generalization, achieving accurate metric depth across diverse camera types--particularly those with large fields of view (FoV) such as fisheye and 360-degree cameras--remains a significant challenge. This paper presents Depth Any Camera…

2025

FogROS2-PLR: Probabilistic Latency-Reliability for Cloud Robotics

ICRA 2025

Cloud robotics enables robots to offload computationally intensive tasks to cloud servers for performance, cost, and ease of management. However, the network and cloud computing infrastructure are not designed for reliable timing guarantees, due to fluctuating Quality-of-Service (QoS). In this work,

Cited by 6SourcecodeScholar
2025

Open Ad-hoc Categorization with Contextualized Feature Learning

CVPR 2025poster

Adaptive categorization of visual scenes is essential for AI agents to handle changing tasks. Unlike fixed common categories for plants or animals, ad-hoc categories, such as things to sell at a garage sale, are created dynamically to achieve specific tasks. We study open ad-hoc categorization, wh…

2025

ProSAM: Enhancing the Robustness of SAM-based Visual Reference Segmentation with Probabilistic Prompts

ICCV 2025poster

The recent advancements in large foundation models have driven the success of open-set image segmentation, a task focused on segmenting objects beyond predefined categories. Among various prompt types (such as points, boxes, texts, and visual references), visual reference segmentation stands out for…

Cited by 0SourcePDFScholar
2025

SMART: Advancing Scalable Map Priors for Driving Topology Reasoning

ICRA 2025

Topology reasoning is crucial for autonomous driving as it enables comprehensive understanding of connec-tivity and relationships between lanes and traffic elements. While recent approaches have shown success in perceiving driving topology using vehicle-mounted sensors, their scalability is hindered

Cited by 9SourcecodeScholar
2025

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads

ICCV 2025poster

Vision foundation models (VFMs) have demonstrated remarkable performance across a wide range of downstream tasks. While several VFM adapters have shown promising results by leveraging the prior knowledge of VFMs, we identify two inefficiencies in these approaches. First, the interaction between conv…

2024

Behind the Veil: Enhanced Indoor 3D Scene Reconstruction with Occluded Surfaces Completion

CVPR 2024poster

In this paper we present a novel indoor 3D reconstruction method with occluded surface completion given a sequence of depth readings. Prior state-of-the-art (SOTA) methods only focus on the reconstruction of the visible areas in a scene neglecting the invisible areas due to the occlusions e.g. the c…

Cited by 1SourcePDFScholar
2024

CLIP-BEVFormer: Enhancing Multi-View Image-Based BEV Detector with Ground Truth Flow

CVPR 2024poster

Autonomous driving stands as a pivotal domain in computer vision shaping the future of transportation. Within this paradigm the backbone of the system plays a crucial role in interpreting the complex environment. However a notable challenge has been the loss of clear supervision when it comes to Bir…

Cited by 11SourcePDFScholar
2024

Enhancing Online Road Network Perception and Reasoning with Standard Definition Maps

IROS 2024poster

Autonomous driving for urban and highway driving applications often requires High Definition (HD) maps to generate a navigation plan. Nevertheless, various challenges arise when generating and maintaining HD maps at scale. While recent online mapping methods have started to emerge, their performance…

Cited by 5SourcecodeScholar
2024

FogROS2-FT: Fault Tolerant Cloud Robotics

IROS 2024poster

Cloud robotics enables robots to offload complex computational tasks to cloud servers for performance and ease of management. However, cloud compute can be costly, cloud services can suffer occasional downtime, and connectivity between the robot and cloud can be prone to variations in network Qualit…

Cited by 0SourceScholar
2024

FogROS2-LS: A Location-Independent Fog Robotics Framework for Latency Sensitive ROS2 Applications

ICRA 2024poster

In Cloud Robotics, long system latency due to varying network conditions can cause instability and collisions. However, this can be minimized in the almost univeral case where there are multiple sources available for cloud servers. By extending anycast routing, we introduce FogROS2-Latency-Sensitive…

Cited by 8SourceScholar
2024

Hyp-OW: Exploiting Hierarchical Structure Learning with Hyperbolic Distance Enhances Open World Object Detection

AAAI 2024technical

Open World Object Detection (OWOD) is a challenging and realistic task that extends beyond the scope of standard Object Detection task. It involves detecting both known and unknown objects while integrating learned knowledge for future tasks. However, the level of "unknownness" varies significantly…

Cited by 23SourcePDFScholar
2024

Lifelong LERF: Local 3D Semantic Inventory Monitoring Using FogROS2

ICRA 2024poster

Inventory monitoring in homes, factories, and retail stores relies on maintaining data despite objects being swapped, added, removed, or moved. We introduce Lifelong LERF, a method that allows a mobile robot with minimal compute to jointly optimize a dense language and geometric representation of it…

Cited by 6SourceScholar
2024

MetaAT: Active Testing for Label-Efficient Evaluation of Dense Recognition Tasks

ECCV 2024poster

"In this study, we investigate the task of active testing for label-efficient evaluation, which aims to estimate a model’s performance on an unlabeled test dataset with a limited annotation budget. Previous approaches relied on deep ensemble models to identify highly informative instances for labeli…

Cited by 0SourcePDFScholar
2024

SUP-NeRF: A Streamlined Unification of Pose Estimation and NeRF for Monocular 3D Object Reconstruction

ECCV 2024poster

"Monocular 3D reconstruction for categorical objects heavily relies on accurately perceiving each object’s pose. While gradient-based optimization in a NeRF framework updates the initial pose, this paper highlights that scale-depth ambiguity in monocular object reconstruction causes failures when th…

2024

SeaBird: Segmentation in Bird's View with Dice Loss Improves Monocular 3D Detection of Large Objects

CVPR 2024poster

Monocular 3D detectors achieve remarkable performance on cars and smaller objects. However their performance drops on larger objects leading to fatal accidents. Some attribute the failures to training data scarcity or the receptive field requirements of large objects. In this paper we highlight this…

2024

TCLC-GS: Tightly Coupled LiDAR-Camera Gaussian Splatting for Autonomous Driving

ECCV 2024poster

"Most 3D Gaussian Splatting (3D-GS) based methods for urban scenes initialize 3D Gaussians directly with 3D LiDAR points, which not only underutilizes LiDAR data capabilities but also overlooks the potential advantages of fusing LiDAR with camera data. In this paper, we design a novel tightly couple…

2024

USE: Universal Segment Embeddings for Open-Vocabulary Image Segmentation

CVPR 2024poster

The open-vocabulary image segmentation task involves partitioning images into semantically meaningful segments and classifying them with flexible text-defined categories. The recent vision-based foundation models such as the Segment Anything Model (SAM) have shown superior performance in generating…

Cited by 5SourcePDFScholar
2024

VLP: Vision Language Planning for Autonomous Driving

CVPR 2024poster

Autonomous driving is a complex and challenging task that aims at safe motion planning through scene understanding and reasoning. While vision-only autonomous driving methods have recently achieved notable performance through enhanced scene understanding several key issues including lack of reasonin…

Cited by 54SourcePDFScholar
2023

3D Copy-Paste: Physically Plausible Object Insertion for Monocular 3D Detection

NeurIPS 2023poster

A major challenge in monocular 3D object detection is the limited diversity and quantity of objects in real datasets. While augmenting real scenes with virtual objects holds promise to improve both the diversity and quantity of the objects, it remains elusive due to the lack of an effective 3D objec…

2023

CLIP-S4: Language-Guided Self-Supervised Semantic Segmentation

CVPR 2023highlight

Existing semantic segmentation approaches are often limited by costly pixel-wise annotations and predefined classes. In this work, we present CLIP-S^4 that leverages self-supervised pixel representation learning and vision-language models to enable various semantic segmentation tasks (e.g., unsuperv…

Cited by 58SourcePDFScholar
2023

CoAug: Combining Augmentation of Labels and Labelling Rules

ACL 2023findings

Collecting labeled data for Named Entity Recognition (NER) tasks is challenging due to the high cost of manual annotations. Instead, researchers have proposed few-shot self-training and rule-augmentation techniques to minimize the reliance on large datasets. However, inductive biases and restricted…

2023

GradOrth: A Simple yet Efficient Out-of-Distribution Detection with Orthogonal Projection of Gradients

NeurIPS 2023poster

Detecting out-of-distribution (OOD) data is crucial for ensuring the safe deployment of machine learning models in real-world applications. However, existing OOD detection approaches primarily rely on the feature maps or the full gradient space information to derive OOD scores neglecting the role of…

Cited by 16SourcePDFScholar
2023

Long-Distance Gesture Recognition Using Dynamic Neural Networks

IROS 2023poster

Gestures form an important medium of communication between humans and machines. An overwhelming majority of existing gesture recognition methods are tailored to a scenario where humans and machines are located very close to each other. This short-distance assumption does not hold true for several ty…

Cited by 2SourceScholar
2023

UP-DP: Unsupervised Prompt Learning for Data Pre-Selection with Vision-Language Models

NeurIPS 2023poster

In this study, we investigate the task of data pre-selection, which aims to select instances for labeling from an unlabeled dataset through a single pass, thereby optimizing performance for undefined downstream tasks with a limited annotation budget. Previous approaches to data pre-selection relied…

Cited by 7SourcePDFScholar
2022

OmniFusion: 360 Monocular Depth Estimation via Geometry-Aware Fusion

CVPR 2022oral

A well-known challenge in applying deep-learning methods to omnidirectional images is spherical distortion. In dense regression tasks such as depth estimation, where structural details are required, using a vanilla CNN layer on the distorted 360 image results in undesired information loss. In this p…

Cited by 91PDFcodeScholar
2022

Self-supervised Semantic Segmentation Grounded in Visual Concepts

IJCAI 2022poster

Unsupervised semantic segmentation requires assigning a label to every pixel without any human annotations. Despite recent advances in self-supervised representation learning for individual images, unsupervised semantic segmentation with pixel-level representations is still a challenging task and r…

Cited by 10SourcePDFScholar
2022

Symmetry and Uncertainty-Aware Object SLAM for 6DoF Object Pose Estimation

CVPR 2022poster

We propose a keypoint-based object-level SLAM framework that can provide globally consistent 6DoF pose estimates for symmetric and asymmetric objects alike. To the best of our knowledge, our system is among the first to utilize the camera pose information from SLAM to provide prior knowledge for tra…

Cited by 51PDFcodeScholar
2021

Unsupervised Discriminative Learning of Sounds for Audio Event Classification

ICASSP 2021accepted

Recent progress in network-based audio event classification has shown the benefit of pre-training models on visual data such as ImageNet. While this process allows knowledge transfer across different domains, training a model on large-scale visual datasets is time consuming. On several audio event c…

Cited by 5SourceScholar
2020

Analytic Combined IMU Integration (ACI2) For Visual Inertial Navigation

ICRA 2020poster

Batch optimization based inertial measurement unit (IMU) and visual sensor fusion enables high rate localization for many robotic tasks. However, it remains a challenge to ensure that the batch optimization is computationally efficient while being consistent for high rate IMU measurements without ma…

Cited by 30SourceScholar