← Search

Hong Cai

22 accepted papers

2026

LidarPainter: One-Step Away from Any Lidar View to Novel Guidance

AAAI 2026technical

Dynamic driving scene reconstruction is of great importance in fields like digital twin system and autonomous driving simulation. However, unacceptable degradation occurs when the view deviates from the input trajectory, leading to corrupted background and vehicle models. To improve reconstruction q

Cited by 0SourcePDFScholar
2026

RoCA: Robust Cross-Domain End-to-End Autonomous Driving

ICML 2026poster

End-to-end (E2E) autonomous driving has recently emerged as a new paradigm, offering significant potential. However, few studies have looked into the practical challenge of deployment across domains (e.g., cities). Although several works have incorporated Large Language Models (LLMs) to leverage the…

Cited by 8SourceScholar
2025

Distilling Multi-modal Large Language Models for Autonomous Driving

CVPR 2025poster

Autonomous driving demands safe motion planning, especially in critical "long-tail" scenarios. Recent end-to-end autonomous driving systems leverage large language models (LLMs) as planners to improve generalizability to rare events. However, using LLMs at test time introduces high computational cos…

Cited by 4SourcePDFScholar
2025

ODG: Occupancy Prediction Using Dual Gaussians

NeurIPS 2025poster

Occupancy prediction infers fine-grained 3D geometry and semantics from camera images of the surrounding environment, making it a critical perception task for autonomous driving. Existing methods either adopt dense grids as scene representation which is difficult to scale to high resolution, or lear…

Cited by 0SourceScholar
2025

PADRe: A Unifying Polynomial Attention Drop-in Replacement for Efficient Vision Transformer

ICLR 2025poster

We present Polynomial Attention Drop-in Replacement (PADRe), a novel and unifying framework designed to replace the conventional self-attention mechanism in transformer models. Notably, several recent alternative attention mechanisms, including Hyena, Mamba, SimA, Conv2Former, and Castling-ViT, can…

Cited by 2SourcePDFScholar
2024

DeCoTR: Enhancing Depth Completion with 2D and 3D Attentions

CVPR 2024poster

In this paper we introduce a novel approach that harnesses both 2D and 3D attentions to enable highly accurate depth completion without requiring iterative spatial propagations. Specifically we first enhance a baseline convolutional depth completion model by applying attention to 2D features in the…

Cited by 3SourcePDFScholar
2024

FutureDepth: Learning to Predict the Future Improves Video Depth Estimation

ECCV 2024poster

"In this paper, we propose a novel video depth estimation approach, , which enables the model to implicitly leverage multi-frame and motion cues to improve depth estimation by making it learn to predict the future at training. More specifically, we propose a future prediction network, F-Net, which t…

Cited by 5SourcePDFScholar
2024

OCAI: Improving Optical Flow Estimation by Occlusion and Consistency Aware Interpolation

CVPR 2024poster

The scarcity of ground-truth labels poses one major challenge in developing optical flow estimation models that are both generalizable and robust. While current methods rely on data augmentation they have yet to fully exploit the rich information available in labeled video sequences. We propose OCAI…

Cited by 3SourcePDFScholar
2023

4D Panoptic Segmentation as Invariant and Equivariant Field Prediction

ICCV 2023poster

In this paper, we develop rotation-equivariant neural networks for 4D panoptic segmentation. 4D panoptic segmentation is a benchmark task for autonomous driving that requires recognizing semantic classes and object instances on the road based on LiDAR scans, as well as assigning temporally consisten…

Cited by 18PDFScholar
2023

DejaVu: Conditional Regenerative Learning To Enhance Dense Prediction

CVPR 2023poster

We present DejaVu, a novel framework which leverages conditional image regeneration as additional supervision during training to improve deep networks for dense prediction tasks such as segmentation, depth estimation, and surface normal prediction. First, we apply redaction to the input image, which…

Cited by 10SourcePDFScholar
2023

DistractFlow: Improving Optical Flow Estimation via Realistic Distractions and Pseudo-Labeling

CVPR 2023poster

We propose a novel data augmentation approach, DistractFlow, for training optical flow estimation models by introducing realistic distractions to the input frames. Based on a mixing ratio, we combine one of the frames in the pair with a distractor image depicting a similar domain, which allows for i…

2023

Factorized Inverse Path Tracing for Efficient and Accurate Material-Lighting Estimation

ICCV 2023oral

Inverse path tracing has recently been applied to joint material and lighting estimation, given geometry and multi-view HDR observations of an indoor scene. However, it has two major limitations: path tracing is expensive to compute, and ambiguities exist between reflection and emission. Our Facto…

Cited by 14PDFcodeScholar
2023

MAMo: Leveraging Memory and Attention for Monocular Video Depth Estimation

ICCV 2023poster

We propose MAMo, a novel memory and attention framework for monocular video depth estimation. MAMo can augment and improve any single-image depth estimation networks into video depth estimation models, enabling them to take advantage of the temporal information to predict more accurate depth. In MAM…

Cited by 17PDFScholar
2023

OpenShape: Scaling Up 3D Shape Representation Towards Open-World Understanding

NeurIPS 2023poster

We introduce OpenShape, a method for learning multi-modal joint representations of text, image, and point clouds. We adopt the commonly used multi-modal contrastive learning framework for representation alignment, but with a specific focus on scaling up 3D representations to enable open-world 3D sha…

Cited by 126SourcePDFScholar
2023

PartSLIP: Low-Shot Part Segmentation for 3D Point Clouds via Pretrained Image-Language Models

CVPR 2023poster

Generalizable 3D part segmentation is important but challenging in vision and robotics. Training deep models via conventional supervised methods requires large-scale 3D datasets with fine-grained part annotations, which are costly to collect. This paper explores an alternative way for low-shot part…

2023

Self-Supervised Geometric Correspondence for Category-Level 6D Object Pose Estimation in the Wild

ICLR 2023poster

While 6D object pose estimation has wide applications across computer vision and robotics, it remains far from being solved due to the lack of annotations. The problem becomes even more challenging when moving to category-level 6D pose, which requires generalization to unseen instances. Current appr…

2023

Transadapt: A Transformative Framework for Online Test Time Adaptive Semantic Segmentation

ICASSP 2023accepted

Test-time adaptive (TTA) semantic segmentation adapts a source pre-trained image semantic segmentation model to unlabeled batches of target domain test images, different from real-world, where samples arrive one-by-one in an online fashion. To tackle online settings, we propose TransAdapt, a framewo…

Cited by 0SourceScholar
2022

Learning Implicit Feature Alignment Function for Semantic Segmentation

ECCV 2022poster

"Integrating high-level context information with low-level details is of central importance in semantic segmentation. Towards this end, most existing segmentation models apply bilinear up-sampling and convolutions to feature maps of different scales, and then align them at the same resolution. Howev…

2022

Panoptic, Instance and Semantic Relations: A Relational Context Encoder To Enhance Panoptic Segmentation

CVPR 2022poster

This paper presents a novel framework to integrate both semantic and instance contexts for panoptic segmentation. In existing works, it is common to use a shared backbone to extract features for both things (countable classes such as vehicles) and stuff (uncountable classes such as roads). This, how…

Cited by 16PDFScholar
2018

PieAPP: Perceptual Image-Error Assessment Through Pairwise Preference

CVPR 2018poster

The ability to estimate the perceptual error between images is an important problem in computer vision with many applications. Although it has been studied extensively, however, no method currently exists that can robustly predict visual differences like humans. Some previous approaches used hand-co…