← Search

Bailan Feng

17 accepted papers

2026

GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation

CVPR 2026

Physics-aware driving world model is essential for drive planning, out-of-distribution data synthesis, and closed-loop evaluation. However, existing methods often rely on a single diffusion model to directly map driving actions to videos, which makes learning difficult and leads to physically incons

Cited by 0SourceScholar
2025

Efficient 3D Perception on Multi-Sweep Point Cloud with Gumbel Spatial Pruning

ICRA 2025

This paper studies point cloud perception within outdoor environments. Existing methods face limitations in recognizing objects located at a distance or occluded, due to the sparse nature of outdoor point clouds. In this work, we observe a significant mitigation of this problem by accumulating multi

Cited by 1SourceScholar
2025

Reliable and Calibrated Semantic Occupancy Prediction by Hybrid Uncertainty Learning

IJCAI 2025

Vision-centric semantic occupancy prediction plays a crucial role in autonomous driving, which requires accurate and reliable predictions from low-cost sensors. Although having notably narrowed the accuracy gap with LiDAR, there is still few research effort to explore the reliability and calibration

Cited by 0SourcePDFScholar
2025

TurboVSR: Fantastic Video Upscalers and Where to Find Them

ICCV 2025poster

Diffusion-based generative models have demonstrated exceptional promise in the video super-resolution (VSR) task, achieving a substantial advancement in detail generation relative to prior methods. However, these approaches face significant computational efficiency challenges. For instance, current…

Cited by 0SourcePDFScholar
2024

"Segment, Lift and Fit: Automatic 3D Shape Labeling from 2D Prompts"

ECCV 2024poster

"This paper proposes an algorithm for automatically labeling 3D objects from 2D point or box prompts, especially focusing on applications in autonomous driving. Unlike previous arts, our auto-labeler predicts 3D shapes instead of bounding boxes and does not require training on a specific dataset. We…

2024

OccGen: Generative Multi-modal 3D Occupancy Prediction for Autonomous Driving

ECCV 2024poster

"Existing 3D semantic occupancy prediction methods typically treat the task as a one-shot 3D voxel-wise segmentation problem, focusing on a single-step mapping between the inputs and occupancy maps, which limits their ability to refine and complete local regions gradually. In this paper, we introduc…

Cited by 21SourcePDFScholar
2024

OctOcc: High-Resolution 3D Occupancy Prediction with Octree

AAAI 2024technical

3D semantic occupancy has garnered considerable attention due to its abundant structural information encompassing the entire scene in autonomous driving. However, existing 3D occupancy prediction methods contend with the constraint of low-resolution 3D voxel features arising from the limitation of c…

Cited by 7SourcePDFScholar
2024

SparseOcc: Rethinking Sparse Latent Representation for Vision-Based Semantic Occupancy Prediction

CVPR 2024poster

Vision-based perception for autonomous driving requires an explicit modeling of a 3D space where 2D latent representations are mapped and subsequent 3D operators are applied. However operating on dense latent spaces introduces a cubic time and space complexity which limits scalability in terms of pe…

Cited by 42SourcePDFScholar
2024

VEON: Vocabulary-Enhanced Occupancy Prediction

ECCV 2024poster

"Perceiving the world as 3D occupancy supports embodied agents to avoid collision with any types of obstacle. While open-vocabulary image understanding has prospered recently, how to bind the predicted 3D occupancy grids with open-world semantics still remains under-explored due to limited open-worl…

2023

AttentionShift: Iteratively Estimated Part-Based Attention Map for Pointly Supervised Instance Segmentation

CVPR 2023poster

Pointly supervised instance segmentation (PSIS) learns to segment objects using a single point within the object extent as supervision. Challenged by the non-negligible semantic variance between object parts, however, the single supervision point causes semantic bias and false segmentation. In this…

Cited by 12SourcePDFScholar
2022

CF-DETR: Coarse-to-Fine Transformers for End-to-End Object Detection

AAAI 2022technical

The recently proposed DEtection TRansformer (DETR) achieves promising performance for end-to-end object detection. However, it has relatively lower detection performance on small objects and suffers from slow convergence. This paper observed that DETR performs surprisingly well even on small objects…

Cited by 41SourcePDFScholar
2022

End-to-End Weakly Supervised Object Detection with Sparse Proposal Evolution

ECCV 2022poster

"Conventional methods for weakly supervised object detection (WSOD) typically enumerate dense proposals and select the discriminative proposals as objects. However, these two-stage “enumerate-and-select” methods suffer object feature ambiguity brought by dense proposals and low detection efficiency…

2022

PANDORA: A Panoramic Detection Dataset for Object with Orientation

ECCV 2022poster

"Panoramic images have become increasingly popular as omnidirectional panoramic technology has advanced. Many datasets and works resort to object detection to better understand the content of the panoramic image. These datasets and detectors use a Bounding Field of View (BFoV) as a bounding box in p…

2022

Unbiased IoU for Spherical Image Object Detection

AAAI 2022technical

As one of the fundamental components of object detection, intersection-over-union (IoU) calculations between two bounding boxes play an important role in samples selection, NMS operation and evaluation of object detection algorithms. This procedure is well-defined and solved for planar images, while…

Cited by 12SourcePDFScholar
2018

A Deep Learning Based No-Reference Image Quality Assessment Model for Single-Image Super-Resolution

ICASSP 2018accepted

Single-image super-resolution (SISR) is a very important and classic problem of the computer vision community. Although a lot of SISR methods have been proposed, few studies have been conducted to address the quality assessment of SISR methods. In this paper, we proposed a deep learning based no-ref…

Cited by 0SourceScholar
2018

HNSR: Highway Networks Based Deep Convolutional Neural Networks Model for Single Image Super-Resolution

ICASSP 2018accepted

Convolutional neural networks (CNNs) have been widely used in computer vision community. Single image super-resolution (SISR) is a classic computer vision problem, which aims to output a high-resolution image from a low-resolution one. In recent years, CNNs-based SISR methods emerged and achieved a…

Cited by 0SourceScholar