← Search

Shizhong Han

10 accepted papers

2026

RoCA: Robust Cross-Domain End-to-End Autonomous Driving

ICML 2026poster

End-to-end (E2E) autonomous driving has recently emerged as a new paradigm, offering significant potential. However, few studies have looked into the practical challenge of deployment across domains (e.g., cities). Although several works have incorporated Large Language Models (LLMs) to leverage the…

Cited by 8SourceScholar
2025

Distilling Multi-modal Large Language Models for Autonomous Driving

CVPR 2025poster

Autonomous driving demands safe motion planning, especially in critical "long-tail" scenarios. Recent end-to-end autonomous driving systems leverage large language models (LLMs) as planners to improve generalizability to rare events. However, using LLMs at test time introduces high computational cos…

Cited by 4SourcePDFScholar
2025

ODG: Occupancy Prediction Using Dual Gaussians

NeurIPS 2025poster

Occupancy prediction infers fine-grained 3D geometry and semantics from camera images of the surrounding environment, making it a critical perception task for autonomous driving. Existing methods either adopt dense grids as scene representation which is difficult to scale to high resolution, or lear…

Cited by 0SourceScholar
2025

PADRe: A Unifying Polynomial Attention Drop-in Replacement for Efficient Vision Transformer

ICLR 2025poster

We present Polynomial Attention Drop-in Replacement (PADRe), a novel and unifying framework designed to replace the conventional self-attention mechanism in transformer models. Notably, several recent alternative attention mechanisms, including Hyena, Mamba, SimA, Conv2Former, and Castling-ViT, can…

Cited by 2SourcePDFScholar
2024

FutureDepth: Learning to Predict the Future Improves Video Depth Estimation

ECCV 2024poster

"In this paper, we propose a novel video depth estimation approach, , which enables the model to implicitly leverage multi-frame and motion cues to improve depth estimation by making it learn to predict the future at training. More specifically, we propose a future prediction network, F-Net, which t…

Cited by 5SourcePDFScholar
2023

4D Panoptic Segmentation as Invariant and Equivariant Field Prediction

ICCV 2023poster

In this paper, we develop rotation-equivariant neural networks for 4D panoptic segmentation. 4D panoptic segmentation is a benchmark task for autonomous driving that requires recognizing semantic classes and object instances on the road based on LiDAR scans, as well as assigning temporally consisten…

Cited by 18PDFScholar
2023

OpenShape: Scaling Up 3D Shape Representation Towards Open-World Understanding

NeurIPS 2023poster

We introduce OpenShape, a method for learning multi-modal joint representations of text, image, and point clouds. We adopt the commonly used multi-modal contrastive learning framework for representation alignment, but with a specific focus on scaling up 3D representations to enable open-world 3D sha…

Cited by 126SourcePDFScholar
2023

PartSLIP: Low-Shot Part Segmentation for 3D Point Clouds via Pretrained Image-Language Models

CVPR 2023poster

Generalizable 3D part segmentation is important but challenging in vision and robotics. Training deep models via conventional supervised methods requires large-scale 3D datasets with fine-grained part annotations, which are costly to collect. This paper explores an alternative way for low-shot part…

2018

Optimizing Filter Size in Convolutional Neural Networks for Facial Action Unit Recognition

CVPR 2018poster

Recognizing facial action units (AUs) during spontaneous facial displays is a challenging problem. Most recently, Convolutional Neural Networks (CNNs) have shown promise for facial AU recognition, where predefined and fixed convolution filter sizes are employed. In order to achieve the best performa…

Cited by 84SourcePDFScholar
2016

Incremental Boosting Convolutional Neural Network for Facial Action Unit Recognition

NeurIPS 2016poster

Recognizing facial action units (AUs) from spontaneous facial expressions is still a challenging problem. Most recently, CNNs have shown promise on facial AU recognition. However, the learned CNNs are often overfitted and do not generalize well to unseen subjects due to limited AU-coded training ima…

Cited by 117SourcePDFScholar