← Search

Lingjun Zhao

8 accepted papers

2026

GaussianFormer3D: Multi-Modal Gaussian-Based Semantic Occupancy Prediction with 3D Deformable Attention

ICRA 2026poster

3D semantic occupancy prediction is essential for achieving safe, reliable autonomous driving and robotic navigation. Compared to camera-only perception systems, multi-modal pipelines, especially LiDAR-camera fusion methods, can produce more accurate and fine-grained predictions. Although voxel-base…

2025

A Necessary Step toward Faithfulness: Measuring and Improving Consistency in Free-Text Explanations

EMNLP 2025

Faithful free-text explanations are important to ensure transparency in high-stakes AI decision-making contexts, but they are challenging to generate by language models and assess by humans. In this paper, we present a measure for Prediction-EXplanation (PEX) consistency, by extending the concept of

Cited by 0SourcePDFScholar
2025

Can Hallucination Correction Improve Video-Language Alignment?

ACL 2025finding

Large Vision-Language Models often generate hallucinated content that is not grounded in its visual inputs. While prior work focuses on mitigating hallucinations, we instead explore leveraging hallucination correction as a training objective to improve video-language alignment. We introduce HACA, a…

Cited by 0SourcePDFScholar
2024

CRKD: Enhanced Camera-Radar Object Detection with Cross-modality Knowledge Distillation

CVPR 2024poster

In the field of 3D object detection for autonomous driving LiDAR-Camera (LC) fusion is the top-performing sensor configuration. Still LiDAR is relatively high cost which hinders adoption of this technology for consumer automobiles. Alternatively camera and radar are commonly deployed on vehicles alr…

2024

LiRaFusion: Deep Adaptive LiDAR-Radar Fusion for 3D Object Detection

ICRA 2024poster

We propose LiRaFusion to tackle LiDAR-radar fusion for 3D object detection to fill the performance gap of existing LiDAR-radar detectors. To improve the feature extraction capabilities from these two modalities, we design an early fusion module for joint voxel feature encoding, and a middle fusion m…

Cited by 15SourcecodeScholar
2024

Successfully Guiding Humans with Imperfect Instructions by Highlighting Potential Errors and Suggesting Corrections

EMNLP 2024main

Language models will inevitably err in situations with which they are unfamiliar. However, by effectively communicating uncertainties, they can still guide humans toward making sound decisions in those contexts. We demonstrate this idea by developing HEAR, a system that can successfully guide humans…

2023

Define, Evaluate, and Improve Task-Oriented Cognitive Capabilities for Instruction Generation Models

ACL 2023findings

Recent work studies the cognitive capabilities of language models through psychological tests designed for humans. While these studies are helpful for understanding the general capabilities of these models, there is no guarantee that a model possessing sufficient capabilities to pass those tests wou…