← Search

Yung-Hui Li

13 accepted papers

2026

Perceiving the Near, Reasoning the Distant: Coherent Long-Horizon Trajectory Prediction for Autonomous Driving

CVPR 2026

Reliable long-horizon trajectory prediction requires both high positional accuracy and physically plausible temporal motion consistency. However, existing methods suffer from two fundamental limitations. First, they overlook the inherent difference in prediction logic: near-future trajectories are p

Cited by 0SourcecodeScholar
2026

When Privacy Meets Recovery: The Overlooked Half of Surrogate-Driven Privacy Preservation for MLLM Editing

AAAI 2026technical

Privacy leakage in Multimodal Large Language Models (MLLMs) has long been an intractable problem. Existing studies, though effectively obscure private information in MLLMs, often overlook the evaluation of authenticity and recovery quality of user privacy. To this end, this work uniquely focuses on

Cited by 0SourcePDFScholar
2025

Global Regulation and Excitation via Attention Tuning for Stereo Matching

ICCV 2025poster

Stereo matching achieves significant progress with iterative algorithms like RAFT-Stereo and IGEV-Stereo. However, these methods struggle in ill-posed regions with occlusions, textureless, or repetitive patterns, due to a lack of global context and geometric information for effective iterative refin…

2025

Memory-Augmented Re-Completion for 3D Semantic Scene Completion

AAAI 2025technical

Semantic Scene Completion (SSC) aims to reconstruct a 3D voxel representation occupied by semantic classes based on ordinary inputs such as 2D RGB images, depth maps, or point clouds. Given the cost-effective and promising applications in autonomous driving, camera-based SSC has attracted considerab…

2025

ModeSeq: Taming Sparse Multimodal Motion Prediction with Sequential Mode Modeling

CVPR 2025poster

Anticipating the multimodality of future events lays the foundation for safe autonomous driving. However, multimodal motion prediction for traffic agents has been clouded by the lack of multimodal ground truth. Existing works predominantly adopt the winner-take-all training strategy to tackle this c…

Cited by 0SourcePDFScholar
2025

PAVLM: Advancing Point Cloud based Affordance Understanding Via Vision-Language Model

IROS 2025

Affordance understanding, the task of identifying actionable regions on 3D objects, plays a vital role in allowing robotic systems to engage with and operate within the physical world. Although Visual Language Models (VLMs) have excelled in high-level reasoning and long-horizon planning for robotic

Cited by 6SourcecodeScholar
2024

BehaviorGPT: Smart Agent Simulation for Autonomous Driving with Next-Patch Prediction

NeurIPS 2024poster

Simulating realistic behaviors of traffic agents is pivotal for efficiently validating the safety of autonomous driving systems. Existing data-driven simulators primarily use an encoder-decoder architecture to encode the historical trajectories before decoding the future. However, the heterogeneity…

Cited by 18SourcePDFScholar
2024

CCTR: Calibrating Trajectory Prediction for Uncertainty-Aware Motion Planning in Autonomous Driving

AAAI 2024technical

Autonomous driving systems rely on precise trajectory prediction for safe and efficient motion planning. Despite considerable efforts to enhance prediction accuracy, inherent uncertainties persist due to data noise and incomplete observations. Many strategies entail formalizing prediction outcomes i…

Cited by 3SourcePDFScholar
2024

SGDCL: Semantic-Guided Dynamic Correlation Learning for Explainable Autonomous Driving

IJCAI 2024poster

By learning expressive representations, deep learning (DL) has revolutionized autonomous driving (AD). Despite significant advancements, the inherent opacity of DL models engenders public distrust, impeding their widespread adoption. For explainable autonomous driving, current studies primarily conc…

2023

Location-Aware Visual Question Generation with Lightweight Models

EMNLP 2023long main

This work introduces a novel task, location-aware visual question generation (LocaVQG), which aims to generate engaging questions from data relevant to a particular geographical location. Specifically, we represent such location-aware information with surrounding images and a GPS coordinate. To tack…

Cited by 0SourcecodeScholar
2022

Selective Mutual Learning: An Efficient Approach for Single Channel Speech Separation

ICASSP 2022accepted

Mutual learning, the related idea to knowledge distillation, is a group of untrained lightweight networks, which simultaneously learn and share knowledge to perform tasks together during training. In this paper, we propose a novel mutual learning approach, namely selective mutual learning. This is t…

Cited by 0SourceScholar
2018

Image Representation Using Supervised and Unsupervised Learning Methods on Complex Domain

ICASSP 2018accepted

Matrix factorization (MF) and its extensions have been intensively studied in computer vision and machine learning. In this paper, unsupervised and supervised learning methods based on MF technique on complex domain are introduced. Projective complex matrix factorization (PCMF) and discriminant proj…

Cited by 0SourceScholar