← Search

Yizhe Li

9 accepted papers

2026

FastViDAR: Real-Time Omnidirectional Depth Estimation Via Alternative Hierarchical Attention

ICRA 2026poster

In this paper, we propose FastViDAR, a novel framework that takes four fisheye camera inputs and produces a full 360 depth map along with per-camera depth, fusion depth, and confidence estimates. Our main contributions are: (1) We introduce Alternative Hierarchical Attention (AHA) mechanism that eff…

2025

Differentially Private Fine-Tuning of Diffusion Models

ICCV 2025poster

Generative AI models, particularly diffusion models (DMs), have demonstrated exceptional capabilities in high-quality image synthesis. However, their large memorization capacity raises significant privacy concerns, especially when trained on sensitive datasets. This paper introduces DP-LoRA, a surpr…

2025

Enhancing Mamba Decoder with Bidirectional Interaction in Multi-Task Dense Prediction

ICCV 2025poster

Sufficient cross-task interaction is crucial for success in multi-task dense prediction. However, sufficient interaction often results in high computational complexity, forcing existing methods to face the trade-off between interaction completeness and computational efficiency. To address this limit…

2025

KD-RIEKF: Kinodynamic Right-Invariant EKF for Legged Robot State Estimation

IROS 2025

We present KD-RIEKF, a novel state estimation framework that incorporates kinodynamic constraints into the Right-Invariant Extended Kalman Filter (RIEKF). Our framework integrates generalized momentum-based contact estimation, centroidal dynamics, and a noise-adaptive module, improving state estimat

Cited by 0SourceScholar
2025

Keypoint-Aware RAG for Robotic Manipulation: In-Context Constraint Learning via Large-Scale Retrieval

IROS 2025

Recent advances in robotic manipulation leverage foundation models pre-trained on internet-scale data, where keypoint-based representations have shown promising results in spatial reasoning. However, existing approaches primarily focus on zero-shot generalization or human-collected demonstrations, w

Cited by 0SourcecodeScholar
2025

SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

NeurIPS 2025poster

Large language models (LLMs) have demonstrated remarkable proficiency in mainstream academic disciplines such as mathematics, physics, and computer science. However, human knowledge encompasses over 200 specialized disciplines, far exceeding the scope of existing benchmarks. The capabilities of LLMs…

Cited by 215SourceScholar
2025

VLR-Driver: Large Vision-Language-Reasoning Models for Embodied Autonomous Driving

ICCV 2025poster

The rise of embodied intelligence and multi-modal large language models has led to exciting advancements in the field of autonomous driving, establishing it as a prominent research focus in both academia and industry. However, when confronted with intricate and ambiguous traffic scenarios, the lack…

Cited by 0SourcePDFScholar
2023

Exploring the Benefits of Visual Prompting in Differential Privacy

ICCV 2023poster

Visual Prompting (VP) is an emerging and powerful technique that allows sample-efficient adaptation to downstream tasks by engineering a well-trained frozen source model. In this work, we explore the benefits of VP in constructing compelling neural network classifiers with differential privacy (DP).…

Cited by 19PDFcodeScholar