← Search

Di Feng

11 accepted papers

2026

SO-Bench: A Structural Output Evaluation of Multimodal LLM

CVPR 2026

Multimodal large language models (MLLMs) are increasingly deployed in real-world, agentic settings where outputs must not only be correct, but also conform to pre-defined data schemas. Despite recent progress in structured generation in textual domain, there is still no benchmark that systematically

Cited by 0SourcecodeScholar
2025

Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms

ICLR 2025poster

Building a generalist model for user interface (UI) understanding is challenging due to various foundational issues, such as platform diversity, resolution variation, and data limitation. In this paper, we introduce Ferret-UI 2, a multimodal large language model (MLLM) designed for universal UI unde…

Cited by 0SourcePDFScholar
2025

UINavBench: A Framework for Comprehensive Evaluation of Interactive Digital Agents

ICCV 2025poster

We build a comprehensive online evaluation benchmark for language-conditioned multi-step task execution on mobile interfaces. Our benchmark strives to evaluate the multi-step planning, reasoning, and visual grounding capabilities of agents, using mobile user interfaces as a concrete testbed. To buil…

Cited by 0SourcePDFScholar
2024

FastOcc: Accelerating 3D Occupancy Prediction by Fusing the 2D Bird’s-Eye View and Perspective View

ICRA 2024poster

In autonomous driving, 3D occupancy prediction outputs voxel-wise status and semantic labels for more comprehensive understandings of 3D scenes compared with traditional perception tasks, such as 3D object detection and bird’s-eye view (BEV) semantic segmentation. Recent researchers have extensively…

Cited by 35SourceScholar
2024

HP3: Hierarchical Prediction-Pretrained Planning for Unprotected Left Turn

IROS 2024poster

Trajectory planning for unprotected left turns poses a significant challenge in autonomous driving. Reinforcement learning (RL) offers potential, but existing methods often rely on scenario-specific state representations, limiting their adaptability. This paper introduces Hierarchical Prediction-Pre…

Cited by 0SourceScholar
2024

LCA-on-the-Line: Benchmarking Out of Distribution Generalization with Class Taxonomies

ICML 2024oral

We tackle the challenge of predicting models' Out-of-Distribution (OOD) performance using in-distribution (ID) measurements without requiring OOD data. Existing evaluations with ``Effective robustness'', which use ID accuracy as an indicator of OOD accuracy, encounter limitations when models are tra…

2024

Mitigating Causal Confusion in Vector-Based Behavior Cloning for Safer Autonomous Planning

ICRA 2024poster

The utilization of vector-based deep learning techniques has great prospects in the realm of autonomous driving, particularly in the domains of prediction and planning tasks. However, the application of vector-based backbones for prediction and planning tasks may lead to the occurrence of causal con…

Cited by 1SourceScholar
2022

DeepFusion: A Robust and Modular 3D Object Detector for Lidars, Cameras and Radars

IROS 2022poster

We propose DeepFusion, a modular multi-modal architecture to fuse lidars, cameras and radars in different combinations for 3D object detection. Specialized feature extractors take advantage of each modality and can be exchanged easily, making the approach simple and flexible. Extracted features are…

Cited by 28SourceScholar
2021

A Simple and Efficient Multi-task Network for 3D Object Detection and Road Understanding

IROS 2021poster

Detecting dynamic objects and predicting static road information such as drivable areas and ground heights are crucial for safe autonomous driving. Previous works studied each perception task separately, and lacked a collective quantitative analysis. In this work, we show that it is possible to perf…

Cited by 28SourcecodeScholar
2020

Inferring Spatial Uncertainty in Object Detection

IROS 2020poster

The availability of real-world datasets is the prerequisite for developing object detection methods for autonomous driving. While ambiguity exists in object labels due to error-prone annotation process or sensor observation noises, current object detection datasets only provide deterministic annotat…

Cited by 34SourceScholar
2017

A Tactile-Based Framework for Active Object Learning and Discrimination using Multimodal Robotic Skin

RA-L 2017

In this letter, we propose a complete probabilistic tactile-based framework to enable robots to autonomously explore unknown workspaces and recognize objects based on their physical properties. Our framework consists of three components: 1) an active pretouch strategy to efficiently explore unknown

Cited by 64SourceScholar