← Search

Dongming Wu

11 accepted papers

2026

AutoFly: Vision-Language-Action Model for UAV Autonomous Navigation in the Wild

ICLR 2026poster

Vision-language navigation (VLN) requires intelligent agents to navigate environments by interpreting linguistic instructions alongside visual observations, serving as a cornerstone task in Embodied AI. Current VLN research for unmanned aerial vehicles (UAVs) relies on detailed, pre-specified instru…

Cited by 0SourceScholar
2026

SpaceVista: All-Scale Visual Spatial Reasoning from mm to km

ICML 2026poster

With the current surge in spatial reasoning, researchers have made significant progress in understanding indoor scenes, but still struggle with more diverse applications. This paper aims to advance all-scale spatial reasoning by tackling two key challenges: 1) the heavy reliance on indoor 3D scans a…

Cited by 0SourcecodeScholar
2025

DrivingSphere: Building a High-fidelity 4D World for Closed-loop Simulation

CVPR 2025poster

Autonomous driving evaluation requires simulation environments that closely replicate actual road conditions, including real-world sensory data and responsive feedback loops. However, many existing simulations need to predict waypoints along fixed routes on public datasets or synthetic photorealisti…

2025

Language Prompt for Autonomous Driving

AAAI 2025technical

A new trend in the computer vision community is to capture objects of interest following flexible human command represented by a natural language prompt. However, the progress of using language prompts in driving scenarios is stuck in a bottleneck due to the scarcity of paired prompt-instance data.…

2025

RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping

ICCV 2025poster

General robotic grasping systems require accurate object affordance perception in diverse open-world scenarios following human instructions. However, current studies suffer from the problem of lacking reasoning-based large-scale affordance prediction data, leading to considerable concern about open-…

Cited by 0SourcePDFScholar
2024

Merlin: Empowering Multimodal LLMs with Foresight Minds

ECCV 2024poster

"Humans can foresee the future based on present observations, a skill we term as foresight minds. However, this capability remains under-explored within existing MLLMs, hindering their capacity to understand intentions behind subjects. To address this, we integrate the future modeling into MLLMs. By…

2024

TopoMLP: A Simple yet Strong Pipeline for Driving Topology Reasoning

ICLR 2024poster

Topology reasoning aims to comprehensively understand road scenes and present drivable routes in autonomous driving. It requires detecting road centerlines (lane) and traffic elements, further reasoning their topology relationship, \textit{i.e.}, lane-lane topology, and lane-traffic topology. In thi…

2023

OnlineRefer: A Simple Online Baseline for Referring Video Object Segmentation

ICCV 2023poster

Referring video object segmentation (RVOS) aims at segmenting an object in a video following human instruction. Current state-of-the-art methods fall into an offline pattern, in which each clip independently interacts with text embedding for cross-modal understanding. They usually present that the o…

Cited by 55PDFcodeScholar
2023

Referring Multi-Object Tracking

CVPR 2023poster

Existing referring understanding tasks tend to involve the detection of a single text-referred object. In this paper, we propose a new and general referring understanding task, termed referring multi-object tracking (RMOT). Its core idea is to employ a language expression as a semantic cue to guide…

2022

Airborne Mimo Radar Transmit-Receive Design Under Spectral Constraint in Signal-Dependent Clutter

ICASSP 2022accepted

This paper considers the joint design of the transmit waveform and receive filter for airborne multiple-input multiple-output (MIMO) radar under spectral constraint in signal-dependent clutter. The spatial-frequency spectral compatibility constraint is imposed in the joint design problem. To tackle…

Cited by 0SourceScholar
2022

Multi-Level Representation Learning With Semantic Alignment for Referring Video Object Segmentation

CVPR 2022poster

Referring video object segmentation (RVOS) is a challenging language-guided video grounding task, which requires comprehensively understanding the semantic information of both video content and language queries for object prediction. However, existing methods adopt multi-modal fusion at a frame-base…

Cited by 64PDFScholar