← Search

Weiming Li

9 accepted papers

2026

DAM-VLA: A Dynamic Action Model-Based Vision-Language-Action Framework for Robot Manipulation

ICRA 2026poster

In dynamic environments such as warehouses, hospitals, and homes, robots must seamlessly transition between gross motion and precise manipulations to complete complex tasks. However, current Vision-Language-Action (VLA) frameworks, largely adapted from pre-trained Vision-Language Models (VLMs), ofte…

2026

DomainCQA: Crafting Knowledge-Intensive QA from Domain-Specific Charts

AAAI 2026technical

Chart Question Answering (CQA) evaluates Multimodal Large Language Models (MLLMs) on visual understanding and reasoning over chart data. However, existing benchmarks mostly test surface-level parsing, such as reading labels and legends, while overlooking deeper scientific reasoning. We propose Domai

Cited by 0SourcePDFScholar
2025

OAMaskFlow: Occlusion-Aware Motion Mask for Scene Flow

AAAI 2025technical

The scene flow estimation methods make significant progress by estimating pixel-wise 3D motion on implicitly learning a motion embedding using an end-to-end differentiable optimization framework. However, the motion embedding learned implicitly is insufficient for grouping pixels into rigid object i…

Cited by 0SourcePDFScholar
2024

DOCTR: Disentangled Object-Centric Transformer for Point Scene Understanding

AAAI 2024technical

Point scene understanding is a challenging task to process real-world scene point cloud, which aims at segmenting each object, estimating its pose, and reconstructing its mesh simultaneously. Recent state-of-the-art method first segments each object and then processes them independently with multipl…

2024

DVI-SLAM: A Dual Visual Inertial SLAM Network

ICRA 2024poster

Recent deep learning based visual simultaneous localization and mapping (SLAM) methods have made significant progress. However, how to make full use of visual information as well as better integrate with inertial measurement unit (IMU) in visual SLAM has potential research value. This paper proposes…

Cited by 14SourceScholar
2024

Is Your HD Map Constructor Reliable under Sensor Corruptions?

NeurIPS 2024poster

Driving systems often rely on high-definition (HD) maps for precise environmental information, which is crucial for planning and navigation. While current HD map constructors perform well under ideal conditions, their resilience to real-world challenges, \eg, adverse weather and sensor failures, is…

Cited by 17SourcePDFScholar
2022

Attention-guided RGB-D Fusion Network for Category-level 6D Object Pose Estimation

IROS 2022poster

This work focuses on estimating 6D poses and sizes of category-level objects from a single RGB-D image. How to exploit the complementary RGB and depth features plays an important role in this task yet remains an open question. Due to the large intra-category texture and shape variations, an object i…

Cited by 5SourceScholar
2021

Learning Generalized Intersection Over Union for Dense Pixelwise Prediction

ICML 2021spotlight

Intersection over union (IoU) score, also named Jaccard Index, is one of the most fundamental evaluation methods in machine learning. The original IoU computation cannot provide non-zero gradients and thus cannot be directly optimized by nowadays deep learning methods. Several recent works generaliz…

Cited by 33SourcePDFScholar
2021

UASNet: Uncertainty Adaptive Sampling Network for Deep Stereo Matching

ICCV 2021poster

Recent studies have shown that cascade cost volume can play a vital role in deep stereo matching to achieve high resolution depth map with efficient hardware usage. However, how to construct good cascade volume as well as effective sampling for them are still under in-depth study. Previous cascade-b…

Cited by 31PDFScholar