← Search

Meng Wei

16 accepted papers

2026

Ground Slow, Move Fast: A Dual-System Foundation Model for Generalizable Vision-Language Navigation

ICLR 2026poster

While recent large vision-language models (VLMs) have improved generalization in vision-language navigation (VLN), existing methods typically rely on end-to-end pipelines that map vision-language inputs directly to short-horizon discrete actions. Such designs often produce fragmented motions, incur…

Cited by 0SourceScholar
2026

NavDP: Learning Sim-To-Real Navigation Diffusion Policy with Privileged Information Guidance

ICRA 2026poster

Learning to navigate in dynamic and complex open-world environments is a critical yet challenging capability for autonomous robots. Existing approaches often rely on cascaded modular frameworks, which require extensive hyperparameter tuning or learning from limited real-world demonstration data. In …

2026

Open-Vocabulary Object-Goal Navigation by Generalizing Semantic Mapping with Dense CLIP

ICRA 2026poster

Object-oriented embodied navigation tasks require agents to locate specific objects, either defined by category or images, in unseen environments. While recent methods have made progress in extending closed-set models to open-vocabulary scenarios with foundation models, they typically rely on traini…

Cited by 0Scholar
2026

SeedVR2: One-Step Video Restoration via Diffusion Adversarial Post-Training

ICLR 2026poster

Recent advances in diffusion-based video restoration (VR) demonstrate significant improvement in visual quality, yet yield a prohibitive computational cost during inference. While several distillation-based approaches have exhibited the potential of one-step image restoration, extending existing app…

Cited by 0SourcecodeScholar
2026

StreamVLN: Streaming Vision-And-Language Navigation Via SlowFast Context Modeling

ICRA 2026poster

Vision-and-Language Navigation (VLN) in real-world settings requires agents to process continuous visual streams and generate actions with low latency grounded in language instructions. While Video-based Large Language Models (Video-LLMs) have driven recent progress, current VLN methods based on Vid…

2026

Unified Camera Positional Encoding for Controlled Video Generation

CVPR 2026

Transformers have emerged as a universal backbone across 3D perception, video generation, and world models for autonomous driving and embodied AI, where understanding camera geometry is essential for grounding visual observations in three-dimensional space. However, existing camera encoding methods

Cited by 0SourcecodeScholar
2026

Where It Moves, It Matters: Referring Surgical Instrument Segmentation via Motion

AAAI 2026technical

Enabling intuitive, language-driven interaction with surgical scenes is a critical step toward intelligent operating rooms and autonomous surgical robotic assistance. However, the task of referring segmentation, localizing surgical instruments based on natural language descriptions, remains underexp

Cited by 0SourcePDFScholar
2025

CameraCtrl II: Dynamic Scene Exploration via Camera-controlled Video Diffusion Models

ICCV 2025poster

This paper introduces CameraCtrl II, a framework that enables continuous and dynamic scene exploration through a camera-controlled video diffusion model. Previous camera-conditioned video generative models suffer from diminished video dynamics and limited range of viewpoints when generating videos w…

Cited by 0SourcePDFScholar
2025

Learning from True-False Labels via Multi-modal Prompt Retrieving

ICML 2025poster

Pre-trained **V**ision-**L**anguage **M**odels (VLMs) exhibit strong zero-shot classification abilities, demonstrating great potential for generating weakly supervised labels. Unfortunately, existing weakly supervised learning methods are short of ability in generating accurate labels via VLMs. In t…

2025

SeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video Restoration

CVPR 2025highlight

Video restoration poses non-trivial challenges in maintaining fidelity while recovering temporally consistent details from unknown degradations in the wild. Despite recent advances in diffusion-based restoration, these methods often face limitations in generation capability and sampling efficiency.…

2024

Normal-GS: 3D Gaussian Splatting with Normal-Involved Rendering

NeurIPS 2024poster

Rendering and reconstruction are long-standing topics in computer vision and graphics. Achieving both high rendering quality and accurate geometry is a challenge. Recent advancements in 3D Gaussian Splatting (3DGS) have enabled high-fidelity novel view synthesis at real-time speeds. However, the noi…

Cited by 2SourcePDFScholar
2023

OV-PARTS: Towards Open-Vocabulary Part Segmentation

NeurIPS 2023poster

Segmenting and recognizing diverse object parts is a crucial ability in applications spanning various computer vision and robotic tasks. While significant progress has been made in object-level Open-Vocabulary Semantic Segmentation (OVSS), i.e., segmenting objects with arbitrary text, the correspond…

2023

Text Promptable Surgical Instrument Segmentation with Vision-Language Models

NeurIPS 2023poster

In this paper, we propose a novel text promptable surgical instrument segmentation approach to overcome challenges associated with diversity and differentiation of surgical instruments in minimally invasive surgeries. We redefine the task as text promptable, thereby enabling a more nuanced comprehen…

2022

Rethinking the Two-Stage Framework for Grounded Situation Recognition

AAAI 2022technical

Grounded Situation Recognition (GSR), i.e., recognizing the salient activity (or verb) category in an image (e.g.,buying) and detecting all corresponding semantic roles (e.g.,agent and goods), is an essential step towards “human-like” event understanding. Since each verb is associated with a specifi…

2021

Vision Transformer With Progressive Sampling

ICCV 2021poster

Transformers with powerful global relation modeling abilities have been introduced to fundamental computer vision tasks recently. As a typical example, the Vision Transformer (ViT) directly applies a pure transformer architecture on image classification, by simply splitting images into tokens with a…

Cited by 126PDFcodeScholar