← Search

Junjie Wen

25 accepted papers

2026

ActiveUMI: Robotic Manipulation with Active Perception from Robot‑Free Human Demonstrations

ICRA 2026poster

We present ActiveUMI, a framework for a data collection system that transfers in-the-wild human demonstrations to robots capable of complex bimanual manipulation. ActiveUMI couples a portable VR teleoperation kit with sensorized controllers that mirror the robot's end-effectors, bridging human-robot…

2026

HumanoidExo: Scalable Whole-Body Humanoid Manipulation Via Wearable Exoskeleton

ICRA 2026poster

A significant bottleneck in humanoid policy learning is the acquisition of large-scale, diverse datasets, as collecting reliable real-world data remains both difficult and cost-prohibitive. To address this limitation, we introduce HumanoidExo, a novel system that transfers human motion to whole-body…

2026

Multi-View Stereo with Geometric Encoding for Large-Scale Dense Scene Reconstruction (I)

ICRA 2026poster

Multi-view stereo (MVS) implicitly encodes photometric and geometric cues into the cost volume for multi-view correspondence matching, transferring insufficient geometric cues essential to depth estimation and reconstruction. This paper proposes GE-MVS, a novel multi-view stereo network with geometr…

Cited by 0Scholar
2026

Open-World Object Manipulation with Vision-Language-Action Models Via Synthetic Multi-Modal Data

ICRA 2026poster

Imitation learning has proven to be highly effective in teaching robots dexterous manipulation skills. However, it typically relies on large amounts of robot data, which limits its scalability and applicability in dynamic, real-world environments. One key challenge in this context is object generali…

Cited by 0Scholar
2026

Scaling Real-World Robot Policy Evaluation via Discrete Diffusion World Model

ICML 2026spotlight

Evaluating generalist robot manipulation policies is costly and difficult to scale in the real world. While emerging world models (e.g., WorldEval, Ctrl-World) offer a promising alternative, the reliability of such evaluation remains a critical bottleneck. Specifically, their visual predictions can …

Cited by 0SourceScholar
2025

ChatVLA-2: Vision-Language-Action Model with Open-World Reasoning

NeurIPS 2025poster

Vision-language-action (VLA) models have emerged as the next generation of models in robotics. However, despite leveraging powerful pre-trained Vision-Language Models (VLMs), existing end-to-end VLA systems often lose key capabilities during fine-tuning as the model adapts to specific robotic tasks.…

Cited by 0SourceScholar
2025

ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

EMNLP 2025

Humans possess a unified cognitive ability to perceive, comprehend, and interact with the physical world. Why can’t large language models replicate this holistic understanding? Through a systematic analysis of existing training paradigms in vision-language-action models (VLA), we identify two key ch

2025

CoA-VLA: Improving Vision-Language-Action Models via Visual-Text Chain-of-Affordance

ICCV 2025poster

Robot foundation models, particularly Vision-Language-Action (VLA) models, have garnered significant attention for their ability to enhance robot policy learning, greatly improving robot's generalization and robustness. OpenAI's recent model, O1, showcased impressive capabilities in solving complex…

Cited by 0SourcePDFScholar
2025

DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control

CoRL 2025poster

Enabling robots to perform diverse tasks across varied environments is a central challenge in robot learning. While vision-language-action (VLA) models have shown promise for generalizable robot skills, realizing their full potential requires addressing limitations in action representation and effic…

Cited by 0SourceScholar
2025

DiffusionVLA: Scaling Robot Foundation Models via Unified Diffusion and Autoregression

ICML 2025poster

In this paper, we present DiffusionVLA, a novel framework that integrates autoregressive reasoning with diffusion policies to address the limitations of existing methods: while autoregressive Vision-Language-Action (VLA) models lack precise and robust action generation, diffusion-based policies inhe…

Cited by 0SourcePDFScholar
2025

Discrete Policy: Learning Disentangled Action Space for Multi-Task Robotic Manipulation

ICRA 2025

Learning visuomotor policy for multi-task robotic manipulation has been a long-standing challenge for the robotics community. The difficulty lies in the diversity of action space: typically, a goal can be accomplished in multiple ways, resulting in a multimodal action distribution for a single task.

Cited by 24SourcecodeScholar
2025

End-to-End Underwater Multi-View Stereo for Dense Scene Reconstruction

ICRA 2025

Recent advancements in learning-based multi-view stereo (MVS) have demonstrated significant improvements over traditional counterpart, primarily due to the extensive availability of multi-view training images with ground-truth metric depths in the terrestrial in-air domain. However, underwater multi

Cited by 3SourcecodeScholar
2025

Lightweight Yet High-Performance Defect Detector for Uav-Based Large-Scale Infrastructure Real-Time Inspection

ICRA 2025

Defect diagnosis in urban infrastructure is crucial for public safety. Traditional manual inspections face significant challenges in terms of accuracy and cost-effectiveness. In this paper, we propose a lightweight and hardware-friendly large-scale infrastructure detector, CUPID, highly suitable for

Cited by 2SourceScholar
2025

Multi-View Stereo with Geometric Encoding for Dense Scene Reconstruction

ICRA 2025

Multi-view stereo (MVS) implicitly encodes photometric and geometric cues into the cost volume for multi-view correspondence matching, transferring insufficient geometric cues essential to depth estimation and reconstruction. This paper proposes GE-MVS, a novel multi-view stereo network with geometr

Cited by 0SourcecodeScholar
2025

Scaling Diffusion Policy in Transformer to 1 Billion Parameters for Robotic Manipulation

ICRA 2025

Diffusion Policy is a powerful technique tool for learning end-to-end visuomotor robot control. It is expected that Diffusion Policy possesses scalability, a key attribute for deep neural networks, typically suggesting that increasing model size would lead to enhanced performance. However, our obser

Cited by 45SourcecodeScholar
2025

TinyVLA: Toward Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation

RA-L 2025

Vision-Language-Action (VLA) models have shown remarkable potential in visuomotor control and instruction comprehension through end-to-end learning processes. However, current VLA models face significant challenges: they are slow during inference and require extensive pre-training on large amounts o

Cited by 303SourceScholar
2024

Det-Recon-Reg: An Intelligent Framework Towards Automated Large-Scale Infrastructure Inspection

IROS 2024poster

Visual inspection plays a predominant role in inspecting infrastructure surface. However, the generalization of existing visual inspection systems to large-scale real-world scenes remains challenging. In this paper, we introduce Det-Recon-Reg, an intelligent framework separating the complex inspecti…

Cited by 1SourcecodeScholar
2024

EnYOLO: A Real-Time Framework for Domain-Adaptive Underwater Object Detection with Image Enhancement

ICRA 2024poster

In recent years, significant progress has been made in the field of underwater image enhancement (UIE). However, its practical utility for high-level vision tasks, such as underwater object detection (UOD) in Autonomous Underwater Vehicles (AUVs), remains relatively unexplored. It may be attributed…

Cited by 2SourceScholar
2024

Language-Conditioned Robotic Manipulation with Fast and Slow Thinking

ICRA 2024poster

The language-conditioned robotic manipulation aims to transfer natural language instructions into executable actions, from simple "pick-and-place" to tasks requiring intent recognition and visual reasoning. Inspired by the dual-process theory in cognitive science—which suggests two parallel systems…

Cited by 17SourceScholar
2024

Object-Centric Instruction Augmentation for Robotic Manipulation

ICRA 2024poster

Humans interpret scenes by recognizing both the identities and positions of objects in their observations. For a robot to perform tasks such as "pick and place", understanding both what the objects are and where they are located is crucial. While the former has been extensively discussed in the lite…

Cited by 14SourceScholar
2023

A Topic-Enhanced Approach for Emotion Distribution Forecasting in Conversations

ICASSP 2023accepted

Emotion Forecasting in Conversations (EFC), the task aims to predict the emotion of next utterance (yet to come), has received more and more attention in recent years. However, this task ignores the one-to-many feature of dialogue and its prediction target is emotion label, which is flawed in most c…

Cited by 0SourceScholar
2023

Category-Level 6D Pose Estimation Using Geometry-Guided Instance-Aware Prior and Multi-Stage Reconstruction

RA-L 2023

Category-level object 6D pose estimation is essential for robotic manipulation, augmented reality and 3D scene understanding. It aims to accurately predict the translation and rotation of arbitrary shape instances from a given set of object classes without models of each instance. However, such esti

Cited by 5SourceScholar
2023

SyreaNet: A Physically Guided Underwater Image Enhancement Framework Integrating Synthetic and Real Images

ICRA 2023poster

Underwater image enhancement (UIE) is vital for high-level vision-related underwater tasks. Although learning-based UIE methods have made remarkable achievements in recent years, it's still challenging for them to consistently deal with various underwater conditions, which could be caused by: 1) the…

Cited by 37SourcecodeScholar
2022

Learning to Detect Noisy Labels Using Model-Based Features

EMNLP 2022finding

Label noise is ubiquitous in various machine learning scenarios such as self-labeling with model predictions and erroneous data annotation. Many existing approaches are based on heuristics such as sample losses, which might not be flexible enough to achieve optimal solutions. Meta learning based met…