← Search

Kun Wu

25 accepted papers

2026

ArtVIP: Articulated Digital Assets of Visual Realism, Modular Interaction, and Physical Fidelity for Robot Learning

ICLR 2026poster

Robot learning increasingly relies on simulation to advance complex ability such as dexterous manipulations and precise interactions, necessitating high-quality digital assets to bridge the sim-to-real gap. However, existing open-source articulated object datasets for simulation are limited by insuf…

Cited by 0SourceScholar
2026

Diffusion Trajectory-Guided Policy for Long-Horizon Robot Manipulation

ICRA 2026poster

Recently, Vision-Language-Action Models (VLA) have advanced robot imitation learning, but high data collection costs and limited demonstrations hinder generalization and current imitation learning methods struggle in out-of-distribution scenarios, especially for long-horizon tasks. A key challenge i…

2026

Histopathology-Genomics Multi-modal Structural Representation Learning for Data-Efficient Precision Oncology

ICLR 2026poster

Fusing histopathology images and genomics data with deep learning has significantly advanced precision oncology. However, genomics data is often missing due to its high acquisition cost and complexity in real-world clinical scenarios. Existing solutions aim to reconstruct genomics data from histopat…

Cited by 0SourcecodeScholar
2026

LaST$_{0}$: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model

ICML 2026spotlight

Vision-Language-Action (VLA) models have recently shown strong generalization, with some approaches seeking to explicitly generate linguistic reasoning traces or predict future observations prior to execution. However, explicit reasoning typically incurs non-negligible inference latency, which const…

Cited by 0SourceScholar
2026

MLA: A Multisensory Language–Action Model for Multimodal Understanding and Forecasting in Robotic Manipulation

ICRA 2026poster

Vision-language-action models (VLAs) have shown generalization capabilities in robotic manipulation tasks by inheriting from vision-language models (VLMs) and learning action generation. Most VLA models focus on interpreting vision and language to generate actions, whereas robots must perceive and i…

2026

SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities for Large Language Models

ICML 2026poster

Evaluating large language models (LLMs) for software engineering has been limited by narrow task coverage, language bias, and insufficient alignment with real-world developer workflows. Existing benchmarks often focus on algorithmic problems or Python-centric bug fixing, leaving critical dimensions …

Cited by 0SourceScholar
2026

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations

ICML 2026oral

Recent progress in large-scale robotic datasets and vision-language models (VLMs) has advanced research on vision-language-action (VLA) models. However, existing VLA models still face two fundamental challenges: (\textit{i}) producing precise low-level actions from high-dimensional observations, (\t…

Cited by 0SourcecodeScholar
2025

Diffusion Trajectory-Guided Policy for Long-Horizon Robot Manipulation

RA-L 2025

Recently, Vision-Language-Action models (VLA) have advanced robot imitation learning, but high data collection costs and limited demonstrations hinder generalization and current imitation learning methods struggle in out-of-distribution scenarios, especially for long-horizon tasks. A key challenge i

Cited by 14SourcecodeScholar
2025

Discrete Policy: Learning Disentangled Action Space for Multi-Task Robotic Manipulation

ICRA 2025

Learning visuomotor policy for multi-task robotic manipulation has been a long-standing challenge for the robotics community. The difficulty lies in the diversity of action space: typically, a goal can be accomplished in multiple ways, resulting in a multimodal action distribution for a single task.

Cited by 24SourcecodeScholar
2025

FreqPolicy: Efficient Flow-based Visuomotor Policy via Frequency Consistency

NeurIPS 2025poster

Generative modeling-based visuomotor policies have been widely adopted in robotic manipulation, attributed to their ability to model multimodal action distributions. However, the high inference cost of multi-step sampling limits its applicability in real-time robotic systems. Existing approaches acc…

Cited by 0SourceScholar
2025

HACTS: a Human-As-Copilot Teleoperation System for Robot Learning

IROS 2025

Teleoperation is essential for autonomous robot learning, especially in manipulation tasks that require human demonstrations or corrections. However, most existing systems only offer unilateral robot control and lack the ability to synchronize the robot’s status with the teleoperation hardware, prev

Cited by 8SourceScholar
2025

Learning From Imperfect Demonstrations With Self-Supervision for Robotic Manipulation

ICRA 2025

Improving data utilization, especially for imperfect data from task failures, is crucial for robotic manipulation due to the challenging, time-consuming, and expensive data collection process in the real world. Current imitation learning (IL) typically discards imperfect data, focusing solely on suc

Cited by 7SourceScholar
2025

Mamba Policy: Towards Efficient 3D Diffusion Policy with Hybrid Selective State Models

IROS 2025

Diffusion models have been widely employed in the field of 3D manipulation due to their efficient capability to learn distributions, allowing for precise prediction of action trajectories. However, diffusion models typically rely on large parameter UNet backbones as policy networks, which can be cha

Cited by 22SourcecodeScholar
2025

Region-based Cluster Discrimination for Visual Representation Learning

ICCV 2025poster

Learning visual representations is foundational for a broad spectrum of downstream tasks. Although recent vision-language contrastive models, such as CLIP and SigLIP, have achieved impressive zero-shot performance via large-scale vision-language alignment, their reliance on global representations co…

2025

RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation

RSS 2025poster

Developing robust and general-purpose manipulation policies is a key goal in robotics. To achieve effective generalization, it is essential to construct comprehensive datasets that encompass a large number of demonstration trajectories and diverse tasks. Unlike vision or language data, which can be…

Cited by 20PDFScholar
2025

ThinkAnswer Loss: Balancing Semantic Similarity and Exact Matching for LLM Reasoning Enhancement

EMNLP 2025

Knowledge distillation for large language models often uses Chain-of-Thought (CoT) and answer pairs, but existing methods struggle with appropriate supervision signals. Uniform constraints (e.g., cross-entropy) on CoT can enforce literal, verbose reasoning and suppress expressive diversity, while so

Cited by 0SourcePDFScholar
2025

TinyVLA: Toward Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation

RA-L 2025

Vision-Language-Action (VLA) models have shown remarkable potential in visuomotor control and instruction comprehension through end-to-end learning processes. However, current VLA models face significant challenges: they are slow during inference and require extensive pre-training on large amounts o

Cited by 303SourceScholar
2025

Training-free Generation of Temporally Consistent Rewards from VLMs

ICCV 2025poster

Recent advances in vision-language models (VLMs) have significantly improved performance in embodied tasks such as goal decomposition and visual comprehension. However, providing accurate rewards for robotic manipulation without fine-tuning VLMs remains challenging due to the absence of domain-speci…

2025

Unifying Within and Across: Intra-Modality Multi-View Fusion and Inter-Modality Alignment for Knowledge Graph Completion

ICASSP 2025accepted

Multi-modal knowledge graph completion (MMKGC) enhances the structural and semantic richness of knowledge graphs by integrating diverse information across modalities. However, existing methods often either overlook the diversity within a single modality or fail to ensure effective cross-modality ali…

Cited by 0SourceScholar
2024

PASUM: A Pre-training Architecture for Social Media User Modeling Based on Text Graph

COLING 2024main

Modeling social media users is the core of social governance in the digital society. Existing works have incorporated different digital traces to better learn the representations of social media users, including text information encoded by pre-trained language models and social network information e…

2024

SoMeLVLM: A Large Vision Language Model for Social Media Processing

ACL 2024findings

The growth of social media, characterized by its multimodal nature, has led to the emergence of diverse phenomena and challenges, which calls for an effective approach to uniformly solve automated tasks. The powerful Large Vision Language Models make it possible to handle a variety of tasks simultan…

Cited by 7SourcePDFScholar
2022

CADRE: A Cascade Deep Reinforcement Learning Framework for Vision-Based Autonomous Urban Driving

AAAI 2022technical

Vision-based autonomous urban driving in dense traffic is quite challenging due to the complicated urban environment and the dynamics of the driving behaviors. Widely-applied methods either heavily rely on hand-crafted rules or learn from limited human experience, which makes them hard to generalize…

2021

Data Augmentation with Hierarchical SQL-to-Question Generation for Cross-domain Text-to-SQL Parsing

EMNLP 2021main

Data augmentation has attracted a lot of research attention in the deep learning era for its ability in alleviating data sparseness. The lack of labeled data for unseen evaluation databases is exactly the major challenge for cross-domain text-to-SQL parsing. Previous works either require human inter…

2021

Hierarchical Graph Attention Network for Few-Shot Visual-Semantic Learning

ICCV 2021poster

Deep learning has made tremendous success in computer vision, natural language processing and even visual-semantic learning, which requires a huge amount of labeled training data. Nevertheless, the goal of human-level intelligence is to enable a model to quickly obtain an in-depth understanding give…

Cited by 13PDFScholar
2020

Knowledge Transfer in Multi-Task Deep Reinforcement Learning for Continuous Control

NeurIPS 2020poster

While Deep Reinforcement Learning (DRL) has emerged as a promising approach to many complex tasks, it remains challenging to train a single DRL agent that is capable of undertaking multiple different continuous control tasks. In this paper, we present a Knowledge Transfer based Multi-task Deep Reinf…

Cited by 53SourcePDFScholar