← Search

Tsung-Wei Ke

13 accepted papers

2026

RobotArena $\infty$: Unlimited Robot Benchmarking via Real-to-Sim Translation

ICLR 2026poster

The pursuit of robot generalists—instructable agents capable of performing diverse tasks across diverse environments—demands rigorous and scalable evaluation. Yet real-world testing of robot policies remains fundamentally constrained: it is labor-intensive, slow, unsafe at scale, and difficult to re…

Cited by 0SourcecodeScholar
2025

See, Point, Fly: A Learning-Free VLM Framework for Universal Unmanned Aerial Navigation

CoRL 2025poster

We present See, Point, Fly (SPF), a training-free aerial vision-and-language navigation (AVLN) framework built atop vision-language models (VLMs). SPF is capable of navigating to any goal based on any type of free-form instructions in any kind of environment. In contrast to existing VLM-based approa…

Cited by 0SourceScholar
2024

3D Diffuser Actor: Policy Diffusion with 3D Scene Representations

CoRL 2024poster

Diffusion policies are conditional diffusion models that learn robot action distributions conditioned on the robot and environment state. They have recently shown to outperform both deterministic and alternative action distribution learning formulations. 3D robot policies use 3D scene feature repres…

Cited by 117SourceScholar
2024

Diffusion-ES: Gradient-free Planning with Diffusion for Autonomous and Instruction-guided Driving

CVPR 2024poster

Diffusion models excel at modeling complex and multimodal trajectory distributions for decision-making and control. Reward-gradient guided denoising has been recently proposed to generate trajectories that maximize both a differentiable reward function and the likelihood under the data distribution…

Cited by 6SourcePDFScholar
2024

Learning Hierarchical Image Segmentation For Recognition and By Recognition

ICLR 2024spotlight

Large vision and language models learned directly through image-text associations often lack detailed visual substantiation, whereas image segmentation tasks are treated separately from recognition, supervisedly learned without interconnections. Our key observation is that, while an image can be r…

2023

ChainedDiffuser: Unifying Trajectory Diffusion and Keypose Prediction for Robotic Manipulation

CoRL 2023poster

We present ChainedDiffuser, a policy architecture that unifies action keypose prediction and trajectory diffusion generation for learning robot manipulation from demonstrations. Our main innovation is to use a global transformer-based action predictor to predict actions at keyframes, a task that req…

Cited by 88SourceScholar
2023

Diffusion-TTA: Test-time Adaptation of Discriminative Models via Generative Feedback

NeurIPS 2023poster

The advancements in generative modeling, particularly the advent of diffusion models, have sparked a fundamental question: how can these models be effectively used for discriminative tasks? In this work, we find that generative models can be great test-time adapters for discriminative models. Our me…

2022

Unsupervised Hierarchical Semantic Segmentation With Multiview Cosegmentation and Clustering Transformers

CVPR 2022oral

Unsupervised semantic segmentation aims to discover groupings within and across images that capture object- and view-invariance of a category without external supervision. Grouping naturally has levels of granularity, creating ambiguity in unsupervised segmentation. Existing methods avoid this ambig…

Cited by 59PDFcodeScholar
2021

Universal Weakly Supervised Segmentation by Pixel-to-Segment Contrastive Learning

ICLR 2021poster

Weakly supervised segmentation requires assigning a label to every pixel based on training instances with partial annotations such as image-level tags, object bounding boxes, labeled points and scribbles. This task is challenging, as coarse annotations (tags, boxes) lack precise pixel localization w…

2019

Adversarial Structure Matching for Structured Prediction Tasks

CVPR 2019poster

Pixel-wise losses, i.e., cross-entropy or L2, have been widely used in structured prediction tasks as a spatial extension of generic image classification or regression. However, its i.i.d. assumption neglects the structural regularity present in natural images. Various attempts have been made to inc…

Cited by 18PDFcodeScholar