← Search

Ömer Erdinç Yağmurlu

4 accepted papers

2025

BEAST: Efficient Tokenization of B-Splines Encoded Action Sequences for Imitation Learning

NeurIPS 2025poster

We present the B-spline Encoded Action Sequence Tokenizer (BEAST), a novel action tokenizer that encodes action sequences into compact discrete or continuous tokens using B-splines. In contrast to existing action tokenizers based on vector quantization or byte pair encoding, BEAST requires no separ…

Cited by 0SourceScholar
2025

FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Flow Models

CoRL 2025poster

Developing efficient Vision-Language-Action (VLA) policies is crucial for practical robotics deployment, yet current approaches face prohibitive computational costs and resource requirements. Existing diffusion-based VLA policies require multi-billion-parameter models and massive datasets to achieve…

Cited by 0SourceScholar
2024

Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

RSS 2024poster

This work introduces the Multimodal Diffusion Transformer (MDT), a novel diffusion policy framework, that excels at learning versatile behavior from multimodal goal specifications with few language annotations. MDT leverages a diffusion based multimodal transformer backbone and two self-supervised a…

2024

Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models

CoRL 2024poster

A central challenge towards developing robots that can relate human language to their perception and actions is the scarcity of natural language annotations in diverse robot datasets. Moreover, robot policies that follow natural language instructions are typically trained on either templated languag…

Cited by 7SourceScholar