← Search

Jan-Nico Zaech

15 accepted papers

2026

AR-VLA: Autoregressive Action Expert for Vision–Language–Action Models

RSS 2026poster

We propose a standalone autoregressive (AR) Action Expert that generates actions as a continuous causal sequence while conditioning on refreshable vision-language prefixes. In contrast to existing Vision-Language-Action (VLA) models and diffusion policies that reset temporal context with each new ob…

Cited by 0SourceScholar
2026

Autonomous Vehicle Path Planning by Searching with Differentiable Simulation

AAAI 2026technical

Planning allows an agent to safely refine its actions before executing them in the real world. In autonomous driving, this is crucial to avoid collisions and navigate in complex, dense traffic scenarios. One way to plan is to search for the best action sequence. However, this is challenging when all

Cited by 0SourcePDFScholar
2026

ConceptPose: Training-Free Zero-Shot Object Pose Estimation using Concept Vectors

CVPR 2026

Object pose estimation is a fundamental task in computer vision and robotics, yet most methods require extensive, dataset-specific training. Concurrently, large-scale vision language models show remarkable zero-shot capabilities. In this work, we bridge these two worlds by introducing ConceptPose, a

Cited by 0SourcecodeScholar
2026

GaussianVLM: Scene-Centric 3D Vision-Language Models Using Language-Aligned Gaussian Splats for Embodied Reasoning and Beyond

ICRA 2026poster

As multimodal language models advance, their application to 3D scene understanding is a fast-growing frontier, driving the development of 3D Vision-Language Models (VLMs). Current methods show strong dependence on object detectors, introducing processing bottlenecks and limitations in taxonomic flex…

2026

SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding

CVPR 2026

Robotic Foundation Models (RFMs) hold great promise as generalist, end-to-end systems for robot control.Yet their ability to generalize across new environments, tasks, and embodiments remains limited.We argue that a major bottleneck lies in their foundations: most RFMs are built by fine-tuning inter

Cited by 0SourceScholar
2026

Unlocking Efficient Vehicle Dynamics Modeling via Analytic World Models

AAAI 2026technical

Differentiable simulators represent an environment’s dynamics as a differentiable function. Within robotics and autonomous driving, this property is used in Analytic Policy Gradients (APG), which relies on backpropagating through the dynamics to train accurate policies for diverse tasks. Here we sho

Cited by 0SourcePDFScholar
2025

Articulate3D: Holistic Understanding of 3D Scenes as Universal Scene Description

ICCV 2025poster

3D scene understanding is a long-standing challenge in computer vision and a key component in enabling mixed reality, wearable computing, and embodied AI. Providing a solution to these applications requires a multifaceted approach that covers scene-centric, object-centric, as well as interaction-cen…

Cited by 0SourcePDFScholar
2025

Generalist Robot Manipulation beyond Action Labeled Data

CoRL 2025poster

Recent advances in generalist robot manipulation leverage pre-trained Vision–Language Models (VLMs) and large-scale robot demonstrations to tackle diverse tasks in a zero-shot manner. A key challenge remains: scaling high-quality, action-labeled robot demonstration data, which existing methods rely…

Cited by 0SourceScholar
2025

LangHOPS: Language Grounded Hierarchical Open-Vocabulary Part Segmentation

NeurIPS 2025poster

We propose LangHOPS, the first Multimodal Large Language Model (MLLM)-based framework for open-vocabulary object–part instance segmentation. Given an image, LangHOPS can jointly detect and segment hierarchical object and part instances from open-vocabulary candidate categories. Unlike prior approach…

Cited by 0SourceScholar
2025

ReVLA: Reverting Visual Domain Limitation of Robotic Foundation Models

ICRA 2025

Recent progress in large language models and access to large-scale robotic datasets has sparked a paradigm shift in robotics models transforming them into generalists able to adapt to various tasks, scenes, and robot modalities. A large step for the community are open Vision Language Action models w

Cited by 22SourceScholar
2024

Probabilistic Sampling of Balanced K-Means using Adiabatic Quantum Computing

CVPR 2024poster

Adiabatic quantum computing (AQC) is a promising approach for discrete and often NP-hard optimization problems. Current AQCs allow to implement problems of research interest which has sparked the development of quantum representations for many computer vision tasks. Despite requiring multiple measur…

Cited by 1SourcePDFScholar
2022

Adiabatic Quantum Computing for Multi Object Tracking

CVPR 2022poster

Multi-Object Tracking (MOT) is most often approached in the tracking-by-detection paradigm, where object detections are associated through time. The association step naturally leads to discrete optimization problems. As these optimization problems are often NP-hard, they can only be solved exactly f…

Cited by 34PDFScholar
2022

Learnable Online Graph Representations for 3D Multi-Object Tracking

RA-L 2022

Autonomous systems that operate in dynamic environments require robust object tracking in 3D as one of their key components. Most recent approaches for 3D multi-object tracking (MOT) from LIDAR use object dynamics together with a set of handcrafted features to match detections of objects across mult

Cited by 78SourceScholar
2021

Decoder Fusion RNN: Context and Interaction Aware Decoders for Trajectory Prediction

IROS 2021poster

Forecasting the future behavior of all traffic agents in the vicinity is a key task to achieve safe and reliable autonomous driving systems. It is a challenging problem as agents adjust their behavior depending on their intentions, the others’ actions, and the road layout. In this paper, we propose…

Cited by 17SourceScholar
2020

Action Sequence Predictions of Vehicles in Urban Environments using Map and Social Context

IROS 2020poster

This work studies the problem of predicting the sequence of future actions for surrounding vehicles in real-world driving scenarios. To this aim, we make three main contributions. The first contribution is an automatic method to convert the trajectories recorded in real-world driving scenarios to ac…

Cited by 13SourceScholar