← Search

Philipp Krähenbühl

21 accepted papers

2026

Latent Chain-of-Thought World Modeling for End-to-End Autonomous Driving

CVPR 2026

Recent Vision-Language-Action (VLA) models for autonomous driving explore inference-time reasoning as a way to improve driving performance and safety in challenging scenarios. Most prior work uses natural language to express chain-of-thought (CoT) reasoning before producing driving actions. However,

Cited by 0SourceScholar
2025

Long-term Traffic Simulation with Interleaved Autoregressive Motion and Scenario Generation

ICCV 2025poster

An ideal traffic simulator replicates the realistic long-term point-to-point trip that a self-driving system experiences during deployment. Prior models and benchmarks focus on closed-loop motion simulation for initial agents in a scene. This is problematic for long-term simulation. Agents enter and…

2023

Learning Video Representations From Large Language Models

CVPR 2023highlight

We introduce LAVILA, a new approach to learning video-language representations by leveraging Large Language Models (LLMs). We repurpose pre-trained LLMs to be conditioned on visual input, and finetune them to create automatic video narrators. Our auto-generated narrations offer a number of advantage…

2023

PartDistillation: Learning Parts From Instance Segmentation

CVPR 2023poster

We present a scalable framework to learn part segmentation from object instance labels. State-of-the-art instance segmentation models contain a surprising amount of part information. However, much of this information is hidden from plain view. For each object instance, the part information is noisy,…

2022

Detecting Twenty-Thousand Classes Using Image-Level Supervision

ECCV 2022poster

"Current object detectors are limited in vocabulary size due to the small scale of detection datasets. Image classifiers, on the other hand, reason about much larger vocabularies, as their datasets are larger and easier to collect. We propose Detic, which simply trains the classifiers of a detector…

2019

Monocular Plan View Networks for Autonomous Driving

IROS 2019poster

Convolutions on monocular dash cam videos capture spatial invariances in the image plane but do not explicitly reason about distances and depth. We propose a simple transformation of observations into a bird's eye view, also known as plan view, for end-to-end control. We detect vehicles and pedestri…

Cited by 95SourceScholar
2018

Compressed Video Action Recognition

CVPR 2018poster

Training robust deep video representations has proven to be much more challenging than learning deep image representations. This is in part due to the enormous size of raw video streams and the high temporal redundancy; the true and interesting signal is often drowned in too much irrelevant data. Mo…

Cited by 428SourcePDFScholar