← Search

Bernt Schiele

161 accepted papers

2026

Align Once to Explain: Feature Alignment for Scalable B-cosification of Foundational Vision Transformers

CVPR 2026

Foundational vision models have become the de facto standard for many vision tasks due to their strong performance. However, they are notoriously opaque and remain hard to interpret. We present ALOE (ALign Once to Explain), a one-time, label-free feature alignment based approach that efficiently con

Cited by 0SourcecodeScholar
2026

AnyUp: Universal Feature Upsampling

ICLR 2026oral

We introduce AnyUp, a method for feature upsampling that can be applied to any vision feature at any resolution, without encoder-specific training. Existing learning-based upsamplers for features like DINO or CLIP need to be re-trained for every feature extractor and thus do not generalize to differ…

Cited by 0SourcecodeScholar
2026

Beyond Accuracy: What Matters in Designing Well-Behaved Image Classification Models?

ICML 2026poster

Deep learning has become an essential part of computer vision, with deep neural networks (DNNs) excelling in predictive performance. However, they often fall short in other critical quality dimensions, such as robustness, calibration, or fairness. While existing studies have focused on a subset of t…

Cited by 0SourceScholar
2026

Certified Circuits: Stability Guarantees for Mechanistic Circuits

ICML 2026poster

Understanding *how* neural networks arrive at their predictions is essential for debugging, auditing, and deployment. Mechanistic interpretability pursues this goal by identifying *circuits*—minimal subnetworks responsible for specific behaviors. However, existing circuit discovery methods are britt…

Cited by 0SourceScholar
2026

DAVE: Distribution-aware Attribution via ViT Gradient Decomposition

ICML 2026spotlight

Vision Transformers (ViTs) have become a dominant architecture in computer vision, yet producing stable and high-resolution attribution maps for these models remains challenging. Architectural components such as patch embeddings and attention routing often introduce structured artifacts in pixel-lev…

Cited by 0SourceScholar
2026

Interpretable 3D Neural Object Volumes for Robust Conceptual Reasoning

ICLR 2026poster

With the rise of deep neural networks, especially in safety-critical applications, robustness and interpretability are crucial to ensure their trustworthiness. Recent advances in 3D-aware classifiers that map image features to volumetric representation of objects, rather than relying solely on 2D ap…

Cited by 0SourcecodeScholar
2026

Rewis3d: Reconstruction Improves Weakly-Supervised Semantic Segmentation

CVPR 2026

We present Rewis3d, a framework that leverages recent advances in feed-forward 3D reconstruction to significantly improve weakly supervised semantic segmentation on 2D images. Obtaining dense, pixel-level annotations remains a costly bottleneck for training segmentation models. Alleviating this issu

Cited by 0SourcecodeScholar
2026

Temporal Concept Dynamics in Diffusion Models via Prompt-Conditioned Interventions

ICLR 2026poster

Diffusion models are usually evaluated by their final outputs, gradually denoising random noise into meaningful images. Yet, generation unfolds along a trajectory, and understanding this dynamic process is crucial for explaining how controllable, reliable, and predictable these models are in terms…

Cited by 0SourcecodeScholar
2026

What is Missing? Explaining Neurons Activated by Absent Concepts

ICML 2026poster

Explainable artificial intelligence (XAI) aims to provide human-interpretable insights into the behavior of deep neural networks (DNNs), typically by estimating a simplified causal structure of the model. In existing work, this causal structure often includes relationships where the presence of a co…

Cited by 0SourceScholar
2025

AIM: Amending Inherent Interpretability via Self-Supervised Masking

ICCV 2025poster

It has been observed that deep neural networks (DNNs) often use both genuine as well as spurious features.In this work, we propose "Amending Inherent Interpretability via Self-Supervised Masking" (AIM), a simple yet surprisingly effective method that promotes the network's utilization of genuine fea…

Cited by 0SourcePDFScholar
2025

FaCT: Faithful Concept Traces for Explaining Neural Network Decisions

NeurIPS 2025poster

Deep networks have shown remarkable performance across a wide range of tasks, yet getting a global concept-level understanding of how they function remains a key challenge. Many post-hoc concept-based approaches have been introduced to understand their workings, yet they are not always faithful to t…

Cited by 0SourceScholar
2025

How to Probe: Simple Yet Effective Techniques for Improving Post-hoc Explanations

ICLR 2025poster

Post-hoc importance attribution methods are a popular tool for “explaining” Deep Neural Networks (DNNs) and are inherently based on the assumption that the explanations can be applied independently of how the models were trained. Contrarily, in this work we bring forward empirical evidence that chal…

2025

KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models

NeurIPS 2025poster

Recent advances in multi-modal generative models have enabled significant progress in instruction-based image editing. However, while these models produce visually plausible outputs, their capacity for knowledge-based reasoning editing tasks remains under-explored. In this paper, We introduce KRIS-B…

Cited by 0SourceScholar
2025

MET3R: Measuring Multi-View Consistency in Generated Images

CVPR 2025poster

We introduce MEt3R, a metric for multi-view consistency in generated images. Large-scale generative models for multi-view image generation are rapidly advancing the field of 3D inference from sparse observations. However, due to the nature of generative modeling, traditional reconstruction metrics a…

Cited by 1SourcePDFScholar
2025

Number it: Temporal Grounding Videos like Flipping Manga

CVPR 2025poster

Video Large Language Models (Vid-LLMs) have made remarkable advancements in comprehending video content for QA dialogue. However, they struggle to extend this visual understanding to tasks requiring precise temporal localization, known as Video Temporal Grounding (VTG). To address this, we introduce…

2025

PersonaHOI: Effortlessly Improving Face Personalization in Human-Object Interaction Generation

CVPR 2025poster

We introduce PersonaHOI, a training- and tuning-free framework that fuses a general StableDiffusion model with a personalized face diffusion (PFD) model to generate identity-consistent human-object interaction (HOI) images. While existing PFD models have advanced significantly, they often overemphas…

2025

Pixel-level Certified Explanations via Randomized Smoothing

ICML 2025poster

Post-hoc attribution methods aim to explain deep learning predictions by highlighting influential input pixels. However, these explanations are highly non-robust: small, imperceptible input perturbations can drastically alter the attribution map while maintaining the same prediction. This vulnerabil…

2025

Samba: Synchronized Set-of-Sequences Modeling for Multiple Object Tracking

ICLR 2025spotlight

Multiple object tracking in complex scenarios - such as coordinated dance performances, team sports, or dynamic animal groups - presents unique challenges. In these settings, objects frequently move in coordinated patterns, occlude each other, and exhibit long-term dependencies in their trajectories…

Cited by 2SourcePDFScholar
2025

Solving Inverse Problems with FLAIR

NeurIPS 2025poster

Flow-based latent generative models such as Stable Diffusion 3 are able to generate images with remarkable quality, even enabling photorealistic text-to-image generation. Their impressive performance suggests that these models should also constitute powerful priors for inverse imaging problems, but…

Cited by 0SourcecodeScholar
2025

Test-Time Visual In-Context Tuning

CVPR 2025poster

Visual in-context learning (VICL), as a new paradigm in computer vision, allows the model to rapidly adapt to various tasks with only a handful of prompts and examples. While effective, the existing VICL paradigm exhibits poor generalizability under distribution shifts. In this work, we propose test…

2025

TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters

ICLR 2025spotlight

Transformers have become the predominant architecture in foundation models due to their excellent performance across various domains. However, the substantial cost of scaling these models remains a significant concern. This problem arises primarily from their dependence on a fixed number of paramete…

2025

Unbiasing through Textual Descriptions: Mitigating Representation Bias in Video Benchmarks

CVPR 2025poster

We propose a new "Unbiased through Textual Description (UTD)" video benchmark based on unbiased subsets of existing video classification and retrieval datasets to enable a more robust assessment of video understanding capabilities. Namely, we tackle the problem that current video benchmarks may suff…

Cited by 0SourcePDFScholar
2025

VITAL: More Understandable Feature Visualization through Distribution Alignment and Relevant Information Flow

ICCV 2025poster

Neural networks are widely adopted to solve complex and challenging tasks. Especially in high-stakes decision-making, understanding their reasoning process is crucial, yet proves challenging for modern deep networks. Feature visualization (FV) is a powerful tool to decode what information neurons ar…

2024

Adaptive Hierarchical Certification for Segmentation using Randomized Smoothing

ICML 2024poster

Certification for machine learning is proving that no adversarial sample can evade a model within a range under certain conditions, a necessity for safety-critical domains. Common certification methods for segmentation use a flat set of fine-grained classes, leading to high abstain rates due to mode…

2024

B-cosification: Transforming Deep Neural Networks to be Inherently Interpretable

NeurIPS 2024poster

B-cos Networks have been shown to be effective for obtaining highly human interpretable explanations of model decisions by architecturally enforcing stronger alignment between inputs and weight. B-cos variants of convolutional networks (CNNs) and vision transformers (ViTs), which primarily replace l…

2024

Discover-then-Name: Task-Agnostic Concept Bottlenecks via Automated Concept Discovery

ECCV 2024poster

"Concept Bottleneck Models (CBMs) have recently been proposed to address the ‘black-box’ problem of deep neural networks, by first mapping images to a human-understandable concept space and then linearly combining concepts for classification. Such models typically require first coming up with a set…

2024

GiT: Towards Generalist Vision Transformer through Universal Language Interface

ECCV 2024oral

"This paper proposes a simple, yet effective framework, called , simultaneously applicable for various vision tasks only with a vanilla ViT. Motivated by the universality of the Multi-layer Transformer architecture (e.g., GPT) widely used in large language models (LLMs), we seek to broaden its scope…

2024

Good Teachers Explain: Explanation-Enhanced Knowledge Distillation

ECCV 2024poster

"Knowledge Distillation (KD) has proven effective for compressing large teacher models into smaller student models. While it is well known that student models can achieve similar accuracies as the teachers, it has also been shown that they nonetheless often do not learn the same function. It is, how…

2024

HowToCaption: Prompting LLMs to Transform Video Annotations at Scale

ECCV 2024poster

"Instructional videos are a common source for learning text-video or even multimodal representations by leveraging subtitles extracted with automatic speech recognition systems (ASR) from the audio signal in the videos. However, in contrast to human-annotated captions, both speech and subtitles natu…

2024

MTA-CLIP: Language-Guided Semantic Segmentation with Mask-Text Alignment

ECCV 2024poster

"Recent approaches have shown that large-scale vision-language models such as CLIP can improve semantic segmentation performance. These methods typically aim for pixel-level vision-language alignment, but often rely on low-resolution image features from CLIP, resulting in class ambiguities along bou…

Cited by 5SourcePDFScholar
2024

On Adversarial Training without Perturbing all Examples

ICLR 2024poster

Adversarial training is the de-facto standard for improving robustness against adversarial examples. This usually involves a multi-step adversarial attack applied on each example during training. In this paper, we explore only constructing adversarial examples (AE) on a subset of the training exampl…

2024

OrCo: Towards Better Generalization via Orthogonality and Contrast for Few-Shot Class-Incremental Learning

CVPR 2024highlight

Few-Shot Class-Incremental Learning (FSCIL) introduces a paradigm in which the problem space expands with limited data. FSCIL methods inherently face the challenge of catastrophic forgetting as data arrives incrementally making models susceptible to overwriting previously acquired knowledge. Moreove…

2024

Scribbles for All: Benchmarking Scribble Supervised Segmentation Across Datasets

NeurIPS 2024spotlight

In this work, we introduce *Scribbles for All*, a label and training data generation algorithm for semantic segmentation trained on scribble labels. Training or fine-tuning semantic segmentation models with weak supervision has become an important topic recently and was subject to significant advanc…

2024

Training Vision Transformers for Semi-Supervised Semantic Segmentation

CVPR 2024poster

We present S4Former a novel approach to training Vision Transformers for Semi-Supervised Semantic Segmentation (S4). At its core S4Former employs a Vision Transformer within a classic teacher-student framework and then leverages three novel technical ingredients: PatchShuffle as a parameter-free per…

2024

Walker: Self-supervised Multiple Object Tracking by Walking on Temporal Object Appearance Graphs

ECCV 2024poster

"The supervision of state-of-the-art multiple object tracking (MOT) methods requires enormous annotation efforts to provide bounding boxes for all frames of all videos, and instance IDs to associate them through time. To this end, we introduce Walker, the first self-supervised tracker that learns fr…

2024

X-MIC: Cross-Modal Instance Conditioning for Egocentric Action Generalization

CVPR 2024poster

Lately there has been growing interest in adapting vision-language models (VLMs) to image and third-person video classification due to their success in zero-shot recognition. However the adaptation of these models to egocentric videos has been largely unexplored. To address this gap we propose a sim…

2024

latentSplat: Autoencoding Variational Gaussians for Fast Generalizable 3D Reconstruction

ECCV 2024poster

"We present latentSplat, a method to predict semantic Gaussians in a 3D latent space that can be splatted and decoded by a light-weight generative 2D architecture. Existing methods for generalizable 3D reconstruction either do not scale to large scenes and resolutions, or are limited to interpolatio…

Cited by 63SourcePDFScholar
2023

A Meta-Learning Approach to Predicting Performance and Data Requirements

CVPR 2023poster

We propose an approach to estimate the number of samples required for a model to reach a target performance. We find that the power law, the de facto principle to estimate model performance, leads to large error when using a small dataset (e.g., 5 samples per class) for extrapolation. This is becaus…

2023

Class-Incremental Exemplar Compression for Class-Incremental Learning

CVPR 2023poster

Exemplar-based class-incremental learning (CIL) finetunes the model with all samples of new classes but few-shot exemplars of old classes in each incremental phase, where the "few-shot" abides by the limited memory budget. In this paper, we break this "few-shot" limit based on a simple yet surprisin…

2023

Continual Detection Transformer for Incremental Object Detection

CVPR 2023poster

Incremental object detection (IOD) aims to train an object detector in phases, each with annotations for new object categories. As other incremental settings, IOD is subject to catastrophic forgetting, which is often addressed by techniques such as knowledge distillation (KD) and exemplar replay (ER…

Cited by 87SourcePDFScholar
2023

DSVT: Dynamic Sparse Voxel Transformer With Rotated Sets

CVPR 2023poster

Designing an efficient yet deployment-friendly 3D backbone to handle sparse point clouds is a fundamental problem in 3D perception. Compared with the customized sparse convolution, the attention mechanism in Transformers is more appropriate for flexibly modeling long-range relationships and is easie…

2023

FreeMatch: Self-adaptive Thresholding for Semi-supervised Learning

ICLR 2023poster

Semi-supervised Learning (SSL) has witnessed great success owing to the impressive performances brought by various methods based on pseudo labeling and consistency regularization. However, we argue that existing methods might fail to utilize the unlabeled data more effectively since they either use…

2023

HGFormer: Hierarchical Grouping Transformer for Domain Generalized Semantic Segmentation

CVPR 2023poster

Current semantic segmentation models have achieved great success under the independent and identically distributed (i.i.d.) condition. However, in real-world applications, test data might come from a different domain than training data. Therefore, it is important to improve model robustness against…

2023

Improving Robustness of Vision Transformers by Reducing Sensitivity To Patch Corruptions

CVPR 2023poster

Despite their success, vision transformers still remain vulnerable to image corruptions, such as noise or blur. Indeed, we find that the vulnerability mainly stems from the unstable self-attention mechanism, which is inherently built upon patch-based inputs and often becomes overly sensitive to the…

2023

In-Style: Bridging Text and Uncurated Videos with Style Transfer for Text-Video Retrieval

ICCV 2023poster

Large-scale noisy web image-text datasets have been proven to be efficient for learning robust vision-language models. However, to transfer them to the task of video retrieval, models still need to be fine-tuned on hand-curated paired text-video data to adapt to the diverse styles of video descripti…

Cited by 4PDFcodeScholar
2023

Learning by Sorting: Self-supervised Learning with Group Ordering Constraints

ICCV 2023poster

Contrastive learning has become an important tool in learning representations from unlabeled data mainly relying on the idea of minimizing distance between positive data pairs, e.g., views from the same images, and maximizing distance between negative data pairs, e.g., views from different images. T…

Cited by 12PDFcodeScholar
2023

Online Hyperparameter Optimization for Class-Incremental Learning

AAAI 2023technical

Class-incremental learning (CIL) aims to train a classification model while the number of classes increases phase-by-phase. An inherent challenge of CIL is the stability-plasticity tradeoff, i.e., CIL models should keep stable to retain old knowledge and keep plastic to absorb new knowledge. However…

2023

SSB: Simple but Strong Baseline for Boosting Performance of Open-Set Semi-Supervised Learning

ICCV 2023poster

Semi-supervised learning (SSL) methods effectively leverage unlabeled data to improve model generalization. However, SSL models often underperform in open-set scenarios, where unlabeled data contain outliers from novel categories that do not appear in the labeled set. In this paper, we study the cha…

Cited by 15PDFcodeScholar
2023

Self-Supervised Pre-Training With Masked Shape Prediction for 3D Scene Understanding

CVPR 2023poster

Masked signal modeling has greatly advanced self-supervised pre-training for language and 2D images. However, it is still not fully explored in 3D scene understanding. Thus, this paper introduces Masked Shape Prediction (MSP), a new framework to conduct masked signal modeling in 3D scenes. MSP uses…

2023

SoftMatch: Addressing the Quantity-Quality Tradeoff in Semi-supervised Learning

ICLR 2023poster

The critical challenge of Semi-Supervised Learning (SSL) is how to effectively leverage the limited labeled data and massive unlabeled data to improve the model's generalization performance. In this paper, we first revisit the popular pseudo-labeling methods via a unified sample weighting formulatio…

2023

Studying How to Efficiently and Effectively Guide Models with Explanations

ICCV 2023poster

Despite being highly performant, deep neural networks might base their decisions on features that spuriously correlate with the provided labels, thus hurting generalization. To mitigate this, 'model guidance' has recently gained popularity, i.e. the idea of regularizing the models' explanations to e…

Cited by 14PDFcodeScholar
2023

Temperature Schedules for self-supervised contrastive methods on long-tail data

ICLR 2023poster

Most approaches for self-supervised learning (SSL) are optimised on curated balanced datasets, e.g. ImageNet, despite the fact that natural data usually exhibits long-tail distributions. In this paper, we analyse the behaviour of one of the most popular variants of SSL, i.e. contrastive methods, on…

2023

Towards Robust Object Detection Invariant to Real-World Domain Shifts

ICLR 2023poster

Safety-critical applications such as autonomous driving require robust object detection invariant to real-world domain shifts. Such shifts can be regarded as different domain styles, which can vary substantially due to environment changes and sensor noises, but deep models only know the training dom…

Cited by 37SourcePDFScholar
2023

UniTR: A Unified and Efficient Multi-Modal Transformer for Bird's-Eye-View Representation

ICCV 2023poster

Jointly processing information from multiple sensors is crucial to achieving accurate and robust perception for reliable autonomous driving systems. However, current 3D perception research follows a modality-specific paradigm, leading to additional computation overheads and inefficient collaboration…

Cited by 78PDFcodeScholar
2023

Unsupervised Open-Vocabulary Object Localization in Videos

ICCV 2023poster

In this paper, we show that recent advances in video representation learning and pre-trained vision-language models allow for substantial improvements in self-supervised video object localization. We propose a method that first localizes objects in videos via a slot attention approach and then assig…

Cited by 7PDFcodeScholar
2023

Visual Coherence Loss for Coherent and Visually Grounded Story Generation

ACL 2023findings

Local coherence is essential for long-form text generation models. We identify two important aspects of local coherence within the visual storytelling task: (1) the model needs to represent re-occurrences of characters within the image sequence in order to mention them correctly in the story; (2) ch…

2023

Weakly-Supervised Domain Adaptive Semantic Segmentation With Prototypical Contrastive Learning

CVPR 2023poster

There has been a lot of effort in improving the performance of unsupervised domain adaptation for semantic segmentation task, however there is still a huge gap in performance when compared with supervised learning. In this work, we propose a common framework to use different weak labels, e.g. image,…

2022

A Unified Query-Based Paradigm for Point Cloud Understanding

CVPR 2022poster

3D point cloud understanding is an important component in autonomous driving and robotics. In this paper, we present a novel Embedding-Querying paradigm (EQ- Paradigm) for 3D understanding tasks including detection, segmentation and classification. EQ-Paradigm is a unified paradigm that enables comb…

Cited by 55PDFcodeScholar
2022

Assaying Out-Of-Distribution Generalization in Transfer Learning

NeurIPS 2022accept

Since out-of-distribution generalization is a generally ill-posed problem, various proxy targets (e.g., calibration, adversarial robustness, algorithmic corruptions, invariance across shifts) were studied across different research programs resulting in different recommendations. While sharing the sa…

2022

Bi-Level Alignment for Cross-Domain Crowd Counting

CVPR 2022poster

Recently, crowd density estimation has received increasing attention. The main challenge for this task is to achieve high-quality manual annotations on a large amount of training data. To avoid reliance on such annotations, previous works apply unsupervised domain adaptation (UDA) techniques by tran…

Cited by 40PDFcodeScholar
2022

Class-Agnostic Object Counting Robust to Intraclass Diversity

ECCV 2022poster

"Most previous works on object counting are limited to pre-defined categories. In this paper, we focus on classagnostic counting, i.e., counting object instances in an image by simply specifying a few exemplar boxes of interest. We start with an analysis on intraclass diversity and point out three f…

2022

CoSSL: Co-Learning of Representation and Classifier for Imbalanced Semi-Supervised Learning

CVPR 2022poster

Standard semi-supervised learning (SSL) using class-balanced datasets has shown great progress to leverage unlabeled data effectively. However, the more realistic setting of class-imbalanced data - called imbalanced SSL - is largely underexplored and standard SSL tends to underperform. In this paper…

Cited by 69PDFcodeScholar
2022

Keypoint Message Passing for Video-Based Person Re-identification

AAAI 2022technical

Video-based person re-identification~(re-ID) is an important technique in visual surveillance systems which aims to match video snippets of people captured by different cameras. Existing methods are mostly based on convolutional neural networks~(CNNs), whose building blocks either process local neig…

2022

Motion Transformer with Global Intention Localization and Local Movement Refinement

NeurIPS 2022accept

Predicting multimodal future behavior of traffic participants is essential for robotic vehicles to make safe decisions. Existing works explore to directly predict future trajectories based on latent features or utilize dense goal candidates to identify agent's destinations, where the former strategy…

2022

Omni-DETR: Omni-Supervised Object Detection With Transformers

CVPR 2022poster

We consider the problem of omni-supervised object detection, which can use unlabeled, fully labeled and weakly labeled annotations, such as image tags, counts, points, etc., for object detection. This is enabled by a unified architecture, Omni-DETR, based on the recent progress on student-teacher fr…

Cited by 63PDFcodeScholar
2022

PoseTrack21: A Dataset for Person Search, Multi-Object Tracking and Multi-Person Pose Tracking

CVPR 2022poster

Current research evaluates person search, multi-object tracking and multi-person pose estimation as separate tasks and on different datasets although these tasks are very akin to each other and comprise similar sub-tasks, e.g. person detection or appearance-based association of detected persons. Con…

Cited by 60PDFcodeScholar
2022

RBGNet: Ray-Based Grouping for 3D Object Detection

CVPR 2022poster

As a fundamental problem in computer vision, 3D object detection is experiencing rapid growth. To extract the point-wise features from the irregularly and sparsely distributed points, previous methods usually take a feature grouping module to aggregate the point features to an object candidate. Howe…

Cited by 75PDFcodeScholar
2022

SHIFT: A Synthetic Driving Dataset for Continuous Multi-Task Domain Adaptation

CVPR 2022poster

Adapting to a continuously evolving environment is a safety-critical challenge inevitably faced by all autonomous-driving systems. Existing image- and video-based driving datasets, however, fall short of capturing the mutable nature of the real world. In this paper, we introduce the largest syntheti…

Cited by 166PDFScholar
2022

USB: A Unified Semi-supervised Learning Benchmark for Classification

NeurIPS 2022accept

Semi-supervised learning (SSL) improves model generalization by leveraging massive unlabeled data to augment limited labeled samples. However, currently, popular SSL evaluation protocols are often constrained to computer vision (CV) tasks. In addition, previous work typically trains deep neural netw…

2022

VGSE: Visually-Grounded Semantic Embeddings for Zero-Shot Learning

CVPR 2022poster

Human-annotated attributes serve as powerful semantic embeddings in zero-shot learning. However, their annotation process is labor-intensive and needs expert supervision. Current unsupervised semantic embeddings, i.e., word embeddings, enable knowledge transfer between classes. However, word embeddi…

Cited by 76PDFcodeScholar
2021

Convolutional Dynamic Alignment Networks for Interpretable Classifications

CVPR 2021poster

We introduce a new family of neural network models called Convolutional Dynamic Alignment Networks (CoDA-Nets), which are performant classifiers with a high degree of inherent interpretability. Their core building blocks are Dynamic Alignment Units (DAUs), which linearly transform their input with w…

Cited by 71PDFcodeScholar
2021

Euro-PVI: Pedestrian Vehicle Interactions in Dense Urban Centers

CVPR 2021poster

Accurate prediction of pedestrian and bicyclist paths is integral to the development of reliable autonomous vehicles in dense urban environments. The interactions between vehicle and pedestrian or bicyclist have a significant impact on the trajectories of traffic participants e.g. stopping or turnin…

Cited by 45PDFScholar
2021

Generalized and Incremental Few-Shot Learning by Explicit Learning and Calibration Without Forgetting

ICCV 2021poster

Both generalized and incremental few-shot learning have to deal with three major challenges: learning novel classes from only few samples per class, preventing catastrophic forgetting of base classes, and classifier calibration across novel and base classes. In this work we propose a three-stage fra…

Cited by 77PDFcodeScholar
2021

Learning Decision Trees Recurrently Through Communication

CVPR 2021poster

Integrated interpretability without sacrificing the prediction accuracy of decision making algorithms has the potential of greatly improving their value to the user. Instead of assigning a label to an image directly, we propose to learn iterative binary sub-decisions, inducing sparsity and transpare…

Cited by 20PDFcodeScholar
2021

Seeking Similarities Over Differences: Similarity-Based Domain Alignment for Adaptive Object Detection

ICCV 2021poster

In order to robustly deploy object detectors across a wide range of scenarios, they should be adaptable to shifts in the input distribution without the need to constantly annotate new data. This has motivated research in Unsupervised Domain Adaptation (UDA) algorithms for detection. UDA methods lear…

Cited by 111PDFcodeScholar
2021

You Only Need Adversarial Supervision for Semantic Image Synthesis

ICLR 2021poster

Despite their recent successes, GAN models for semantic image synthesis still suffer from poor image quality when trained with only adversarial supervision. Historically, additionally employing the VGG-based perceptual loss has helped to overcome this issue, significantly improving the synthesis qua…

2020

Attribute Prototype Network for Zero-Shot Learning

NeurIPS 2020poster

From the beginning of zero-shot learning research, visual attributes have been shown to play an important role. In order to better transfer attribute-based knowledge from known to unknown classes, we argue that an image representation with integrated attribute localization ability would be beneficia…

Cited by 378SourcePDFScholar
2020

Confidence-Calibrated Adversarial Training: Generalizing to Unseen Attacks

ICML 2020poster

Adversarial training yields robust models against a specific threat model, e.g., $L_\infty$ adversarial examples. Typically robustness does not generalize to previously unseen threat models, e.g., other $L_p$ norms, or larger perturbations. Our confidence-calibrated adversarial training (CCAT) tackl…

2020

Deep Wiener Deconvolution: Wiener Meets Deep Learning for Image Deblurring

NeurIPS 2020oral

We present a simple and effective approach for non-blind image deblurring, combining classical techniques and deep learning. In contrast to existing methods that deblur the image directly in the standard image space, we propose to perform an explicit deconvolution process in a feature space by integ…

Cited by 156SourcePDFScholar
2020

Mnemonics Training: Multi-Class Incremental Learning Without Forgetting

CVPR 2020oral

Multi-Class Incremental Learning (MCIL) aims to learn new concepts by incrementally updating a model trained on previous concepts. However, there is an inherent trade-off to effectively learning new concepts without catastrophic forgetting of previous ones. To alleviate this issue, it has been propo…

Cited by 454PDFcodeScholar
2020

Normalizing Flows With Multi-Scale Autoregressive Priors

CVPR 2020poster

Flow-based generative models are an important class of exact inference models that admit efficient inference and sampling for image synthesis. Owing to the efficiency constraints on the design of the flow layers, e.g. split coupling flow layers in which approximately half the pixels do not undergo f…

Cited by 14PDFcodeScholar
2020

Prediction Poisoning: Towards Defenses Against DNN Model Stealing Attacks

ICLR 2020poster

High-performance Deep Neural Networks (DNNs) are increasingly deployed in many real-world applications e.g., cloud prediction APIs. Recent advances in model functionality stealing attacks via black-box access (i.e., inputs in, predictions out) threaten the business model of such applications, which…

Cited by 228SourceScholar
2020

Segmentations-Leak: Membership Inference Attacks and Defenses in Semantic Image Segmentation

ECCV 2020poster

Today's success of state of the art methods for semantic segmentation is driven by large datasets. Data is considered an important asset that needs to be protected, as the collection and annotation of such datasets comes at significant efforts and associated costs. In addition, visual data might con…

2020

Towards Automated Testing and Robustification by Semantic Adversarial Data Generation

ECCV 2020poster

Widespread application of computer vision systems in real world tasks is currently hindered by their unexpected behavior on unseen examples. This occurs due to limitations of empirical testing on finite test sets and lack of systematic methods to identify the breaking points of a trained model. In t…

Cited by 5SourcePDFScholar
2019

Bayesian Prediction of Future Street Scenes using Synthetic Likelihoods

ICLR 2019poster

For autonomous agents to successfully operate in the real world, the ability to anticipate future scene states is a key competence. In real-world scenarios, future states become increasingly uncertain and multi-modal, particularly on long time horizons. Dropout based Bayesian inference provides a co…

Cited by 56SourcePDFScholar
2019

F-VAEGAN-D2: A Feature Generating Framework for Any-Shot Learning

CVPR 2019poster

When labeled training data is scarce, a promising data augmentation approach is to generate visual features of unknown classes using their attributes. To learn the class conditional distribution of CNN features, these models rely on pairs of image features and class attributes. Hence, they can not m…

Cited by 647PDFScholar
2019

Learning to Self-Train for Semi-Supervised Few-Shot Classification

NeurIPS 2019poster

Few-shot classification (FSC) is challenging due to the scarcity of labeled training data (e.g. only one labeled data point per class). Meta-learning has shown to achieve promising results by learning to initialize a classification model for FSC. In this paper we propose a novel semi-supervised meta…

2019

Not Using the Car to See the Sidewalk -- Quantifying and Controlling the Effects of Context in Classification and Segmentation

CVPR 2019poster

Importance of visual context in scene understanding tasks is well recognized in the computer vision community. However, to what extent the computer vision models are dependent on the context to make their predictions is unclear. A model overly relying on context will fail when encountering objects…

Cited by 101PDFScholar
2019

Semantic Projection Network for Zero- and Few-Label Semantic Segmentation

CVPR 2019poster

Semantic segmentation is one of the most fundamental problems in computer vision and pixel-level labelling in this context is particularly expensive. Hence, there have been several attempts to reduce the annotation effort such as learning from image level labels and bounding box annotations. In this…

Cited by 294PDFScholar
2018

A Hybrid Model for Identity Obfuscation by Face Replacement

ECCV 2018poster

As more and more personal photos are shared and tagged in social media, avoiding privacy risks such as unintended recognition, becomes increasingly challenging. We propose a new hybrid approach to obfuscate identities in photos by head replacement. Our approach combines state of the art parametric f…

Cited by 143SourcePDFScholar
2018

Accurate and Diverse Sampling of Sequences Based on a “Best of Many” Sample Objective

CVPR 2018poster

For autonomous agents to successfully operate in the real world, anticipation of future events and states of their environment is a key competence. This problem has been formalized as a sequence extrapolation problem, where a number of observations are used to predict the sequence into the future. R…

Cited by 140SourcePDFScholar
2018

Adversarial Scene Editing: Automatic Object Removal from Weak Supervision

NeurIPS 2018poster

While great progress has been made recently in automatic image manipulation, it has been limited to object centric images like faces or structured scene datasets. In this work, we take a step towards general scene-level image editing by developing an automatic interaction-free object removal model.…

Cited by 112SourcePDFScholar
2018

Connecting Pixels to Privacy and Utility: Automatic Redaction of Private Information in Images

CVPR 2018poster

Images convey a broad spectrum of personal information. If such images are shared on social media platforms, this personal information is leaked which conflicts with the privacy of depicted persons. Therefore, we aim for automated approaches to redact such private information and thereby protect pr…

2018

Disentangled Person Image Generation

CVPR 2018poster

Generating novel, yet realistic, images of persons is a challenging task due to the complex interplay between the different image factors, such as the foreground, background and pose information. In this work, we aim at generating such images based on a novel, two-stage reconstruction pipeline that…

Cited by 541SourcePDFScholar
2018

Diverse Conditional Image Generation by Stochastic Regression with Latent Drop-Out Codes

ECCV 2018poster

Recent advances in Deep Learning and probabilistic modeling have let to strong improvements in generative models for images. On the one hand, GANs have contributed a highly effective adversarial learning procedure, but still suffer from stability issues. On the other hand, CVAE models provide a soun…

2018

Long-Term On-Board Prediction of People in Traffic Scenes Under Uncertainty

CVPR 2018poster

Progress towards advanced systems for assisted and autonomous driving is leveraging recent advances in recognition and segmentation methods. Yet, we are still facing challenges in bringing reliable driving to inner cities, as those are composed of highly dynamic scenes observed from a moving platfo…

Cited by 288SourcePDFScholar
2018

Multimodal Explanations: Justifying Decisions and Pointing to the Evidence

CVPR 2018poster

Deep models that are both effective and explainable are desirable in many settings; prior explainable models have been unimodal, offering either image-based visualization of attention weights or text-based generation of post-hoc justifications. We propose a multimodal approach to explanation, and…

2018

Natural and Effective Obfuscation by Head Inpainting

CVPR 2018poster

As more and more personal photos are shared online, being able to obfuscate identities in such photos is becoming a necessity for privacy protection. People have largely resorted to blacking out or blurring head regions, but they result in poor user experience while being surprisingly ineffective ag…

Cited by 266SourcePDFScholar
2018

PoseTrack: A Benchmark for Human Pose Estimation and Tracking

CVPR 2018poster

Existing systems for video-based pose estimation and tracking struggle to perform well on realistic videos with multiple people and often fail to output body-pose trajectories consistent over time. To address this shortcoming this paper introduces PoseTrack which is a new large-scale benchmark for v…

Cited by 621SourcePDFScholar
2017

Adversarial Image Perturbation for Privacy Protection -- A Game Theory Perspective

ICCV 2017poster

Users like sharing personal photos with others through social media. At the same time, they might want to make automatic identification in such photos difficult or even impossible. Classic obfuscation methods such as blurring are not only unpleasant but also not as effective as one would expect. Rec…

Cited by 183PDFScholar
2017

ArtTrack: Articulated Multi-Person Tracking in the Wild

CVPR 2017oral

In this paper we propose an approach for articulated tracking of multiple people in unconstrained videos. Our starting point is a model that resembles existing architectures for single-frame pose estimation but is substantially faster. We achieve this in two ways: (1) by simplifying and sparsifying…

Cited by 379PDFScholar
2017

Exploiting Saliency for Object Segmentation From Image Level Labels

CVPR 2017poster

There have been remarkable improvements in the semantic labelling task in the recent years. However, the state of the art methods rely on large-scale pixel-level annotations. This paper studies the problem of training a pixel-wise semantic labeller network from image-level annotations of the present…

Cited by 236PDFScholar
2017

Generating Descriptions With Grounded and Co-Referenced People

CVPR 2017poster

Learning how to generate descriptions of images or videos received major interest both in the Computer Vision and Natural Language Processing communities. While a few works have proposed to learn a grounding during the generation process in an unsupervised way (via an attention mechanism), it remain…

Cited by 76PDFScholar
2017

Joint Graph Decomposition & Node Labeling: Problem, Algorithms, Applications

CVPR 2017poster

We state a combinatorial optimization problem whose feasible solutions define both a decomposition and a node labeling of a given graph. This problem offers a common mathematical abstraction of seemingly unrelated computer vision tasks, including instance-separating semantic segmentation, articulate…

Cited by 131PDFcodeScholar
2017

Learning Video Object Segmentation From Static Images

CVPR 2017spotlight

Inspired by recent advances of deep learning in instance segmentation and object tracking, we introduce the concept of convnet-based guidance applied to video object segmentation. Our model proceeds on a per-frame basis, guided by the output of the previous frame towards the object of interest in th…

Cited by 651PDFScholar
2017

Multiple People Tracking by Lifted Multicut and Person Re-Identification

CVPR 2017poster

Tracking multiple persons in a monocular video of a crowded scene is a challenging task. Humans can master it even if they loose track of a person locally by re-identifying the same person based on their appearance. Care must be taken across long distances, as similar-looking persons need not be ide…

Cited by 703PDFScholar
2017

Pose Guided Person Image Generation

NeurIPS 2017poster

This paper proposes the novel Pose Guided Person Generation Network (PG$^2$) that allows to synthesize person images in arbitrary poses, based on an image of that person and a novel pose. Our generation framework PG$^2$ utilizes the pose information explicitly and consists of two key stages: pose in…

Cited by 1080SourcePDFScholar
2017

Simple Does It: Weakly Supervised Instance and Semantic Segmentation

CVPR 2017poster

Semantic labelling and instance segmentation are two tasks that require particularly costly annotations. Starting from weak supervision in the form of bounding box detection annotations, we propose a new approach that does not require modification of the segmentation training procedure. We show that…

Cited by 971PDFScholar
2017

Speaking the Same Language: Matching Machine to Human Captions by Adversarial Training

ICCV 2017poster

While strong progress has been made in image captioning recently, machine and human captions are still quite distinct. This is primarily due to the deficiencies in the generated word distribution, vocabulary size, and strong bias in the generators towards frequent captions. Furthermore, humans -- ri…

Cited by 311PDFScholar
2017

Towards a Visual Privacy Advisor: Understanding and Predicting Privacy Risks in Images

ICCV 2017poster

With an increasing number of users sharing information online, privacy implications entailing such actions are a major concern. For explicit content, such as user profile or GPS data, devices (e.g. mobile phones) as well as web services (e.g. facebook) offer to set privacy settings in order to enfor…

Cited by 194PDFScholar
2016

DeepCut: Joint Subset Partition and Labeling for Multi Person Pose Estimation

CVPR 2016spotlight

This paper considers the task of articulated human pose estimation of multiple people in real world images. We propose an approach that jointly solves the tasks of detection and pose estimation: it infers the number of persons in a scene, identifies occluded body parts, and disambiguates body parts…

Cited by 1446PDFScholar
2016

Generative Adversarial Text to Image Synthesis

ICML 2016poster

Automatic synthesis of realistic images from text would be interesting and useful, but current AI systems are still far from this goal. However, in recent years generic and powerful recurrent neural network architectures have been developed to learn discriminative text feature representations. Meanw…

2016

How Far Are We From Solving Pedestrian Detection?

CVPR 2016poster

Encouraged by the recent progress in pedestrian detection, we investigate the gap between current state-of-the-art methods and the "perfect single frame detector". We enable our analysis by creating a human baseline for pedestrian detection (over the Caltech dataset), and by manually clustering the…

Cited by 597PDFScholar
2016

Latent Embeddings for Zero-Shot Classification

CVPR 2016spotlight

We present a novel latent embedding model for learning a compatibility function between image and class embeddings, in the context of zero-shot classification. The proposed method augments the state-of-the-art bilinear compatibility model by incorporating latent variables. Instead of learning a sing…

Cited by 888PDFScholar
2016

Learning Deep Representations of Fine-Grained Visual Descriptions

CVPR 2016spotlight

State-of-the-art methods for zero-shot visual recognition formulate learning as a joint embedding problem of images and side information. In these formulations the current best complement to visual features are attributes: manually-encoded vectors describing shared characteristics among categories.…

Cited by 1105PDFcodeScholar
2016

Learning What and Where to Draw

NeurIPS 2016oral

Generative Adversarial Networks (GANs) have recently demonstrated the capability to synthesize compelling real-world images, such as room interiors, album covers, manga, faces, birds, and flowers. While existing models can synthesize images based on global constraints such as a class label or captio…

2016

The Cityscapes Dataset for Semantic Urban Scene Understanding

CVPR 2016spotlight

Visual understanding of complex urban street scenes is an enabling factor for a wide range of applications. Object detection has benefited enormously from large-scale datasets, especially in the context of deep learning. For semantic urban scene understanding, however, no current dataset adequately…

Cited by 15494PDFScholar
2015

Classifier Based Graph Construction for Video Segmentation

CVPR 2015poster

Video segmentation has become an important and active research area with a large diversity of proposed approaches. Graph-based methods, enabling topperformance on recent benchmarks, consist of three essential components: 1. powerful features account for object appearance and motion similarities; 2.…

Cited by 85SourcePDFScholar
2015

Efficient ConvNet-Based Marker-Less Motion Capture in General Scenes With a Low Number of Cameras

CVPR 2015poster

We present a novel method for accurate marker-less capture of articulated skeleton motion of several subjects in general scenes, indoors and outdoors, even from input filmed with as few as two cameras. Our approach unites a discriminative image-based joint detection method with a model-based generat…

Cited by 192SourcePDFScholar
2015

Efficient Output Kernel Learning for Multiple Tasks

NeurIPS 2015poster

The paradigm of multi-task learning is that one can achieve better generalization by learning tasks jointly and thus exploiting the similarity between the tasks rather than learning them independently of each other. While previously the relationship between tasks had to be user-defined in the form o…

Cited by 39SourcePDFScholar
2015

Evaluation of Output Embeddings for Fine-Grained Image Classification

CVPR 2015poster

Image classification has advanced significantly in recent years with the availability of large-scale image sets. However, fine-grained classification remains a major challenge due to the annotation cost of large numbers of fine-grained categories. This project shows that compelling classification pe…

Cited by 1290SourcePDFScholar