← Search

Zeynep Akata

92 accepted papers

2026

Are Reasoning LLMs Robust to Interventions on their Chain-of-Thought?

ICLR 2026poster

Reasoning LLMs (RLLMs) generate step-by-step chains of thought (CoTs) before giving an answer, which improves performance on complex tasks and makes reasoning transparent. But how robust are these reasoning traces to disruptions that occur within them? To address this question, we introduce a contro…

Cited by 0SourcecodeScholar
2026

Attentive Multi-Layer Fusion for Vision Transformers

ICML 2026poster

With the rise of large-scale foundation models, efficiently adapting them to downstream tasks remains a central challenge. Linear probing, which freezes the backbone and trains a lightweight head, is computationally efficient but often restricted to last-layer representations. We show that task-rele…

Cited by 0SourceScholar
2026

Diamond Maps: Efficient Reward Alignment via Stochastic Flow Maps

ICML 2026poster

Flow and diffusion models produce high-quality samples, but adapting them to user preferences or constraints post-training remains costly and brittle, a challenge commonly called reward alignment. We argue that efficient reward alignment should be a property of the generative model itself, not an af…

Cited by 0SourceScholar
2026

Dissecting Multimodal In-Context Learning: Modality Asymmetries and Circuit Dynamics in modern Transformers

ICML 2026spotlight

Transformer-based multimodal large language models often exhibit in-context learning (ICL) capabilities. Motivated by this phenomenon, we ask: how do transformers learn to associate information across modalities from in-context examples? We investigate this through controlled experiments on small tr…

Cited by 1SourceScholar
2026

Explaining CLIP Zero-shot Predictions Through Concepts

CVPR 2026

Large-scale vision-language models such as CLIP have achieved remarkable success in zero-shot image recognition, yet their predictions remain largely opaque to human understanding. In contrast, Concept Bottleneck Models provide interpretable intermediate representations by reasoning through human-de

Cited by 0SourcecodeScholar
2026

FINER: MLLMs Hallucinate under Fine-grained Negative Queries

CVPR 2026

Multimodal large language models (MLLMs) struggle with hallucinations, particularly with fine-grained queries, a challenge underrepresented by existing benchmarks that focus on coarse image-related questions. We introduce **FI**ne-grained **NE**gative que**R**ies (**FINER**), alongside two benchmark

Cited by 0SourcecodeScholar
2026

Mantis: Lightweight Foundation Model for Time Series Classification

ICML 2026poster

While foundation models have revolutionized various domains, their application to time series classification remains rather under-explored, with existing literature predominantly focused on forecasting. To bridge this gap, we introduce \textbf{Mantis}, a transformer-based foundation model pre-traine…

Cited by 0SourceScholar
2026

Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models

ICLR 2026poster

Vision-language models trained on large-scale multimodal datasets show strong demographic biases, but the role of training data in producing these biases remains unclear. A major barrier has been the lack of demographic annotations in web-scale datasets such as LAION-400M. We address this gap by cre…

Cited by 0SourceScholar
2026

Post-hoc Probabilistic Vision-Language Models

ICLR 2026poster

Vision-language models (VLMs), such as CLIP and SigLIP, have found remarkable success in classification, retrieval, and generative tasks. For this, VLMs deterministically map images and text descriptions to a joint latent space in which their similarity is assessed using the cosine similarity. Howev…

Cited by 0SourcecodeScholar
2026

Rethinking Concept Bottleneck Models: From Pitfalls to Solutions

CVPR 2026

Concept Bottleneck Models (CBMs) ground predictions in human-understandable concepts but face fundamental limitations: the absence of a metric to pre-evaluate concept relevance, the "linearity problem" causing recent CBMs to bypass the concept bottleneck entirely, an accuracy gap compared to opaque

Cited by 0SourcecodeScholar
2026

SOTAlign: Semi-Supervised Alignment of Unimodal Vision and Language Models via Optimal Transport

ICML 2026poster

The Platonic Representation Hypothesis posits that neural networks trained on different modalities converge toward a shared statistical model of the world. Recent work exploits this convergence by aligning frozen pretrained vision and language models with lightweight alignment layers, but typically …

Cited by 0SourceScholar
2026

The Latent Color Subspace: Emergent Order in High-Dimensional Chaos

ICML 2026poster

Text-to-image generation models have advanced rapidly, yet achieving fine-grained control over generated images remains difficult, largely due to limited understanding of how semantic information is encoded. We develop an interpretation of the color representation in the Variational Autoencoder late…

Cited by 0SourceScholar
2026

TimeSAE: Sparse Decoding for Faithful Explanations of Black-Box Time Series Models

ICML 2026poster

As black box models and pretrained models gain traction in time series applications, understanding and explaining their predictions becomes increasingly vital, especially in high-stakes domains where interpretability and trust are essential. However, most of the existing methods involve only in-dist…

Cited by 0SourceScholar
2025

Building, Reusing, and Generalizing Abstract Representations from Concrete Sequences

ICLR 2025poster

Humans excel at learning abstract patterns across different sequences, filtering out irrelevant details, and transferring these generalized concepts to new sequences. In contrast, many sequence learning models lack the ability to abstract, which leads to memory inefficiency and poor transfer. We int…

Cited by 0SourcePDFScholar
2025

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training

CVPR 2025poster

Vision-Language Models (VLMs) trained with contrastive loss have achieved significant advancements in various vision and language tasks. However, the global nature of the contrastive loss makes VLMs focus predominantly on foreground objects, neglecting other crucial information in the image, which l…

2025

Concept-Guided Interpretability via Neural Chunking

NeurIPS 2025poster

Neural networks are often described as black boxes, reflecting the significant challenge of understanding their internal workings and interactions. We propose a different perspective that challenges the prevailing view: rather than being inscrutable, neural networks exhibit patterns in their raw po…

Cited by 0SourcecodeScholar
2025

Context-Aware Multimodal Pretraining

CVPR 2025highlight

Large-scale multimodal representation learning successfully optimizes for zero-shot transfer at test time. Yet the standard pretraining paradigm (contrastive learning on large amounts of image-text data) does not explicitly encourage representations to support few-shot adaptation. In this work, we p…

2025

Disentangled Representation Learning with the Gromov-Monge Gap

ICLR 2025poster

Learning disentangled representations from unlabelled data is a fundamental challenge in machine learning. Solving it may unlock other problems, such as generalization, interpretability, or fairness. Although remarkably challenging to solve in theory, disentanglement is often achieved in practice th…

Cited by 0SourcePDFScholar
2025

FLAIR: VLM with Fine-grained Language-informed Image Representations

CVPR 2025poster

CLIP has shown impressive results in aligning images and text at scale. However, its ability to capture detailed visual features remains limited because CLIP matches images and texts at a global level. To address this issue, we propose FLAIR, Fine-grained Language-informed Image Representations, an…

2025

How to Merge Your Multimodal Models Over Time?

CVPR 2025poster

Model merging combines expert models---each finetuned from a shared foundation model on diverse tasks and domains---into a single, more capable base model. However, existing model merging approaches assume all experts to be available simultaneously. In reality, new tasks and domains emerge continuou…

2025

Noise Hypernetworks: Amortizing Test-Time Compute in Diffusion Models

NeurIPS 2025poster

The new paradigm of test-time scaling has yielded remarkable breakthroughs in Large Language Models (LLMs) (e.g. reasoning models) and in generative vision models, allowing models to allocate additional computation during inference to effectively tackle increasingly complex problems. Despite the imp…

Cited by 0SourceScholar
2025

Revealing and Reducing Gender Biases in Vision and Language Assistants (VLAs)

ICLR 2025poster

Pre-trained large language models (LLMs) have been reliably integrated with visual input for multimodal tasks. The widespread adoption of instruction-tuned image-to-text vision-language assistants (VLAs) like LLaVA and InternVL necessitates evaluating gender biases. We study gender bias in 22 popula…

2025

SUB: Benchmarking CBM Generalization via Synthetic Attribute Substitutions

ICCV 2025poster

Concept Bottleneck Models (CBMs) and other concept-based interpretable models show great promise for making AI applications more transparent, which is essential in fields like medicine. Despite their success, we demonstrate that CBMs struggle to reliably identify the correct concepts under distribut…

2025

Scalable Ranked Preference Optimization for Text-to-Image Generation

ICCV 2025poster

Direct Preference Optimization (DPO) has emerged as a powerful approach to align text-to-image (T2I) models with human feedback. Unfortunately, successful application of DPO to T2I models requires a huge amount of resources to collect and label large-scale datasets, e.g., millions of generated paire…

Cited by 0SourcePDFScholar
2025

Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models

NeurIPS 2025poster

Sparse Autoencoders (SAEs) have recently gained attention as a means to improve the interpretability and steerability of Large Language Models (LLMs), both of which are essential for AI safety. In this work, we extend the application of SAEs to Vision-Language Models (VLMs), such as CLIP, and introd…

Cited by 0SourcecodeScholar
2025

WikiBigEdit: Understanding the Limits of Lifelong Knowledge Editing in LLMs

ICML 2025poster

Keeping large language models factually up-to-date is crucial for deployment, yet costly retraining remains a challenge. Knowledge editing offers a promising alternative, but methods are only tested on small-scale or synthetic edit benchmarks. In this work, we aim to bridge research into lifelong kn…

Cited by 0SourcePDFScholar
2024

A Practitioner's Guide to Real-World Continual Multimodal Pretraining

NeurIPS 2024poster

Multimodal foundation models serve numerous applications at the intersection of vision and language. Still, despite being pretrained on extensive data, they become outdated over time. To keep models updated, research into continual pretraining mainly explores scenarios with either (1) infrequent, in…

2024

DataDream: Few-shot Guided Dataset Generation

ECCV 2024poster

"While text-to-image diffusion models have been shown to achieve state-of-the-art results in image synthesis, they have yet to prove their effectiveness in downstream applications. Previous work has proposed to generate data for image classifier training given limited real data access. However, thes…

2024

ETHER: Efficient Finetuning of Large-Scale Models with Hyperplane Reflections

ICML 2024poster

Parameter-efficient finetuning (PEFT) has become ubiquitous to adapt foundation models to downstream task requirements while retaining their generalization ability. However, the amount of additionally introduced parameters and compute for successful adaptation and hyperparameter searches can explode…

2024

EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval

ECCV 2024poster

"In Composed Video Retrieval, a video and a textual description which modifies the video content are provided as inputs to the model. The aim is to retrieve the relevant video with the modified content from a database of videos. In this challenging task, the first step is to acquire large-scale trai…

2024

Fantastic Gains and Where to Find Them: On the Existence and Prospect of General Knowledge Transfer between Any Pretrained Model

ICLR 2024spotlight

Training deep networks requires various design decisions regarding for instance their architecture, data augmentation, or optimization. In this work, we find these training variations to result in networks learning unique feature sets from the data. Using public model libraries comprising thousands…

2024

Geometry Fidelity for Spherical Images

ECCV 2024poster

"Spherical or omni-directional images offer an immersive visual format appealing to a wide range of computer vision applications. However, geometric properties of spherical images pose a major challenge for models and metrics designed for ordinary 2D images. Here, we show that direct application of…

Cited by 2SourcePDFScholar
2024

Improving Intervention Efficacy via Concept Realignment in Concept Bottleneck Models

ECCV 2024poster

"Concept Bottleneck Models (CBMs) ground image classification on human-understandable concepts to allow for interpretable model decisions as well as human interventions, in which expert users can modify misaligned concept choices to interpretably influence the decision of the model. However, existin…

2024

ReNO: Enhancing One-step Text-to-Image Models through Reward-based Noise Optimization

NeurIPS 2024poster

Text-to-Image (T2I) models have made significant advancements in recent years, but they still struggle to accurately capture intricate details specified in complex compositional prompts. While fine-tuning T2I models with reward objectives has shown promise, it suffers from "reward hacking" and may n…

2024

Unbalancedness in Neural Monge Maps Improves Unpaired Domain Translation

ICLR 2024poster

In optimal transport (OT), a Monge map is known as a mapping that transports a source distribution to a target distribution in the most cost-efficient way. Recently, multiple neural estimators for Monge maps have been developed and applied in diverse unpaired domain translation tasks, e.g. in single…

2024

Vision-by-Language for Training-Free Compositional Image Retrieval

ICLR 2024poster

Given an image and a target modification (e.g an image of the Eiffel tower and the text “without people and at night-time”), Compositional Image Retrieval (CIR) aims to retrieve the relevant target image in a database. While supervised approaches rely on annotating triplets that is costly (i.e. quer…

2023

Bridging the Gap Between Model Explanations in Partially Annotated Multi-Label Classification

CVPR 2023poster

Due to the expensive costs of collecting labels in multi-label classification datasets, partially annotated multi-label classification has become an emerging field in computer vision. One baseline approach to this task is to assume unobserved labels as negative labels, but this assumption induces la…

2023

Disentanglement of Correlated Factors via Hausdorff Factorized Support

ICLR 2023poster

A grand goal in deep learning research is to learn representations capable of generalizing across distribution shifts. Disentanglement is one promising direction aimed at aligning a model's representation with the underlying factors generating the data (e.g. color or background). Existing disentangl…

2023

Image-Free Classifier Injection for Zero-Shot Classification

ICCV 2023poster

Zero-shot learning models achieve remarkable results on image classification for samples from classes that were not seen during training. However, such models must be trained from scratch with specialised methods: therefore, access to a training dataset is required when the need for zero-shot classi…

Cited by 16PDFcodeScholar
2023

In-Context Impersonation Reveals Large Language Models' Strengths and Biases

NeurIPS 2023spotlight

In everyday conversations, humans can take on different roles and adapt their vocabulary to their chosen roles. We explore whether LLMs can take on, that is impersonate, different roles when they generate text in-context. We ask LLMs to assume different personas before solving vision and language ta…

2023

Iterative Superquadric Recomposition of 3D Objects from Multiple Views

ICCV 2023poster

Humans are good at recomposing novel objects, i.e they can identify commonalities between unknown objects from general structure to finer detail, an ability difficult to replicate by machines. We propose a framework, ISCO, to recompose an object using 3D superquadrics as semantic parts directly from…

Cited by 10PDFcodeScholar
2023

Meta-in-context learning in large language models

NeurIPS 2023poster

Large language models have shown tremendous performance in a variety of tasks. In-context learning -- the ability to improve at a task after being provided with a number of demonstrations -- is seen as one of the main contributors to their success. In the present paper, we demonstrate that the in-…

2023

PDiscoNet: Semantically consistent part discovery for fine-grained recognition

ICCV 2023poster

Fine-grained classification often requires recognizing specific object parts, such as beak shape and wing patterns for birds. Encouraging a fine-grained classification model to first detect such parts and then using them to infer the class could help us gauge whether the model is indeed looking at t…

Cited by 16PDFcodeScholar
2023

ProbVLM: Probabilistic Adapter for Frozen Vison-Language Models

ICCV 2023poster

Large-scale vision-language models (VLMs) like CLIP successfully find correspondences between images and text. Through the standard deterministic mapping process, an image or a text sample is mapped to a single vector in the embedding space. This is problematic: as multiple samples (images or text)…

Cited by 31PDFScholar
2023

Transitivity Recovering Decompositions: Interpretable and Robust Fine-Grained Relationships

NeurIPS 2023poster

Recent advances in fine-grained representation learning leverage local-to-global (emergent) relationships for achieving state-of-the-art results. The relational representations relied upon by such methods, however, are abstract. We aim to deconstruct this abstraction by expressing them as interpreta…

2023

USIM-DAL: Uncertainty-aware Statistical Image Modeling-based Dense Active Learning for Super-resolution

UAI 2023poster

Dense regression is a widely used approach in computer vision for tasks such as image super-resolution, enhancement, depth estimation, etc. However, the high cost of annotation and labeling makes it challenging to achieve accurate results. We propose incorporating active learning into dense regressi…

Cited by 3SourcePDFScholar
2023

Waffling Around for Performance: Visual Classification with Random Words and Broad Concepts

ICCV 2023poster

The visual classification performance of vision-language models such as CLIP has been shown to benefit from additional semantic knowledge from large language models (LLMs) such as GPT-3. In particular, averaging over LLM-generated class descriptors, e.g. "waffle, which has a round shape", can notabl…

Cited by 86PDFcodeScholar
2022

A Non-Isotropic Probabilistic Take On Proxy-Based Deep Metric Learning

ECCV 2022poster

"Proxy-based Deep Metric Learning (DML) learns deep metric spaces by embedding images and class representatives (proxies) close to one another during training, as commonly measured by the angle between them. However, this disregards the embedding norm, which can carry additional beneficial context s…

2022

Abstracting Sketches through Simple Primitives

ECCV 2022poster

"Humans show high-level of abstraction capabilities in games that require quickly communicating object information. They decompose the message content into multiple parts and communicate them in an interpretable protocol. Toward equipping machines with such capabilities, we propose the Primitive-bas…

2022

Audio-Visual Generalised Zero-Shot Learning With Cross-Modal Attention and Language

CVPR 2022poster

Learning to classify video data from classes not included in the training data, i.e. video-based zero-shot learning, is challenging. We conjecture that the natural alignment between the audio and visual modalities in video data provides a rich training signal for learning discriminative multi-modal…

Cited by 72PDFcodeScholar
2022

BayesCap: Bayesian Identity Cap for Calibrated Uncertainty in Frozen Neural Networks

ECCV 2022poster

"High-quality calibrated uncertainty estimates are crucial for numerous real-world applications, especially for deep learning-based deployed ML systems. While Bayesian deep learning techniques allow uncertainty estimation, training them with large-scale datasets is an expensive process that does not…

2022

KG-SP: Knowledge Guided Simple Primitives for Open World Compositional Zero-Shot Learning

CVPR 2022poster

The goal of open-world compositional zero-shot learning(OW-CZSL) is to recognize compositions of state and objects in images, given only a subset of them during training and no prior on the unseen compositions. In this setting, models operate on a huge output space, containing all possible state-obj…

Cited by 63PDFcodeScholar
2022

Large Loss Matters in Weakly Supervised Multi-Label Classification

CVPR 2022poster

Weakly supervised multi-label classification (WSML) task, which is to learn a multi-label classification using partially observed labels per image, is becoming increasingly important due to its huge annotation cost. In this work, we first regard unobserved labels as negative labels, casting the WSML…

Cited by 79PDFcodeScholar
2022

PlanT: Explainable Planning Transformers via Object-Level Representations

CoRL 2022poster

Planning an optimal route in a complex environment requires efficient reasoning about the surrounding scene. While human drivers prioritize important objects and ignore details not relevant to the decision, learning-based planners typically extract features from dense, high-dimensional grid represen…

Cited by 114SourcecodeScholar
2022

Relational Proxies: Emergent Relationships as Fine-Grained Discriminators

NeurIPS 2022accept

Fine-grained categories that largely share the same set of parts cannot be discriminated based on part information alone, as they mostly differ in the way the local parts relate to the overall global structure of the object. We propose Relational Proxies, a novel approach that leverages the relation…

2022

Temporal and Cross-Modal Attention for Audio-Visual Zero-Shot Learning

ECCV 2022poster

"Audio-visual generalised zero-shot learning for video classification requires understanding the relations between the audio and visual information in order to be able to recognise samples from novel, previously unseen classes at test time. The natural semantic and temporal alignment between audio a…

2022

VGSE: Visually-Grounded Semantic Embeddings for Zero-Shot Learning

CVPR 2022poster

Human-annotated attributes serve as powerful semantic embeddings in zero-shot learning. However, their annotation process is labor-intensive and needs expert supervision. Current unsupervised semantic embeddings, i.e., word embeddings, enable knowledge transfer between classes. However, word embeddi…

Cited by 76PDFcodeScholar
2021

Distilling Audio-Visual Knowledge by Compositional Contrastive Learning

CVPR 2021poster

Having access to multi-modal cues (e.g. vision and audio) empowers some cognitive tasks to be done faster compared to learning from a single modality. In this work, we propose to transfer knowledge across heterogeneous modalities, even though these data modalities may not be semantically correlated.…

Cited by 95PDFcodeScholar
2021

E-ViL: A Dataset and Benchmark for Natural Language Explanations in Vision-Language Tasks

ICCV 2021poster

Recently, there has been an increasing number of efforts to introduce models capable of generating natural language explanations (NLEs) for their predictions on vision-language (VL) tasks. Such models are appealing, because they can provide human-friendly and comprehensive explanations. However, the…

Cited by 111PDFcodeScholar
2021

Fine-Grained Zero-Shot Learning with DNA as Side Information

NeurIPS 2021poster

Fine-grained zero-shot learning task requires some form of side-information to transfer discriminative information from seen to unseen classes. As manually annotated visual attributes are extremely costly and often impractical to obtain for a large number of classes, in this study we use DNA as a si…

2021

Learning Decision Trees Recurrently Through Communication

CVPR 2021poster

Integrated interpretability without sacrificing the prediction accuracy of decision making algorithms has the potential of greatly improving their value to the user. Instead of assigning a label to an image directly, we propose to learn iterative binary sub-decisions, inducing sparsity and transpare…

Cited by 20PDFcodeScholar
2021

Learning Graph Embeddings for Compositional Zero-Shot Learning

CVPR 2021poster

In compositional zero-shot learning, the goal is to recognize unseen compositions (e.g. old dog) of observed visual primitives states (e.g. old, cute) and objects (e.g. car, dog)in the training set. This is challenging because the same state can for example alter the visual appearance of a dog drast…

Cited by 199PDFcodeScholar
2021

Open World Compositional Zero-Shot Learning

CVPR 2021poster

Compositional Zero-Shot learning (CZSL) requires to recognize state-object compositions unseen during training. In this work, instead of assuming prior knowledge about the unseen compositions, we operate in the open world setting, where the search space includes a large number of unseen compositions…

Cited by 168PDFcodeScholar
2020

Attribute Prototype Network for Zero-Shot Learning

NeurIPS 2020poster

From the beginning of zero-shot learning research, visual attributes have been shown to play an important role. In order to better transfer attribute-based knowledge from known to unknown classes, we argue that an image representation with integrated attribute localization ability would be beneficia…

Cited by 378SourcePDFScholar
2020

Evaluating Weakly Supervised Object Localization Methods Right

CVPR 2020poster

Weakly-supervised object localization (WSOL) has gained popularity over the last years for its promise to train localization models with only image-level labels. Since the seminal WSOL work of class activation mapping (CAM), the field has focused on how to expand the attention regions to cover objec…

Cited by 237PDFcodeScholar
2020

Learning Robust Representations via Multi-View Information Bottleneck

ICLR 2020poster

The information bottleneck principle provides an information-theoretic method for representation learning, by training an encoder to retain all information which is relevant for predicting the label while minimizing the amount of other, excess information in the representation. The original formulat…

Cited by 323SourcecodeScholar
2020

Towards Recognizing Unseen Categories in Unseen Domains

ECCV 2020poster

Current deep visual recognition systems suffer from severe performance degradation when they encounter new images from classes and scenarios unseen during training. Hence, the core challenge of Zero-Shot Learning (ZSL) is to cope with the semantic-shift whereas the main challenge of Domain Adaptatio…

2019

Combining Generative and Discriminative Models for Hybrid Inference

NeurIPS 2019spotlight

A graphical model is a structured representation of the data generating process. The traditional method to reason over random variables is to perform inference in this graphical model. However, in many cases the generating process is only a poor approximation of the much more complex true data gener…

2019

F-VAEGAN-D2: A Feature Generating Framework for Any-Shot Learning

CVPR 2019poster

When labeled training data is scarce, a promising data augmentation approach is to generate visual features of unknown classes using their attributes. To learn the class conditional distribution of CNN features, these models rely on pairs of image features and class attributes. Hence, they can not m…

Cited by 647PDFScholar
2019

Generalized Zero- and Few-Shot Learning via Aligned Variational Autoencoders

CVPR 2019poster

Many approaches in generalized zero-shot learning rely on cross-modal mapping between the image feature space and the class embedding space. As labeled images are expensive, one direction is to augment the dataset by generating either images or image features. However, the former misses fine-grained…

Cited by 834PDFcodeScholar
2019

Modeling Conceptual Understanding in Image Reference Games

NeurIPS 2019spotlight

An agent who interacts with a wide population of other agents needs to be aware that there may be variations in their understanding of the world. Furthermore, the machinery which they use to perceive may be inherently different, as is the case between humans and machines. In this work, we present b…

2019

Semantic Projection Network for Zero- and Few-Label Semantic Segmentation

CVPR 2019poster

Semantic segmentation is one of the most fundamental problems in computer vision and pixel-level labelling in this context is particularly expensive. Hence, there have been several attempts to reduce the annotation effort such as learning from image level labels and bounding box annotations. In this…

Cited by 294PDFScholar
2018

Multimodal Explanations: Justifying Decisions and Pointing to the Evidence

CVPR 2018poster

Deep models that are both effective and explainable are desirable in many settings; prior explainable models have been unimodal, offering either image-based visualization of attention weights or text-based generation of post-hoc justifications. We propose a multimodal approach to explanation, and…

2018

Textual Explanations for Self-Driving Vehicles

ECCV 2018poster

Deep neural perception and control networks have become key components of self-driving vehicles. User acceptance is likely to benefit from easy-to-interpret textual explanations which allow end-users to understand what triggered a particular behavior. Explanations may be triggered by the neural cont…

2017

Exploiting Saliency for Object Segmentation From Image Level Labels

CVPR 2017poster

There have been remarkable improvements in the semantic labelling task in the recent years. However, the state of the art methods rely on large-scale pixel-level annotations. This paper studies the problem of training a pixel-wise semantic labeller network from image-level annotations of the present…

Cited by 236PDFScholar
2016

Generative Adversarial Text to Image Synthesis

ICML 2016poster

Automatic synthesis of realistic images from text would be interesting and useful, but current AI systems are still far from this goal. However, in recent years generic and powerful recurrent neural network architectures have been developed to learn discriminative text feature representations. Meanw…

2016

Latent Embeddings for Zero-Shot Classification

CVPR 2016spotlight

We present a novel latent embedding model for learning a compatibility function between image and class embeddings, in the context of zero-shot classification. The proposed method augments the state-of-the-art bilinear compatibility model by incorporating latent variables. Instead of learning a sing…

Cited by 888PDFScholar
2016

Learning Deep Representations of Fine-Grained Visual Descriptions

CVPR 2016spotlight

State-of-the-art methods for zero-shot visual recognition formulate learning as a joint embedding problem of images and side information. In these formulations the current best complement to visual features are attributes: manually-encoded vectors describing shared characteristics among categories.…

Cited by 1105PDFcodeScholar
2016

Learning What and Where to Draw

NeurIPS 2016oral

Generative Adversarial Networks (GANs) have recently demonstrated the capability to synthesize compelling real-world images, such as room interiors, album covers, manga, faces, birds, and flowers. While existing models can synthesize images based on global constraints such as a class label or captio…

2015

Evaluation of Output Embeddings for Fine-Grained Image Classification

CVPR 2015poster

Image classification has advanced significantly in recent years with the availability of large-scale image sets. However, fine-grained classification remains a major challenge due to the annotation cost of large numbers of fine-grained categories. This project shows that compelling classification pe…

Cited by 1290SourcePDFScholar