← Search

Caner Hazirbas

9 accepted papers

2024

UniBench: Visual Reasoning Requires Rethinking Vision-Language Beyond Scaling

NeurIPS 2024poster

Significant research efforts have been made to scale and improve vision-language model (VLM) training approaches. Yet, with an ever-growing number of benchmarks, researchers are tasked with the heavy burden of implementing each protocol, bearing a non-trivial computational cost, and making sense of…

2023

A Whac-a-Mole Dilemma: Shortcuts Come in Multiples Where Mitigating One Amplifies Others

CVPR 2023poster

Machine learning models have been found to learn shortcuts---unintended decision rules that are unable to generalize---undermining models' reliability. Previous works address this problem under the tenuous assumption that only a single shortcut exists in the training data. Real-world images are rife…

2023

Exploring Why Object Recognition Performance Degrades Across Income Levels and Geographies with Factor Annotations

NeurIPS 2023spotlight

Despite impressive advances in object-recognition, deep learning systems’ performance degrades significantly across geographies and lower income levels---raising pressing concerns of inequity. Addressing such performance gaps remains a challenge, as little is understood about why performance degrade…

Cited by 4SourcePDFScholar
2023

ImageNet-X: Understanding Model Mistakes with Factor of Variation Annotations

ICLR 2023top-25%

Deep learning vision systems are widely deployed across applications where reliability is critical. However, even today's best models can fail to recognize an object when its pose, lighting, or background varies. While existing benchmarks surface examples challenging for models, they do not explain…

Cited by 50SourcePDFScholar
2022

Generating High Fidelity Data From Low-Density Regions Using Diffusion Models

CVPR 2022poster

Our work focuses on addressing sample deficiency from low-density regions of data manifold in common image datasets. We leverage diffusion process based generative models to synthesize novel images from low-density regions. We observe that uniform sampling from diffusion models predominantly samples…

Cited by 71PDFcodeScholar
2022

Towards Measuring Fairness in Speech Recognition: Casual Conversations Dataset Transcriptions

ICASSP 2022accepted

The problem of machine learning systems demonstrating bias towards specific groups of individuals has been studied extensively, particularly in the Facial Recognition area, but much less so in Automatic Speech Recognition (ASR). This paper presents initial Speech Recognition results on “Casual Conve…

Cited by 53SourceScholar
2017

Image-Based Localization Using LSTMs for Structured Feature Correlation

ICCV 2017poster

In this work we propose a new CNN+LSTM architecture for camera pose regression for indoor and outdoor scenes. CNNs allow us to learn suitable feature representations for localization that are robust against motion blur and illumination changes. We make use of LSTM units on the CNN output, which play…

Cited by 658PDFScholar
2017

Learning Proximal Operators: Using Denoising Networks for Regularizing Inverse Imaging Problems

ICCV 2017poster

While variational methods have been among the most powerful tools for solving linear inverse problems in imaging, deep (convolutional) neural networks have recently taken the lead in many challenging benchmarks. A remaining drawback of deep learning approaches is their requirement for an expensive r…

Cited by 443PDFcodeScholar
2015

FlowNet: Learning Optical Flow With Convolutional Networks

ICCV 2015poster

Convolutional neural networks (CNNs) have recently been very successful in a variety of computer vision tasks, especially on those linked to recognition. Optical flow estimation has not been among the tasks CNNs succeeded at. In this paper we construct CNNs which are capable of solving the optical f…

Cited by 4909PDFScholar