← Search

Hakan Bilen

42 accepted papers

2026

Beyond Pixel Context Windows: Neural World Simulators with Persistent 3D State

ICML 2026poster

Interactive world models continually generate video by responding to a user's actions, enabling open-ended generation capabilities. However, existing models typically lack a 3D representation of the environment, meaning 3D consistency must be implicitly learned from data, and spatial memory is restr…

Cited by 0SourceScholar
2025

DepthCues: Evaluating Monocular Depth Perception in Large Vision Models

CVPR 2025poster

Large-scale pre-trained vision models are becoming increasingly prevalent, offering expressive and generalizable visual representations that benefit various downstream tasks. Recent studies on the emergent properties of these models have revealed their high-level geometric understanding, in particul…

Cited by 3SourcePDFScholar
2025

Interactive Anomaly Detection for Articulated Objects via Motion Anticipation

NeurIPS 2025poster

This paper presents a novel problem, interactive anomaly detection (AD) for articulated objects, and introduces a tailored solution that detects functional anomalies by integrating vision, interaction, and anticipation. Unlike traditional AD methods that rely on passive visual observations, our appr…

Cited by 0SourceScholar
2025

Jamais Vu: Exposing the Generalization Gap in Supervised Semantic Correspondence

NeurIPS 2025poster

Semantic correspondence (SC) aims to establish semantically meaningful matches across different instances of an object category. We illustrate how recent supervised SC methods remain limited in their ability to generalize beyond sparsely annotated training keypoints, effectively acting as keypoint d…

Cited by 0SourceScholar
2024

Articulate your NeRF: Unsupervised articulated object modeling via conditional view synthesis

NeurIPS 2024poster

We propose a novel unsupervised method to learn pose and part-segmentation of articulated objects with rigid parts. Given two observations of an object in different articulation states, our method learns the geometry and appearance of object parts by using an implicit model from the first observ…

Cited by 3SourcePDFScholar
2024

Improving Semantic Correspondence with Viewpoint-Guided Spherical Maps

CVPR 2024poster

Recent self-supervised models produce visual features that are not only effective at encoding image-level but also pixel-level semantics. They have been reported to obtain impressive results for dense visual semantic correspondence estimation even outperforming fully-supervised methods. Nevertheless…

Cited by 12SourcePDFScholar
2024

Multi-task Learning with 3D-Aware Regularization

ICLR 2024poster

Deep neural networks have become the standard solution for designing models that can perform multiple dense computer vision tasks such as depth estimation and semantic segmentation thanks to their ability to capture complex correlations in high dimensional feature space across tasks. However, the cr…

2024

RECANTFormer: Referring Expression Comprehension with Varying Numbers of Targets

EMNLP 2024main

The Generalized Referring Expression Comprehension (GREC) task extends classic REC by generating image bounding boxes for objects referred to in natural language expressions, which may indicate zero, one, or multiple targets. This generalization enhances the practicality of REC models for diverse re…

2023

Learning Action Changes by Measuring Verb-Adverb Textual Relationships

CVPR 2023poster

The goal of this work is to understand the way actions are performed in videos. That is, given a video, we aim to predict an adverb indicating a modification applied to the action (e.g. cut "finely"). We cast this problem as a regression task. We measure textual relationships between verbs and adver…

2023

RenderDiffusion: Image Diffusion for 3D Reconstruction, Inpainting and Generation

CVPR 2023poster

Diffusion models currently achieve state-of-the-art performance for both conditional and unconditional image generation. However, so far, image diffusion models do not support tasks required for 3D understanding, such as view-consistent 3D generation or single-view object reconstruction. In this pap…

2023

Semi-supervised multimodal coreference resolution in image narrations

EMNLP 2023long main

In this paper, we study multimodal coreference resolution, specifically where a longer descriptive text, i.e., a narration is paired with an image. This poses significant challenges due to fine-grained image-text alignment, inherent ambiguity present in narrative language, and unavailability of larg…

Cited by 0SourcecodeScholar
2023

Who Are You Referring To? Coreference Resolution In Image Narrations

ICCV 2023poster

Coreference resolution aims to identify words and phrases which refer to the same entity in a text, a core task in natural language processing. In this paper, we extend this task to resolving coreferences in long-form narrations of visual scenes. First, we introduce a new dataset with annotated core…

Cited by 4PDFScholar
2022

3D Equivariant Graph Implicit Functions

ECCV 2022poster

"In recent years, neural implicit representations have made remarkable progress in modeling of 3D shapes with arbitrary topology. In this work, we address two key limitations of such representations, in failing to capture local 3D geometric fine details, and to learn from and generalize to shapes wi…

2022

CAFE: Learning To Condense Dataset by Aligning Features

CVPR 2022poster

Dataset condensation aims at reducing the network training effort through condensing a cumbersome training set into a compact synthetic one. State-of-the-art approaches largely rely on learning the synthetic data by matching the gradients between the real and synthetic data batches. Despite the intu…

Cited by 277PDFcodeScholar
2022

Distilling Representations from GAN Generator via Squeeze and Span

NeurIPS 2022accept

In recent years, generative adversarial networks (GANs) have been an actively studied topic and shown to successfully produce high-quality realistic images in various domains. The controllable synthesis ability of GAN generators suggests that they maintain informative, disentangled, and explainable…

2022

Learning to Annotate Part Segmentation with Gradient Matching

ICLR 2022poster

The success of state-of-the-art deep neural networks heavily relies on the presence of large-scale labelled datasets, which are extremely expensive and time-consuming to annotate. This paper focuses on tackling semi-supervised part segmentation tasks by generating high-quality images with a pre-trai…

2022

Not All Relations Are Equal: Mining Informative Labels for Scene Graph Generation

CVPR 2022poster

Scene graph generation (SGG) aims to capture a wide variety of interactions between pairs of objects, which is essential for full scene understanding. Existing SGG methods trained on the entire set of relations fail to acquire complex reasoning about visual and textual correlations due to various bi…

Cited by 38PDFScholar
2021

Neural Feature Matching in Implicit 3D Representations

ICML 2021spotlight

Recently, neural implicit functions have achieved impressive results for encoding 3D shapes. Conditioning on low-dimensional latent codes generalises a single implicit function to learn shared representation space for a variety of shapes, with the advantage of smooth interpolation. While the benefit…

2020

Self-Supervised Learning of Interpretable Keypoints From Unlabelled Videos

CVPR 2020oral

We propose a new method for recognizing the pose of objects from a single image that for learning uses only unlabelled videos and a weak empirical prior on the object poses. Video frames differ primarily in the pose of the objects they contain, so our method distils the pose information by analyzing…

Cited by 100PDFScholar
2019

Unsupervised Learning of Landmarks by Descriptor Vector Exchange

ICCV 2019poster

Equivariance to random image transformations is an effective method to learn landmarks of object categories, such as the eyes and the nose in faces, without manual supervision. However, this method does not explicitly guarantee that the learned landmarks are consistent with changes between different…

Cited by 84PDFcodeScholar
2018

Efficient Parametrization of Multi-Domain Deep Neural Networks

CVPR 2018poster

A practical limitation of deep neural networks is their high degree of specialization to a single task and visual domain. In complex applications such as mobile platforms, this requires juggling several large models with detrimental effect on speed and battery life. Recently, inspired by the success…

Cited by 387SourcePDFScholar
2018

Modelling and unsupervised learning of symmetric deformable object categories

NeurIPS 2018poster

We propose a new approach to model and learn, without manual supervision, the symmetries of natural objects, such as faces or flowers, given only images as input. It is well known that objects that have a symmetric structure do not usually result in symmetric images due to articulation and perspecti…

Cited by 15SourcePDFScholar
2018

Unsupervised Learning of Object Landmarks through Conditional Image Generation

NeurIPS 2018poster

We propose a method for learning landmark detectors for visual objects (such as the eyes and the nose in a face) without any manual supervision. We cast this as the problem of generating images that combine the appearance of the object as seen in a first example image with the geometry of the object…

Cited by 285SourcePDFScholar
2017

Learning multiple visual domains with residual adapters

NeurIPS 2017spotlight

There is a growing interest in learning data representations that work well for many different types of problems and data. In this paper, we look in particular at the task of learning a single visual representation that can be successfully utilized in the analysis of very different types of images,…

2017

Self-Supervised Video Representation Learning With Odd-One-Out Networks

CVPR 2017poster

We propose a new self-supervised CNN pre-training technique based on a novel auxiliary task called odd-one-out learning. In this task, the machine is asked to identify the unrelated or odd element from a set of otherwise related elements. We apply this technique to self-supervised video representati…

Cited by 562PDFScholar
2017

Unsupervised learning of object frames by dense equivariant image labelling

NeurIPS 2017oral

One of the key challenges of visual perception is to extract abstract models of 3D objects and object categories from visual measurements, which are affected by complex nuisance factors such as viewpoint, occlusion, motion, and deformations. Starting from the recent idea of viewpoint factorization,…

Cited by 113SourcePDFScholar
2016

Dynamic Image Networks for Action Recognition

CVPR 2016oral

We introduce the concept of dynamic image, a novel compact representation of videos useful for video analysis especially when convolutional neural networks (CNNs) are used. The dynamic image is based on the rank pooling concept and is obtained through the parameters of a ranking machine that encodes…

Cited by 727PDFcodeScholar