← Search

Konstantinos G Derpanis

30 accepted papers

2026

Face2Scene: Using Facial Degradation as an Oracle for Diffusion-Based Scene Restoration

CVPR 2026

Recent advances in image restoration have enabled high-fidelity recovery of faces from degraded inputs using reference-based face restoration models (Ref-FR). However, such methods focus solely on facial regions, neglecting degradation across the full scene, including body and background, which limi

Cited by 0SourceScholar
2025

Revisiting Image Fusion for Multi-Illuminant White-Balance Correction

ICCV 2025poster

White balance (WB) correction in scenes with multiple illuminants remains a persistent challenge in computer vision. Recent methods explored fusion-based approaches, where a neural network linearly blends multiple sRGB versions of an input image, each processed with predefined WB presets. However, w…

Cited by 0SourcePDFScholar
2025

Universal Sparse Autoencoders: Interpretable Cross-Model Concept Alignment

ICML 2025poster

We present Universal Sparse Autoencoders (USAEs), a framework for uncovering and aligning interpretable concepts spanning multiple pretrained deep neural networks. Unlike existing concept-based interpretability methods, which focus on a single model, USAEs jointly learn a universal concept space tha…

Cited by 4SourcePDFScholar
2024

PolyOculus: Simultaneous Multi-view Image-based Novel View Synthesis

ECCV 2024poster

"This paper considers the problem of generative novel view synthesis (GNVS), generating novel, plausible views of a scene given a limited number of known views. Here, we propose a set-based generative model that can simultaneously generate multiple, self-consistent new views, conditioned on any numb…

2024

Understanding Video Transformers via Universal Concept Discovery

CVPR 2024highlight

This paper studies the problem of concept-based interpretability of transformer representations for videos. Concretely we seek to explain the decision-making process of video transformers based on high-level spatiotemporal concepts that are automatically discovered. Prior research on concept-based i…

Cited by 6SourcePDFScholar
2024

Visual Concept Connectome (VCC): Open World Concept Discovery and their Interlayer Connections in Deep Models

CVPR 2024highlight

Understanding what deep network models capture in their learned representations is a fundamental challenge in computer vision. We present a new methodology to understanding such vision models the Visual Concept Connectome (VCC) which discovers human interpretable concepts and their interlayer connec…

Cited by 8SourcePDFScholar
2024

Watch Your Steps: Local Image and Scene Editing by Text Instructions

ECCV 2024oral

"The success of denoising diffusion models in generating and editing images has sparked interest in using diffusion models for editing 3D scenes represented via neural radiance fields (NeRFs). However, current 3D editing methods lack a way to both pinpoint the edit location and limit changes to the…

Cited by 35SourcePDFScholar
2023

GePSAn: Generative Procedure Step Anticipation in Cooking Videos

ICCV 2023poster

We study the problem of future step anticipation in procedural videos. Given a video of an ongoing procedural activity, we predict a plausible next procedure step described in rich natural language. While most previous work focus on the problem of data scarcity in procedural video datasets, another…

Cited by 11PDFcodeScholar
2023

Long-Term Photometric Consistent Novel View Synthesis with Diffusion Models

ICCV 2023poster

Novel view synthesis from a single input image is a challenging task, where the goal is to generate a new view of a scene from a desired camera pose that may be separated by a large motion. The highly uncertain nature of this synthesis task due to unobserved elements within the scene (i.e. occlusion…

Cited by 39PDFcodeScholar
2023

Reference-guided Controllable Inpainting of Neural Radiance Fields

ICCV 2023poster

The popularity of Neural Radiance Fields (NeRFs) for view synthesis has led to a desire for NeRF editing tools. Here, we focus on inpainting regions in a view-consistent and controllable manner. In addition to the typical NeRF inputs and masks delineating the unwanted region in each view, we require…

Cited by 42PDFcodeScholar
2023

SPIn-NeRF: Multiview Segmentation and Perceptual Inpainting With Neural Radiance Fields

CVPR 2023poster

Neural Radiance Fields (NeRFs) have emerged as a popular approach for novel view synthesis. While NeRFs are quickly being adapted for a wider set of applications, intuitively editing NeRF scenes is still an open challenge. One important editing task is the removal of unwanted objects from a 3D scene…

2023

StepFormer: Self-Supervised Step Discovery and Localization in Instructional Videos

CVPR 2023poster

Instructional videos are an important resource to learn procedural tasks from human demonstrations. However, the instruction steps in such videos are typically short and sparse, with most of the video being irrelevant to the procedure. This motivates the need to temporally localize the instruction s…

Cited by 29SourcePDFScholar
2022

A Deeper Dive Into What Deep Spatiotemporal Networks Encode: Quantifying Static vs. Dynamic Information

CVPR 2022poster

Deep spatiotemporal models are used in a variety of computer vision tasks, such as action recognition and video object segmentation. Currently, there is a limited understanding of what information is captured by these models in their intermediate representations. For example, while it has been obser…

Cited by 22PDFcodeScholar
2022

P3IV: Probabilistic Procedure Planning From Instructional Videos With Weak Supervision

CVPR 2022oral

In this paper, we study the problem of procedure planning in instructional videos. Here, an agent must produce a plausible sequence of actions that can transform the environment from a given start to a desired goal state. When learning procedure planning from instructional videos, most recent work l…

Cited by 52PDFcodeScholar
2021

Drop-DTW: Aligning Common Signal Between Sequences While Dropping Outliers

NeurIPS 2021poster

In this work, we consider the problem of sequence-to-sequence alignment for signals containing outliers. Assuming the absence of outliers, the standard Dynamic Time Warping (DTW) algorithm efficiently computes the optimal alignment between two (generally) variable-length sequences. While DTW is ro…

Cited by 59SourcePDFScholar
2021

Global Pooling, More Than Meets the Eye: Position Information Is Encoded Channel-Wise in CNNs

ICCV 2021poster

In this paper, we challenge the common assumption that collapsing the spatial dimensions of a 3D (spatial-channel) tensor in a convolutional neural network (CNN) into a vector via global pooling removes all spatial information. Specifically, we demonstrate that positional information is encoded base…

Cited by 46PDFcodeScholar
2021

Learning Multi-Scale Photo Exposure Correction

CVPR 2021poster

Capturing photographs with wrong exposures remains a major source of errors in camera-based imaging. Exposure problems are categorized as either: (i) overexposed, where the camera exposure was too long, resulting in bright and washed-out image regions, or (ii) underexposed, where the exposure was to…

Cited by 241PDFcodeScholar
2021

Representation Learning via Global Temporal Alignment and Cycle-Consistency

CVPR 2021poster

We introduce a weakly supervised method for representation learning based on aligning temporal sequences (e.g., videos) of the same process (e.g., human action). The main idea is to use the global temporal ordering of latent correspondences across sequence pairs as a supervisory signal. In particula…

Cited by 71PDFcodeScholar
2021

Shape or Texture: Understanding Discriminative Features in CNNs

ICLR 2021poster

Contrasting the previous evidence that neurons in the later layers of a Convolutional Neural Network (CNN) respond to complex object shapes, recent studies have shown that CNNs actually exhibit a 'texture bias': given an image with both texture and shape cues (e.g., a stylized image), a CNN is biase…

Cited by 89SourcePDFScholar
2021

Stochastic Image-to-Video Synthesis Using cINNs

CVPR 2021poster

Video understanding calls for a model to learn the characteristic interplay between static scene content and its dynamics: Given an image, the model must be able to predict a future progression of the portrayed scene and, conversely, a video should be explained in terms of its static image content a…

Cited by 67PDFcodeScholar
2020

RankMI: A Mutual Information Maximizing Ranking Loss

CVPR 2020poster

We introduce an information-theoretic loss function, RankMI, and an associated training algorithm for deep representation learning for image retrieval. Our proposed framework consists of alternating updates to a network that estimates the divergence between distance distributions of matching and non…

Cited by 53PDFScholar
2020

Wavelet Flow: Fast Training of High Resolution Normalizing Flows

NeurIPS 2020poster

Normalizing flows are a class of probabilistic generative models which allow for both fast density computation and efficient sampling and are effective at modelling complex distributions like images. A drawback among current methods is their significant training cost, sometimes requiring months of G…

2019

End-to-End Learning of Representations for Asynchronous Event-Based Data

ICCV 2019poster

Event cameras are vision sensors that record asynchronous streams of per-pixel brightness changes, referred to as "events". They have appealing advantages over frame based cameras for computer vision, including high temporal resolution, high dynamic range, and no motion blur. Due to the sparse, non-…

Cited by 418PDFcodeScholar
2019

Learning what you can do before doing anything

ICLR 2019poster

Intelligent agents can learn to represent the action spaces of other agents simply by observing them act. Such representations help agents quickly learn to predict the effects of their own actions on the environment and to plan complex action sequences. In this work, we address the problem of learni…

2018

Two-Stream Convolutional Networks for Dynamic Texture Synthesis

CVPR 2018poster

We introduce a two-stream model for dynamic texture synthesis. Our model is based on pre-trained convolutional networks (ConvNets) that target two independent tasks: (i) object recognition, and (ii) optical flow prediction. Given an input dynamic texture, statistics of filter responses from the obje…

Cited by 64SourcePDFScholar
2017

6-DoF object pose from semantic keypoints

ICRA 2017poster

This paper presents a novel approach to estimating the continuous six degree of freedom (6-DoF) pose (3D translation and rotation) of an object from a single RGB image. The approach combines semantic keypoints predicted by a convolutional network (convnet) with a deformable shape model. Unlike prior…

Cited by 540SourceScholar
2017

Coarse-To-Fine Volumetric Prediction for Single-Image 3D Human Pose

CVPR 2017spotlight

This paper addresses the challenge of 3D human pose estimation from a single color image. Despite the general success of the end-to-end learning paradigm, top performing approaches employ a two-step solution consisting of a Convolutional Network (ConvNet) for 2D joint localization and a subsequent o…

Cited by 1177PDFScholar
2017

Harvesting Multiple Views for Marker-Less 3D Human Pose Annotations

CVPR 2017spotlight

Recent advances with Convolutional Networks (ConvNets) have shifted the bottleneck for many computer vision tasks to annotated data collection. In this paper, we present a geometry-driven approach to automatically collect annotations for human pose prediction tasks. Starting from a generic ConvNet f…

Cited by 248PDFScholar
2017

Segmentation-Aware Convolutional Networks Using Local Attention Masks

ICCV 2017poster

We introduce an approach to integrate segmentation information within a convolutional neural network (CNN). This counter-acts the tendency of CNNs to smooth information across regions and increases their spatial precision. To obtain segmentation information, we set up a CNN to provide an embedding s…

Cited by 188PDFcodeScholar
2016

Sparseness Meets Deepness: 3D Human Pose Estimation From Monocular Video

CVPR 2016spotlight

This paper addresses the challenge of 3D full-body human pose estimation from a monocular image sequence. Here, two cases are considered: (i) the image locations of the human joints are provided and (ii) the image locations of joints are unknown. In the former case, a novel approach is introduced th…

Cited by 545PDFScholar