← Search

Dimitris Samaras

69 accepted papers

2026

Generating metamers of human scene understanding

ICLR 2026oral

Human vision combines low-resolution “gist” information from the visual periphery with sparse but high-resolution information from fixated locations to construct a coherent understanding of a visual scene. In this paper, we introduce MetamerGen, a tool for generating scenes that are aligned with lat…

Cited by 0SourceScholar
2026

Personalized Image Descriptions from Attention Sequences

CVPR 2026

People can view the same image differently: they focus on different regions, objects, and details in varying orders and describe them in distinct linguistic styles. This leads to substantial variability in image descriptions. However, existing models for personalized image description focus on lingu

Cited by 0SourcecodeScholar
2025

2DMamba: Efficient State Space Model for Image Representation with Applications on Giga-Pixel Whole Slide Image Classification

CVPR 2025poster

Efficiently modeling large 2D contexts is essential for various fields including Giga-Pixel Whole Slide Imaging (WSI) and remote sensing. Transformer-based models offer high parallelism but face challenges due to their quadratic complexity for handling long sequences. Recently, Mamba introduced a se…

2025

AV-Flow: Transforming Text to Audio-Visual Human-like Interactions

ICCV 2025poster

We introduce AV-Flow, an audio-visual generative model that animates photo-realistic 4D talking avatars given only text input. In contrast to prior work that assumes an existing speech signal, we synthesize speech and vision jointly. We demonstrate human-like speech synthesis, synchronized lip motio…

Cited by 0SourcePDFScholar
2025

Fast constrained sampling in pre-trained diffusion models

NeurIPS 2025poster

Large denoising diffusion models, such as Stable Diffusion, have been trained on billions of image-caption pairs to perform text-conditioned image generation. As a byproduct of this training, these models have acquired general knowledge about image statistics, which can be useful for other inference…

Cited by 0SourcecodeScholar
2025

Few-shot Personalized Scanpath Prediction

CVPR 2025poster

A personalized model for scanpath prediction provides insights into the visual preferences and attention patterns of individual subjects. However, existing methods for training scanpath prediction models are data-intensive and cannot be effectively personalized to new individuals with only a few ava…

2025

GECKO: Gigapixel Vision-Concept Contrastive Pretraining in Histopathology

ICCV 2025poster

Pretraining a Multiple Instance Learning (MIL) aggregator enables the derivation of Whole Slide Image (WSI)-level embeddings from patch-level representations without supervision. While recent multimodal MIL pretraining approaches leveraging auxiliary modalities have demonstrated performance gains ov…

2025

Hummingbird: High Fidelity Image Generation via Multimodal Context Alignment

ICLR 2025poster

While diffusion models are powerful in generating high-quality, diverse synthetic data for object-centric tasks, existing methods struggle with scene-aware tasks such as Visual Question Answering (VQA) and Human-Object Interaction (HOI) Reasoning, where it is critical to preserve scene attributes in…

2025

Importance-Based Token Merging for Efficient Image and Video Generation

ICCV 2025poster

Token merging can effectively accelerate various vision systems by processing groups of similar tokens only once and sharing the results across them. However, existing token grouping methods are often ad hoc and random, disregarding the actual content of the samples. We show that preserving high-inf…

Cited by 0SourcePDFScholar
2025

Leveraging Registers in Vision Transformers for Robust Adaptation

ICASSP 2025accepted

Vision Transformers (ViTs) have shown success across a variety of tasks due to their ability to capture global image representations. Recent studies have identified the existence of high-norm tokens in ViTs, which can interfere with unsupervised object discovery. To address this, the use of "registe…

Cited by 3SourceScholar
2025

Low-Rank Head Avatar Personalization with Registers

NeurIPS 2025poster

We introduce a novel method for low-rank personalization of a generic model for head avatar generation. Prior work proposes generic models that achieve high-quality face animation by leveraging large-scale datasets of multiple identities. However, such generic models usually fail to synthesize uniqu…

Cited by 0SourceScholar
2025

Multi-view Gaze Target Estimation

ICCV 2025poster

This paper presents a method that utilizes multiple camera views for the gaze target estimation (GTE) task. The approach integrates information from different camera views to improve accuracy and expand applicability, addressing limitations in existing single-view methods that face challenges such a…

Cited by 0SourcePDFScholar
2025

TopoCellGen: Generating Histopathology Cell Topology with a Diffusion Model

CVPR 2025poster

Accurately modeling multi-class cell topology is crucial in digital pathology, as it provides critical insights into tissue structure and pathology. The synthetic generation of cell topology enables realistic simulations of complex tissue environments, enhances downstream tasks by augmenting trainin…

2025

ZoomLDM: Latent Diffusion Model for Multi-scale Image Generation

CVPR 2025poster

Diffusion models have revolutionized image generation, yet several challenges restrict their application to large-image domains, such as digital pathology and satellite imagery. Given that it is infeasible to directly train a model on 'whole' images from domains with potential gigapixel sizes, diffu…

2024

Beyond Pixels: Semi-Supervised Semantic Segmentation with a Multi-scale Patch-based Multi-Label Classifier

ECCV 2024poster

"Incorporating pixel contextual information is critical for accurate segmentation. In this paper, we show that an effective way to incorporate contextual information is through a patch-based classifier. This patch classifier is trained to identify classes present within an image region, which facili…

2024

Diffusion-Refined VQA Annotations for Semi-Supervised Gaze Following

ECCV 2024poster

"Training gaze following models requires a large number of images with gaze target coordinates annotated by human annotators, which is a laborious and inherently ambiguous process. We propose the first semi-supervised method for gaze following by introducing two novel priors to the task. We obtain t…

2024

Learned Representation-Guided Diffusion Models for Large-Image Generation

CVPR 2024poster

To synthesize high-fidelity samples diffusion models typically require auxiliary data to guide the generation process. However it is impractical to procure the painstaking patch-level annotation effort required in specialized domains like histopathology and satellite imagery; it is often performed b…

2024

Look Hear: Gaze Prediction for Speech-directed Human Attention

ECCV 2024poster

"For computer systems to effectively interact with humans using spoken language, they need to understand how the words being generated affect the users’ moment-by-moment attention. Our study focuses on the incremental prediction of attention as a person is seeing an image and hearing a referring exp…

2024

MIGS: Multi-Identity Gaussian Splatting via Tensor Decomposition

ECCV 2024oral

"We introduce (Multi-Identity Gaussian Splatting), a novel method that learns a single neural representation for multiple identities, using only monocular videos. Recent 3D Gaussian Splatting (3DGS) approaches for human avatars require per-identity optimization. However, learning a multi-identity re…

2024

SI-MIL: Taming Deep MIL for Self-Interpretability in Gigapixel Histopathology

CVPR 2024poster

Introducing interpretability and reasoning into Multiple Instance Learning (MIL) methods for Whole Slide Image (WSI) analysis is challenging given the complexity of gigapixel slides. Traditionally MIL interpretability is limited to identifying salient regions deemed pertinent for downstream tasks of…

2024

Self-supervised co-salient object detection via feature correspondences at multiple scales

ECCV 2024poster

"Our paper introduces a novel two-stage self-supervised approach for detecting co-occurring salient objects (CoSOD) in image groups without requiring segmentation annotations. Unlike existing unsupervised methods that rely solely on patch-level information (clustering patch descriptors) or on comput…

2024

Unifying Top-down and Bottom-up Scanpath Prediction Using Transformers

CVPR 2024poster

Most models of visual attention aim at predicting either top-down or bottom-up control as studied using different visual search and free-viewing tasks. In this paper we propose the Human Attention Transformer (HAT) a single model that predicts both forms of attention control. HAT uses a novel transf…

2024

Weighting Pseudo-Labels via High-Activation Feature Index Similarity and Object Detection for Semi-Supervised Segmentation

ECCV 2024poster

"Semi-supervised semantic segmentation methods leverage unlabeled data by pseudo-labeling them. Thus the success of these methods hinges on the reliability of the pseudo-labels. Existing methods mostly choose high-confidence pixels in an effort to avoid erroneous pseudo-labels. However, high confide…

2024

∞-Brush: Controllable Large Image Synthesis with Diffusion Models in Infinite Dimensions

ECCV 2024poster

"Synthesizing high-resolution images from intricate, domain-specific information remains a significant challenge in generative modeling, particularly for applications in large-image domains such as digital histopathology and remote sensing. Existing methods face critical limitations: conditional dif…

2023

Gazeformer: Scalable, Effective and Fast Prediction of Goal-Directed Human Attention

CVPR 2023poster

Predicting human gaze is important in Human-Computer Interaction (HCI). However, to practically serve HCI applications, gaze prediction models must be scalable, fast, and accurate in their spatial and temporal gaze predictions. Recent scanpath prediction models focus on goal-directed attention (sear…

2023

Generating Features With Increased Crop-Related Diversity for Few-Shot Object Detection

CVPR 2023poster

Two-stage object detectors generate object proposals and classify them to detect objects in images. These proposals often do not perfectly contain the objects but overlap with them in many possible ways, exhibiting great variability in the difficulty levels of the proposals. Training a robust classi…

Cited by 45SourcePDFScholar
2023

Learning Probabilistic Topological Representations Using Discrete Morse Theory

ICLR 2023top-25%

Accurate delineation of fine-scale structures is a very important yet challenging problem. Existing methods use topological information as an additional training loss, but are ultimately making pixel-wise predictions. In this paper, we propose a novel deep learning based method to learn topological/…

Cited by 19SourcePDFScholar
2023

S-VolSDF: Sparse Multi-View Stereo Regularization of Neural Implicit Surfaces

ICCV 2023poster

Neural rendering of implicit surfaces performs well in 3D vision applications. However, it requires dense input views as supervision. When only sparse input images are available, output quality drops significantly due to the shape-radiance ambiguity problem. We note that this ambiguity can be constr…

Cited by 18PDFScholar
2023

Topology-Guided Multi-Class Cell Context Generation for Digital Pathology

CVPR 2023poster

In digital pathology, the spatial context of cells is important for cell classification, cancer diagnosis and prognosis. To model such complex cell context, however, is challenging. Cells form different mixtures, lineages, clusters and holes. To model such structural patterns in a learnable fashion,…

Cited by 15SourcePDFScholar
2022

Diffusion Models as Plug-and-Play Priors

NeurIPS 2022accept

We consider the problem of inferring high-dimensional data $x$ in a model that consists of a prior $p(x)$ and an auxiliary differentiable constraint $c(x,y)$ on $x$ given some additional information $y$. In this paper, the prior is an independently trained denoising diffusion generative model. The a…

2022

Learning an Isometric Surface Parameterization for Texture Unwrapping

ECCV 2022poster

"In this paper, we present a novel approach to learn texture mapping for an isometrically deformed 3D surface and apply it for texture unwrapping of documents or other objects. Recent work on differentiable rendering techniques for implicit surfaces has shown high-quality 3D scene reconstruction and…

2022

Target-Absent Human Attention

ECCV 2022poster

"The prediction of human gaze behavior is important for building human-computer interactive systems that can anticipate a user’s attention. Computer vision models have been developed to predict the fixations made by people as they search for target objects. But what about when the image has no targe…

2021

End-to-End Piece-Wise Unwarping of Document Images

ICCV 2021poster

Document unwarping attempts to undo the physical deformation of the paper and recover a 'flatbed' scanned document-image for downstream tasks such as OCR. Current state-of-the-art relies on global unwarping of the document which is not robust to local deformation changes. Moreover, a global unwarpin…

Cited by 36PDFScholar
2021

Localization in the Crowd with Topological Constraints

AAAI 2021technical

We address the problem of crowd localization, i.e., the prediction of dots corresponding to people in a crowded scene. Due to various challenges, a localization method is prone to spatial semantic errors, i.e., predicting multiple dots within a same person or collapsing multiple dots in a cluttered…

2021

Modeling Deep Learning Based Privacy Attacks on Physical Mail

AAAI 2021technical

Mail privacy protection aims to prevent unauthorized access to hidden content within an envelope since normal paper envelopes are not as safe as we think. In this paper, for the first time, we show that with a well designed deep learning model, the hidden content may be largely recovered without ope…

2021

Multi-Class Cell Detection Using Spatial Context Representation

ICCV 2021poster

In digital pathology, both detection and classification of cells are important for automatic diagnostic and prognostic tasks. Classifying cells into subtypes, such as tumor cells, lymphocytes or stromal cells is particularly challenging. Existing methods focus on morphological appearance of individu…

Cited by 42PDFcodeScholar
2021

Topology-Aware Segmentation Using Discrete Morse Theory

ICLR 2021spotlight

In the segmentation of fine-scale structures from natural and biomedical images, per-pixel accuracy is not the only metric of concern. Topological correctness, such as vessel connectivity and membrane closure, is crucial for downstream analysis tasks. In this paper, we propose a new approach to trai…

Cited by 112SourcePDFScholar
2021

Variational Feature Disentangling for Fine-Grained Few-Shot Classification

ICCV 2021poster

Data augmentation is an intuitive step towards solving the problem of few-shot classification. However, ensuring both discriminability and diversity in the augmented samples is challenging. To address this, we propose a feature disentanglement framework that allows us to augment features with random…

Cited by 76PDFcodeScholar
2020

Distribution Matching for Crowd Counting

NeurIPS 2020spotlight

In crowd counting, each training image contains multiple people, where each person is annotated by a dot. Existing crowd counting methods need to use a Gaussian to smooth each annotated dot or to estimate the likelihood of every pixel given the annotated point. In this paper, we show that imposing G…

2020

Learning Visual Emotion Representations From Web Data

CVPR 2020poster

We present a scalable approach for learning powerful visual features for emotion recognition. A critical bottleneck in emotion recognition is the lack of large scale datasets that can be used for learning visual emotion features. To this end, we curate a webly derived large scale dataset, StockEmoti…

Cited by 50PDFScholar
2020

Predicting Goal-Directed Human Attention Using Inverse Reinforcement Learning

CVPR 2020oral

Human gaze behavior prediction is important for behavioral vision and for computer vision applications. Most models mainly focus on predicting free-viewing behavior using saliency maps, but do not generalize to goal-directed behavior, such as when a person searches for a visual target object. We pro…

Cited by 136PDFcodeScholar
2019

DewarpNet: Single-Image Document Unwarping With Stacked 3D and 2D Regression Networks

ICCV 2019poster

Capturing document images with hand-held devices in unstructured environments is a common practice nowadays. However, "casual" photos of documents are usually unsuitable for automatic information extraction, mainly due to physical distortion of the document paper, as well as various camera positions…

Cited by 92PDFScholar
2019

Label super-resolution networks

ICLR 2019poster

We present a deep learning-based method for super-resolving coarse (low-resolution) labels assigned to groups of image pixels into pixel-level (high-resolution) labels, given the joint distribution between those low- and high-resolution labels. This method involves a novel loss function that minimiz…

Cited by 37SourcePDFScholar
2019

Robust Histopathology Image Analysis: To Label or to Synthesize?

CVPR 2019oral

Detection, segmentation and classification of nuclei are fundamental analysis operations in digital pathology. Existing state-of-the-art approaches demand extensive amount of supervised training data from pathologists and may still perform poorly in images from unseen tissue types. We propose an uns…

Cited by 160PDFScholar
2018

A+D Net: Training a Shadow Detector with Adversarial Shadow Attenuation

ECCV 2018poster

We propose a novel GAN-based framework for detecting shadows in images, in which a shadow detection network (D-Net) is trained together with a shadow attenuation network (A-Net) that generates adversarial training examples. The A-Net modifies the original training images constrained by a simplified…

Cited by 142SourcePDFScholar
2018

Deforming Autoencoders: Unsupervised Disentangling of Shape and Appearance

ECCV 2018poster

In this work we introduce the Deforming Autoencoder, a generative model for images that disentangles shape from appearance in a latent representation space that is learned in a fully unsupervised manner. As in the deformable template paradigm, shape is represented as a diffeomorphism between a canon…

Cited by 249SourcePDFScholar
2018

Good View Hunting: Learning Photo Composition From Dense View Pairs

CVPR 2018poster

Finding views with good photo composition is a challenging task for machine learning methods. A key difficulty is the lack of well annotated large scale datasets. Most existing datasets only provide a limited number of annotations for good views, while ignoring the comparative nature of view select…

Cited by 110SourcePDFScholar
2018

Sequence-to-Segment Networks for Segment Detection

NeurIPS 2018poster

Detecting segments of interest from an input sequence is a challenging problem which often requires not only good knowledge of individual target segments, but also contextual understanding of the entire input sequence and the relationships between the target segments. To address this problem, we pr…

Cited by 20SourcePDFScholar
2017

ConvNets with Smooth Adaptive Activation Functions for Regression

AISTATS 2017poster

Within Neural Networks (NN), the parameters of Adaptive Activation Functions (AAF) control the shapes of activation functions. These parameters are trained along with other parameters in the NN. AAFs have improved performance of Convolutional Neural Networks (CNN) in multiple classification tasks. I…

Cited by 57SourcePDFScholar
2017

Neural Face Editing With Intrinsic Image Disentangling

CVPR 2017oral

Traditional face editing methods often require a number of sophisticated and task specific algorithms to be applied one after the other --- a process that is tedious, fragile, and computationally intensive. In this paper, we propose an end-to-end generative adversarial network that infers a face-spe…

Cited by 340PDFcodeScholar
2017

Shadow Detection With Conditional Generative Adversarial Networks

ICCV 2017oral

We introduce scGAN, a novel extension of conditional Generative Adversarial Networks (GAN) tailored for the challenging problem of shadow detection in images. Previous methods for shadow detection focus on learning the local appearance of shadow regions, while using limited local context reasoning i…

Cited by 245PDFScholar
2016

Learned Region Sparsity and Diversity Also Predicts Visual Attention

NeurIPS 2016poster

Learned region sparsity has achieved state-of-the-art performance in classification tasks by exploiting and integrating a sparse set of local information into global decisions. The underlying mechanism resembles how people sample information from an image with their eye movements when making similar…

Cited by 22SourcePDFScholar
2016

Patch-Based Convolutional Neural Network for Whole Slide Tissue Image Classification

CVPR 2016spotlight

Convolutional Neural Networks (CNN) are state-of-the-art models for many image classification tasks. However, to recognize cancer subtypes automatically, training a CNN on gigapixel resolution Whole Slide Tissue Images (WSI) is currently computationally impossible. The differentiation of cancer subt…

Cited by 1032PDFScholar
2015

Efficient Video Segmentation Using Parametric Graph Partitioning

ICCV 2015poster

Video segmentation is the task of grouping similar pixels in the spatio-temporal domain, and has become an important preprocessing step for subsequent video analysis. Most video segmentation and supervoxel methods output a hierarchy of segmentations, but while this provides useful multiscale informa…

Cited by 35PDFScholar