← Search

Tom Gedeon

16 accepted papers

2026

Bipartite Mode Matching for Vision Training Set Search from a Hierarchical Data Server

AAAI 2026technical

We explore a situation in which the target domain is accessible, but real-time data annotation is not feasible. Instead, we would like to construct an alternative training set from a large-scale data server so that a competitive model can be obtained. For this problem, because the target domain usua

Cited by 0SourcePDFScholar
2025

Learnable Expansion of Graph Operators for Multi-Modal Feature Fusion

ICLR 2025poster

In computer vision tasks, features often come from diverse representations, domains (e.g., indoor and outdoor), and modalities (e.g., text, images, and videos). Effectively fusing these features is essential for robust performance, especially with the availability of powerful pre-trained models like…

Cited by 1SourcePDFScholar
2025

Ranked from Within: Ranking Large Multimodal Models Without Labels

ICML 2025poster

Can the relative performance of a pre-trained large multimodal model (LMM) be predicted without access to labels? As LMMs proliferate, it becomes increasingly important to develop efficient ways to choose between them when faced with new data or tasks. The usual approach does the equivalent of givin…

Cited by 0SourcePDFScholar
2025

Unsupervised Search for Ethnic Minorities' Medical Segmentation Training Set

ICASSP 2025accepted

This paper investigates the critical issue of dataset bias in medical imaging, with a particular emphasis on racial disparities caused by uneven population distribution in dataset collection. Our analysis reveals that medical segmentation datasets are significantly biased, primarily influenced by th…

Cited by 0SourceScholar
2024

Advancing Video Anomaly Detection: A Concise Review and a New Dataset

NeurIPS 2024poster

Video Anomaly Detection (VAD) finds widespread applications in security surveillance, traffic monitoring, industrial monitoring, and healthcare. Despite extensive research efforts, there remains a lack of concise reviews that provide insightful guidance for researchers. Such reviews would serve as q…

Cited by 14SourcePDFScholar
2024

An Empirical Study Into What Matters for Calibrating Vision-Language Models

ICML 2024poster

Vision-Language Models (VLMs) have emerged as the dominant approach for zero-shot recognition, adept at handling diverse scenarios and significant distribution changes. However, their deployment in risk-sensitive areas requires a deeper understanding of their uncertainty estimation capabilities, a r…

Cited by 8SourcePDFScholar
2024

Visual Prompting in LLMs for Enhancing Emotion Recognition

EMNLP 2024main

Vision Large Language Models (VLLMs) are transforming the intersection of computer vision and natural language processing; however, the potential of using visual prompts for emotion recognition in these models remains largely unexplored and untapped. Traditional methods in VLLMs struggle with spatia…

2023

A Bag-of-Prototypes Representation for Dataset-Level Applications

CVPR 2023poster

This work investigates dataset vectorization for two dataset-level tasks: assessing training set suitability and test set difficulty. The former measures how suitable a training set is for a target domain, while the latter studies how challenging a test set is for a learned model. Central of the two…

Cited by 12SourcePDFScholar
2023

A Closer Look at the Robustness of Contrastive Language-Image Pre-Training (CLIP)

NeurIPS 2023poster

Contrastive Language-Image Pre-training (CLIP) models have demonstrated remarkable generalization capabilities across multiple challenging distribution shifts. However, there is still much to be explored in terms of their robustness to the variations of specific visual factors. In real-world applica…

Cited by 40SourcePDFScholar
2022

How to Synthesize a Large-Scale and Trainable Micro-Expression Dataset?

ECCV 2022poster

"This paper does not contain technical novelty but introduces our key discoveries in a data generation protocol, a database and insights. We aim to address the lack of large-scale datasets in micro-expression (MiE) recognition due to the prohibitive cost of data collection, which renders large-scale…

2021

Invertible Denoising Network: A Light Solution for Real Noise Removal

CVPR 2021poster

Invertible networks have various benefits for image denoising since they are lightweight, information-lossless, and memory-saving during back-propagation. However, applying invertible models to remove noise is challenging because the input is noisy, and the reversed output is clean, following two di…

Cited by 200PDFcodeScholar
2020

Simulating Content Consistent Vehicle Datasets with Attribute Descent

ECCV 2020poster

This paper uses a graphic engine to simulate a large amount of training data with free annotations. Between synthetic and real data, there is a two-level domain gap, i.e., content level and appearance level. While the latter has been widely studied, we focus on reducing the content gap in attributes…