← Search

Thomas Seidl

8 accepted papers

2026

CHB: A Diagnostic Toolkit for Hardness-Aware Clustering Evaluation

ICML 2026poster

Clustering is commonly compared through leaderboards that collapse performance into a single aggregate ranking. Such summaries obscure why methods succeed, which data properties align with failure, and how conclusions shift under representation changes and realistic tuning constraints. We present CH…

Cited by 0SourceScholar
2026

Human Uncertainty-Aware Data Selection and Automatic Labeling in Visual Question Answering

ICLR 2026poster

Large vision-language models (VLMs) achieve strong performance in Visual Question Answering but still rely heavily on supervised fine-tuning (SFT) with massive labeled datasets, which is costly due to human annotations. Crucially, real-world datasets often exhibit *human uncertainty* (**HU**) — var…

Cited by 0SourceScholar
2025

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos

CVPR 2025poster

Large language models (LLMs) excel at retrieving information from lengthy text, but their vision-language counterparts (VLMs) face difficulties with hour-long videos, especially for temporal grounding. Specifically, these VLMs are constrained by frame limitations, often losing essential temporal det…

2024

Autoregressive Policy Optimization for Constrained Allocation Tasks

NeurIPS 2024poster

Allocation tasks represent a class of problems where a limited amount of resources must be allocated to a set of entities at each time step. Prominent examples of this task include portfolio optimization or distributing computational workloads across servers. Allocation tasks are typically bound by…

2024

RGNet: A Unified Clip Retrieval and Grounding Network for Long Videos

ECCV 2024poster

"Locating specific moments within long videos (20–120 minutes) presents a significant challenge, akin to finding a needle in a haystack. Adapting existing short video (5–30 seconds) grounding methods to this problem yields poor performance. Since most real-life videos, such as those on YouTube and A…

2023

InstanceFormer: An Online Video Instance Segmentation Framework

AAAI 2023technical

Recent transformer-based offline video instance segmentation (VIS) approaches achieve encouraging results and significantly outperform online approaches. However, their reliance on the whole video and the immense computational complexity caused by full Spatio-temporal attention limit them in real-li…

2021

Argument Mining Driven Analysis of Peer-Reviews

AAAI 2021technical

Peer reviewing is a central process in modern research and essential for ensuring high quality and reliability of published work. At the same time, it is a time-consuming process and increasing interest in emerging fields often results in a high review workload, especially for senior researchers in…