← Search

Thorsten Bagdonat

2 accepted papers

2025

FOCUS: Internal MLLM Representations for Efficient Fine-Grained Visual Question Answering

NeurIPS 2025poster

While Multimodal Large Language Models (MLLMs) offer strong perception and reasoning capabilities for image-text input, Visual Question Answering (VQA) focusing on small image details still remains a challenge. Although visual cropping techniques seem promising, recent approaches have several limita…

Cited by 0SourceScholar
2024

Distributed Semantic Segmentation with Efficient Joint Source and Task Decoding

ECCV 2024poster

"Distributed computing in the context of deep neural networks (DNNs) implies the execution of one part of the network on edge devices and the other part typically on a large-scale cloud platform. Conventional methods propose to employ a serial concatenation of a learned image and source encoder, the…