← Search

Sara Atito

11 accepted papers

2026

Building Robust Vision Encoders for Cross-Dataset Evaluation in Immunofluorescent Microscopy

CVPR 2026

Immunofluorescence (IF) images reveal detailed information about structures and functions at the subcellular level. However, unlike RGB images, IF datasets pose challenges for deep learning models due to their inconsistencies in channel count and configuration, stemming from varying staining protoco

Cited by 0SourcecodeScholar
2026

CoLA: Cross-Modal Low-rank Adaptation for Multimodal Downstream Tasks

ICML 2026poster

Foundation models have revolutionized AI, but adapting them efficiently for multimodal tasks, particularly in dual-stream architectures composed of unimodal encoders, such as DINO and BERT, remains a significant challenge. Parameter-Efficient Fine-Tuning (PEFT) methods like Low-Rank Adaptation (LoRA…

Cited by 0SourceScholar
2026

Object-Centric Refinement for Enhanced Zero-Shot Segmentation

ICLR 2026poster

Zero-shot semantic segmentation aims to recognize, pixel-wise, unseen categories without annotated masks, typically by leveraging vision-language models such as CLIP. However, the patch representations obtained by the CLIP's vision encoder lack object-centric structure, making it difficult to locali…

Cited by 0SourceScholar
2025

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning

NeurIPS 2025poster

Despite significant advances in inference-time search for vision–language models (VLMs), existing approaches remain both computationally expensive and prone to unpenalized, low-confidence generations which often lead to persistent hallucinations. We introduce \textbf{Value-guided Inference with Marg…

Cited by 0SourcecodeScholar
2025

Enhanced Weakly Supervised Few-shot Classification & Segmentation

ICASSP 2025accepted

The emergence of vision-language foundation models has enabled the integration of textual information into vision-based applications. However, in few-shot classification and segmentation (FS-CS), this potential remains underutilised. Commonly, self-supervised vision models have been employed, partic…

Cited by 0SourceScholar
2025

One Model for ALL: Low-Level Task Interaction Is a Key to Task-Agnostic Image Fusion

CVPR 2025poster

Advanced image fusion methods mostly prioritise high-level missions, where task interaction struggles with semantic gaps, requiring complex bridging mechanisms. In contrast, we propose to leverage low-level vision tasks from digital photography fusion, allowing for effective feature interaction thro…

2025

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes

ICLR 2025poster

Self-supervised pre-trained audio networks have seen widespread adoption in real-world systems, particularly in multi-modal large language models. These networks are often employed in a frozen state, under the assumption that the self-supervised pre-training has sufficiently equipped them to handle…

2025

Text Augmented Correlation Transformer For Few-shot Classification & Segmentation

CVPR 2025poster

Foundation models like CLIP and ALIGN have transformed few-shot and zero-shot vision applications by fusing visual and textual data, yet the integrative few-shot classification and segmentation (FS-CS) task primarily leverages visual cues, overlooking the potential of textual support. In FS-CS scena…

Cited by 0SourcePDFScholar
2024

C2C: Component-to-Composition Learning for Zero-Shot Compositional Action Recognition

ECCV 2024oral

"Compositional actions consist of dynamic (verbs) and static (objects) concepts. Humans can easily recognize unseen compositions using the learned concepts. For machines, solving such a problem requires a model to recognize unseen actions composed of previously observed verbs and objects, thus requi…

2024

Improved Image Captioning Via Knowledge Graph-Augmented Models

ICASSP 2024accepted

Multimodal foundation models, pre-trained on large-scale data, effectively capture vast amounts of factual and commonsense knowledge. However, these models store all their knowledge within their parameters, requiring increasingly larger models and training data to capture more knowledge. To address…

Cited by 0SourceScholar