← Search

Richard McCreadie

4 accepted papers

2024

LaCViT: A Label-Aware Contrastive Fine-Tuning Framework for Vision Transformers

ICASSP 2024accepted

Vision Transformers (ViTs) have emerged as popular models in computer vision, demonstrating state-of-the-art performance across various tasks. This success typically follows a two-stage strategy involving pre-training on large-scale datasets using self-supervised signals, such as masked random patch…

Cited by 0SourceScholar
2024

Multiway-Adapter: Adapting Multimodal Large Language Models for Scalable Image-Text Retrieval

ICASSP 2024accepted

As Multimodal Large Language Models (MLLMs) grow in size, adapting them to specialized tasks becomes increasingly challenging due to high computational and memory demands. Indeed, traditional fine-tuning methods are costly, due to the need for extensive, task-specific training. While efficient adapt…

Cited by 0SourceScholar
2024

RoboLLM: Robotic Vision Tasks Grounded on Multimodal Large Language Models

ICRA 2024poster

Robotic vision applications often necessitate a wide range of visual perception tasks, such as object detection, segmentation, and identification. While there have been substantial advances in these individual tasks, integrating specialized models into a unified vision pipeline presents significant…

Cited by 21SourcecodeScholar
2024

Understanding and Mitigating Human-Labelling Errors in Supervised Contrastive Learning

ECCV 2024poster

"Human-annotated vision datasets inevitably contain a fraction of human-mislabelled examples. While the detrimental effects of such mislabelling on supervised learning are well-researched, their influence on Supervised Contrastive Learning (SCL) remains largely unexplored. In this paper, we show tha…

Cited by 3SourcePDFScholar