← Search

Noel E. O'Connor

7 accepted papers

2025

Harnessing Frozen Unimodal Encoders for Flexible Multimodal Alignment

CVPR 2025poster

Recent contrastive multimodal vision-language models like CLIP have demonstrated robust open-world semantic understanding, becoming the standard image backbones for vision-language applications. However, recent findings suggest high semantic similarity between well-trained unimodal encoders, which r…

2024

Do Vision and Language Encoders Represent the World Similarly?

CVPR 2024poster

Aligned text-image encoders such as CLIP have become the de-facto model for vision-language tasks. Furthermore modality-specific encoders achieve impressive performances in their respective domains. This raises a central question: does an alignment exist between uni-modal vision and language encoder…

2024

Identifying Expert Behavior in Offline Training Datasets Improves Behavioral Cloning of Robotic Manipulation Policies

RA-L 2024

This letter presents our solution for the Real Robot Challenge III <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> , aiming to address dexterous robotic manipulation tasks through learning from offline data. In this competition, participants wer

Cited by 13SourcecodeScholar
2021

Multi-Objective Interpolation Training for Robustness To Label Noise

CVPR 2021poster

Deep neural networks trained with standard cross-entropy loss memorize noisy labels, which degrades their performance. Most research to mitigate this memorization proposes new robust classification loss functions. Conversely, we propose a Multi-Objective Interpolation Training (MOIT) approach that j…

Cited by 159PDFcodeScholar
2021

Unsupervised Contrastive Learning of Sound Event Representations

ICASSP 2021accepted

Self-supervised representation learning can mitigate the limitations in recognition tasks with few manually labeled data but abundant unlabeled data—a common scenario in sound event research. In this work, we explore unsupervised contrastive learning as a way to learn sound event representations. To…

Cited by 0SourceScholar
2018

People, Penguins and Petri Dishes: Adapting Object Counting Models to New Visual Domains and Object Types Without Forgetting

CVPR 2018poster

In this paper we propose a technique to adapt a convolutional neural network (CNN) based object counter to additional visual domains and object types while still preserving the original counting function. Domain-specific normalisation and scaling operators are trained to allow the model to adjust t…

2016

Shallow and Deep Convolutional Networks for Saliency Prediction

CVPR 2016poster

The prediction of salient areas in images has been traditionally addressed with hand-crafted features based on neuroscience principles. This paper, however, addresses the problem with a completely data-driven approach by training a convolutional neural network (convnet). The learning process is form…

Cited by 587PDFcodeScholar