← Search

Daniela Massiceti

12 accepted papers

2025

Investigating Dictionary Expansion for Video-based Sign Language Dictionaries

EMNLP 2025

Like most languages, sign languages evolve over time. It is important that sign language dictionaries’ vocabularies are updated over time to reflect these changes, such as by adding new signs. However, most dictionary retrieval methods based upon machine learning models only work with fixed vocabula

Cited by 0SourcePDFScholar
2024

Distilling Knowledge from Text-to-Image Generative Models Improves Visio-Linguistic Reasoning in CLIP

EMNLP 2024main

Image-text contrastive models like CLIP have wide applications in zero-shot classification, image-text retrieval, and transfer learning. However, they often struggle on compositional visio-linguistic tasks (e.g., attribute-binding or object-relationships) where their performance is no better than ra…

Cited by 1SourcePDFScholar
2024

Explaining CLIP's Performance Disparities on Data from Blind/Low Vision Users

CVPR 2024poster

Large multi-modal models (LMMs) hold the potential to usher in a new era of automated visual assistance for people who are blind or low vision (BLV). Yet these models have not been systematically evaluated on data captured by BLV users. We address this by empirically assessing CLIP a widely-used LMM…

Cited by 7SourcePDFScholar
2024

Strong Baselines for Parameter-Efficient Few-Shot Fine-Tuning

AAAI 2024technical

Few-shot classification (FSC) entails learning novel classes given only a few examples per class after a pre-training (or meta-training) phase on a set of base classes. Recent works have shown that simply fine-tuning a pre-trained Vision Transformer (ViT) on new test classes is a strong approach for…

Cited by 31SourcePDFScholar
2024

Understanding Information Storage and Transfer in Multi-Modal Large Language Models

NeurIPS 2024poster

Understanding the mechanisms of information storage and transfer in Transformer-based models is important for driving model understanding progress. Recent work has studied these mechanisms for Large Language Models (LLMs), revealing insights on how information is stored in a model's parameters and h…

Cited by 12SourcePDFScholar
2023

Hard-Meta-Dataset++: Towards Understanding Few-Shot Performance on Difficult Tasks

ICLR 2023poster

Few-shot classification is the ability to adapt to any new classification task from only a few training examples. The performance of current top-performing few-shot classifiers varies widely across different tasks where they often fail on a subset of `difficult' tasks. This phenomenon has real-world…

Cited by 6SourcePDFScholar
2023

NP-SemiSeg: When Neural Processes meet Semi-Supervised Semantic Segmentation

ICML 2023poster

Semi-supervised semantic segmentation involves assigning pixel-wise labels to unlabeled images at training time. This is useful in a wide range of real-world applications where collecting pixel-wise labels is not feasible in time or cost. Current approaches to semi-supervised semantic segmentation w…

2022

NP-Match: When Neural Processes meet Semi-Supervised Learning

ICML 2022spotlight

Semi-supervised learning (SSL) has been widely explored in recent years, and it is an effective way of leveraging unlabeled data to reduce the reliance on labeled data. In this work, we adjust neural processes (NPs) to the semi-supervised image classification task, resulting in a new method named NP…

2021

Memory Efficient Meta-Learning with Large Images

NeurIPS 2021poster

Meta learning approaches to few-shot classification are computationally efficient at test time, requiring just a few optimization steps or single forward pass to learn a new task, but they remain highly memory-intensive to train. This limitation arises because a task's entire support set, which can…

Cited by 26SourcePDFScholar
2021

ORBIT: A Real-World Few-Shot Dataset for Teachable Object Recognition

ICCV 2021poster

Object recognition has made great advances in the last decade, but predominately still relies on many high-quality training examples per object category. In contrast, learning new objects from only a few examples could enable many impactful applications from robotics to user personalization. Most fe…

Cited by 57PDFcodeScholar
2018

FlipDial: A Generative Model for Two-Way Visual Dialogue

CVPR 2018poster

We present FlipDial, a generative model for Visual Dialogue that simultaneously plays the role of both participants in a visually-grounded dialogue. Given context in the form of an image and an associated caption summarising the contents of the image, FlipDial learns both to answer questions and put…

Cited by 48SourcePDFScholar
2017

Random forests versus Neural Networks — What's best for camera localization?

ICRA 2017poster

This work addresses the task of camera localization in a known 3D scene given a single input RGB image. State-of-the-art approaches accomplish this in two steps: firstly, regressing for every pixel in the image its 3D scene coordinate and subsequently, using these coordinates to estimate the final 6…

Cited by 90SourceScholar