← Search

Serena Yeung

24 accepted papers

2026

SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models

ICML 2026poster

Large Multimodal Models (LMMs) have achieved remarkable progress across various capabilities; however, complex video reasoning in the scientific domain remains a significant and challenging frontier. Current video benchmarks predominantly target general scenarios where perception/recognition is heav…

Cited by 0SourceScholar
2024

Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data

ICLR 2024poster

Building cross-modal applications is challenging due to limited paired multi-modal data. Recent works have shown that leveraging a pre-trained multi-modal contrastive representation space enables cross-modal tasks to be learned from uni-modal data. This is based on the assumption that contrastive op…

2023

Beyond Positive Scaling: How Negation Impacts Scaling Trends of Language Models

ACL 2023findings

Language models have been shown to exhibit positive scaling, where performance improves as models are scaled up in terms of size, compute, or data. In this work, we introduce NeQA, a dataset consisting of questions with negation in which language models do not exhibit straightforward positive scalin…

2023

DataPerf: Benchmarks for Data-Centric AI Development

NeurIPS 2023poster

Machine learning research has long focused on models rather than datasets, and prominent datasets are used for common ML tasks without regard to the breadth, difficulty, and faithfulness of the underlying problems. Neglecting the fundamental importance of data has given rise to inaccuracy, bias, and…

2023

Diagnosing and Rectifying Vision Models using Language

ICLR 2023poster

Recent multi-modal contrastive learning models have demonstrated the ability to learn an embedding space suitable for building strong vision classifiers, by leveraging the rich information in large-scale image-caption datasets. Our work highlights a distinct advantage of this multi-modal embedding s…

2023

INSPECT: A Multimodal Dataset for Pulmonary Embolism Diagnosis and Prognosis

NeurIPS 2023poster

Synthesizing information from various data sources plays a crucial role in the practice of modern medicine. Current applications of artificial intelligence in medicine often focus on single-modality data due to a lack of publicly available, multimodal medical datasets. To address this limitation, we…

Cited by 11SourcePDFScholar
2023

LOVM: Language-Only Vision Model Selection

NeurIPS 2023poster

Pre-trained multi-modal vision-language models (VLMs) are becoming increasingly popular due to their exceptional performance on downstream vision applications, particularly in the few- and zero-shot settings. However, selecting the best-performing VLM for some downstream applications is non-trivial,…

2023

NeMo: Learning 3D Neural Motion Fields From Multiple Video Instances of the Same Action

CVPR 2023highlight

The task of reconstructing 3D human motion has wide-ranging applications. The gold standard Motion capture (MoCap) systems are accurate but inaccessible to the general public due to their cost, hardware, and space constraints. In contrast, monocular human mesh recovery (HMR) methods are much more ac…

Cited by 9SourcePDFScholar
2022

Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

NeurIPS 2022accept

We present modality gap, an intriguing geometric phenomenon of the representation space of multi-modal models. Specifically, we show that different data modalities (e.g. images and text) are embedded at arm's length in their shared representation in multi-modal models such as CLIP. Our systematic an…

2021

Capturing implicit hierarchical structure in 3D biomedical images with self-supervised hyperbolic representations

NeurIPS 2021poster

We consider the task of representation learning for unsupervised segmentation of 3D voxel-grid biomedical images. We show that models that capture implicit hierarchical relationships between subvolumes are better suited for this task. To that end, we consider encoder-decoder architectures with a hyp…

Cited by 35SourcePDFScholar
2021

DARCNN: Domain Adaptive Region-Based Convolutional Neural Network for Unsupervised Instance Segmentation in Biomedical Images

CVPR 2021poster

In the biomedical domain, there is an abundance of dense, complex data where objects of interest may be challenging to detect or constrained by limits of human knowledge. Labelled domain specific datasets for supervised tasks are often expensive to obtain, and furthermore discovery of novel distinct…

Cited by 43PDFcodeScholar
2021

GLoRIA: A Multimodal Global-Local Representation Learning Framework for Label-Efficient Medical Image Recognition

ICCV 2021poster

In recent years, the growing number of medical imaging studies is placing an ever-increasing burden on radiologists. Deep learning provides a promising solution for automatic medical image analysis and clinical decision support. However, large-scale manually labeled datasets required for training de…

Cited by 404PDFcodeScholar
2021

Personalized Federated Learning with First Order Model Optimization

ICLR 2021poster

While federated learning traditionally aims to train a single global model across decentralized local datasets, one model may not always be ideal for all participating clients. Here we propose an alternative, where each client only federates with other relevant clients to obtain a stronger model per…

2021

Unsupervised Discovery of the Long-Tail in Instance Segmentation Using Hierarchical Self-Supervision

CVPR 2021poster

Instance segmentation is an active topic in computer vision that is usually solved by using supervised learning approaches over very large datasets composed of object level masks. Obtaining such a dataset for any new domain can be very expensive and time-consuming. In addition, models trained on cer…

Cited by 47PDFScholar
2018

Dynamic Task Prioritization for Multitask Learning

ECCV 2018poster

We propose dynamic task prioritization for multitask learning. This allows a model to dynamically prioritize difficult tasks during training, where difficulty is inversely proportional to performance, and where difficulty changes over time. In contrast to curriculum learning, where easy tasks are pr…

Cited by 470SourcePDFScholar
2018

Neural Graph Matching Networks for Fewshot 3D Action Recognition

ECCV 2018poster

We propose Neural Graph Matching (NGM) Networks, a novel framework that can learn to recognize a previous unseen 3D action class with only a few examples. We achieve this by leveraging the inherent structure of 3D data through a graphical representation. This allows us to modularize our model and le…

Cited by 132SourcePDFScholar
2018

Temporal Modular Networks for Retrieving Complex Compositional Activities in Videos

ECCV 2018poster

A major challenge in computer vision is scaling activity understanding to the long tail of complex activities without requiring collecting large quantities of data for new actions. The task of video retrieval using natural language descriptions seeks to address this through rich, unconstrained super…

Cited by 94SourcePDFScholar
2017

Jointly Learning Energy Expenditures and Activities Using Egocentric Multimodal Signals

CVPR 2017poster

Physiological signals such as heart rate can provide valuable information about an individual's state and activity. However, existing work on computer vision has not yet explored leveraging these signals to enhance egocentric video understanding. In this work, we propose a model for reasoning on mul…

Cited by 88PDFScholar
2017

Learning to Learn From Noisy Web Videos

CVPR 2017poster

Understanding the simultaneously very diverse and intricately fine-grained set of possible human actions is a critical open problem in computer vision. Manually labeling training videos is feasible for some action classes but doesn't scale to the full long-tailed distribution of actions. A promising…

Cited by 36PDFScholar
2016

End-To-End Learning of Action Detection From Frame Glimpses in Videos

CVPR 2016poster

In this work we introduce a fully end-to-end approach for action detection in videos that learns to directly predict the temporal bounds of actions. Our intuition is that the process of detecting actions is naturally one of observation and refinement: observing moments in video, and refining hypothe…

Cited by 763PDFScholar