← Search

Oisin Mac Aodha

47 accepted papers

2026

BiMotion: B-spline Motion for Text-guided Dynamic 3D Character Generation

CVPR 2026

Text-guided dynamic 3D character generation has advanced rapidly, yet producing high-quality motion that faithfully reflects rich textual descriptions remains challenging. Existing methods tend to generate limited sub-actions or incoherent motion due to fixed-length temporal inputs and discrete fram

Cited by 0SourcecodeScholar
2026

MotionPhysics: Learnable Motion Distillation for Text-Guided Simulation

AAAI 2026technical

Accurately simulating existing 3D objects and a wide variety of materials often demands expert knowledge and time-consuming physical parameter tuning to achieve the desired dynamic behavior. We introduce MotionPhysics, an end‑to‑end differentiable framework that infers plausible physical parameters

Cited by 0SourcePDFScholar
2026

The Temporal Trap: Entanglement in Pre-Trained Visual Representations for Visuomotor Policy Learning

ICRA 2026poster

The integration of pre-trained visual representations (PVRs) has significantly advanced visuomotor policy learning. However, effectively leveraging these models remains a challenge. We identify temporal entanglement as a critical, inherent issue when using these time-invariant models in sequential d…

2025

CleverBirds: A Multiple-Choice Benchmark for Fine-grained Human Knowledge Tracing

NeurIPS 2025poster

Mastering fine-grained visual recognition, essential in many expert domains, can require that specialists undergo years of dedicated training. Modeling the progression of such expertize in humans remains challenging, and accurately inferring a human learner’s knowledge state is a key step toward und…

Cited by 0SourceScholar
2025

CrossSDF: 3D Reconstruction of Thin Structures From Cross-Sections

CVPR 2025poster

Reconstructing complex structures from planar cross-sections is a challenging problem, with wide-reaching applications in medical imaging, manufacturing, and topography. Out-of-the-box point cloud reconstruction methods can often fail due to the data sparsity between slicing planes, while current be…

Cited by 0SourcePDFScholar
2025

DepthCues: Evaluating Monocular Depth Perception in Large Vision Models

CVPR 2025poster

Large-scale pre-trained vision models are becoming increasingly prevalent, offering expressive and generalizable visual representations that benefit various downstream tasks. Recent studies on the emergent properties of these models have revealed their high-level geometric understanding, in particul…

Cited by 3SourcePDFScholar
2025

Enhancing Tactile-based Reinforcement Learning for Robotic Control

NeurIPS 2025poster

Achieving safe, reliable real-world robotic manipulation requires agents to evolve beyond vision and incorporate tactile sensing to overcome sensory deficits and reliance on idealised state information. Despite its potential, the efficacy of tactile sensing in reinforcement learning (RL) remains inc…

Cited by 0SourcecodeScholar
2025

Feedforward Few-shot Species Range Estimation

ICML 2025poster

Knowing where a particular species can or cannot be found on Earth is crucial for ecological research and conservation efforts. By mapping the spatial ranges of all species, we would obtain deeper insights into how global biodiversity is affected by climate change and habitat loss. However, accurat…

Cited by 0SourcePDFScholar
2025

Jamais Vu: Exposing the Generalization Gap in Supervised Semantic Correspondence

NeurIPS 2025poster

Semantic correspondence (SC) aims to establish semantically meaningful matches across different instances of an object category. We illustrate how recent supervised SC methods remain limited in their ability to generalize beyond sparsely annotated training keypoints, effectively acting as keypoint d…

Cited by 0SourceScholar
2025

MVSAnywhere: Zero-Shot Multi-View Stereo

CVPR 2025poster

Computing accurate depth from multiple views is a fundamental and longstanding challenge in computer vision.However, most existing approaches do not generalize well across different domains and scene types (e.g. indoor vs outdoor). Training a general-purpose multi-view stereo model is challenging an…

2025

Representational Similarity via Interpretable Visual Concepts

ICLR 2025poster

How do two deep neural networks differ in how they arrive at a decision? Measuring the similarity of deep networks has been a long-standing open question. Most existing methods provide a single number to measure the similarity of two networks at a given layer, but give no insight into what makes th…

2025

The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements

NeurIPS 2025poster

Rapidly improving large language models (LLMs) have the potential to assist in scientific progress. One critical skill in this endeavor is the ability to faithfully reproduce existing work. To evaluate the capability of AI agents to reproduce complex code in an active research area, we introduce the…

Cited by 0SourcecodeScholar
2025

WildSAT: Learning Satellite Image Representations from Wildlife Observations

ICCV 2025poster

Species distributions encode valuable ecological and environmental information, yet their potential for guiding representation learning in remote sensing remains underexplored. We introduce WildSAT, which pairs satellite images with millions of geo-tagged wildlife observations readily-available on c…

2024

AirPlanes: Accurate Plane Estimation via 3D-Consistent Embeddings

CVPR 2024poster

Extracting planes from a 3D scene is useful for downstream tasks in robotics and augmented reality. In this paper we tackle the problem of estimating the planar surfaces in a scene from posed images. Our first finding is that a surprisingly competitive baseline results from combining popular cluster…

Cited by 1SourcePDFScholar
2024

Click to Grasp: Zero-Shot Precise Manipulation via Visual Diffusion Descriptors

IROS 2024poster

Precise manipulation that is generalizable across scenes and objects remains a persistent challenge in robotics. Current approaches for this task heavily depend on having a significant number of training instances to handle objects with pronounced visual and/or geometric part ambiguities. Our work e…

Cited by 4SourcecodeScholar
2024

Combining Observational Data and Language for Species Range Estimation

NeurIPS 2024poster

Species range maps (SRMs) are essential tools for research and policy-making in ecology, conservation, and environmental management. However, traditional SRMs rely on the availability of environmental covariates and high-quality observational data, both of which can be challenging to obtain due to g…

2024

From Coarse to Fine-Grained Open-Set Recognition

CVPR 2024poster

Open-set recognition (OSR) methods aim to identify whether or not a test example belongs to a category ob- served during training. Depending on how visually sim- ilar a test example is to the training categories the OSR task can be easy or extremely challenging. However the vast majority of previous…

2024

INQUIRE: A Natural World Text-to-Image Retrieval Benchmark

NeurIPS 2024poster

We introduce INQUIRE, a text-to-image retrieval benchmark designed to challenge multimodal vision-language models on expert-level queries. INQUIRE includes iNaturalist 2024 (iNat24), a new dataset of five million natural world images, along with 250 expert-level retrieval queries. These queries are…

2024

Improving Semantic Correspondence with Viewpoint-Guided Spherical Maps

CVPR 2024poster

Recent self-supervised models produce visual features that are not only effective at encoding image-level but also pixel-level semantics. They have been reported to obtain impressive results for dense visual semantic correspondence estimation even outperforming fully-supervised methods. Nevertheless…

Cited by 12SourcePDFScholar
2023

Active Learning-Based Species Range Estimation

NeurIPS 2023poster

We propose a new active learning approach for efficiently estimating the geographic range of a species from a limited number of on the ground observations. We model the range of an unmapped species of interest as the weighted combination of estimated ranges obtained from a set of different species.…

2023

Spatial Implicit Neural Representations for Global-Scale Species Mapping

ICML 2023poster

Estimating the geographical range of a species from sparse observations is a challenging and important geospatial prediction problem. Given a set of locations where a species has been observed, the goal is to build a model to predict whether the species is present or absent at any location. This pro…

2023

Virtual Occlusions Through Implicit Depth

CVPR 2023poster

For augmented reality (AR), it is important that virtual assets appear to 'sit among' real world objects. The virtual element should variously occlude and be occluded by real matter, based on a plausible depth ordering. This occlusion should be consistent over time as the viewer's camera moves. Unfo…

2022

Exploring Fine-Grained Audiovisual Categorization with the SSW60 Dataset

ECCV 2022poster

"We present a new benchmark dataset, Sapsucker Woods 60 (SSW60), for advancing research on audiovisual fine-grained categorization. While our community has made great strides in fine-grained visual categorization on images, the counterparts in audio and video fine-grained categorization are relative…

2022

On Label Granularity and Object Localization

ECCV 2022poster

"Weakly supervised object localization (WSOL) aims to learn representations that encode object location using only image-level category labels. However, many objects can be labeled at different levels of granularity. Is it an animal, a bird, or a great horned owl? Which image-level labels should we…

2022

When Does Contrastive Visual Representation Learning Work?

CVPR 2022poster

Recent self-supervised representation learning techniques have largely closed the gap between supervised and unsupervised learning on ImageNet classification. While the particulars of pretraining on ImageNet are now relatively well understood, the field still lacks widely accepted best practices for…

Cited by 152PDFScholar
2021

Benchmarking Representation Learning for Natural World Image Collections

CVPR 2021poster

Recent progress in self-supervised learning has resulted in models that are capable of extracting rich representations from image collections without requiring any explicit label supervision. However, to date the vast majority of these approaches have restricted themselves to training on standard be…

Cited by 191PDFcodeScholar
2021

Focus on the Positives: Self-Supervised Learning for Biodiversity Monitoring

ICCV 2021poster

We address the problem of learning self-supervised representations from unlabeled image collections. Unlike existing approaches that attempt to learn useful features by maximizing similarity between augmented versions of each input image or by speculatively picking negative samples, we instead also…

Cited by 32PDFcodeScholar
2021

Multi-Label Learning From Single Positive Labels

CVPR 2021poster

Predicting all applicable labels for a given image is known as multi-label classification. Compared to the standard multi-class case (where each image has only one label), it is considerably more challenging to annotate training data for multi-label classification. When the number of potential label…

Cited by 133PDFcodeScholar
2021

The Temporal Opportunist: Self-Supervised Multi-Frame Monocular Depth

CVPR 2021poster

Self-supervised monocular depth estimation networks are trained to predict scene depth using nearby frames as a supervision signal during training. However, for many applications, sequence information in the form of video frames is also available at test time. The vast majority of monocular networks…

Cited by 346PDFcodeScholar
2020

Learning Stereo from Single Images

ECCV 2020poster

Supervised deep networks are among the best methods for finding correspondences in stereo image pairs. Like all supervised approaches, these networks require ground truth data during training. However, collecting large quantities of accurate dense correspondence data is very challenging. We propose…

2019

Digging Into Self-Supervised Monocular Depth Estimation

ICCV 2019poster

Per-pixel ground-truth depth data is challenging to acquire at scale. To overcome this limitation, self-supervised learning has emerged as a promising alternative for training models to perform monocular depth estimation. In this paper, we propose a set of improvements, which together result in both…

Cited by 2896PDFcodeScholar
2019

Teaching Multiple Concepts to a Forgetful Learner

NeurIPS 2019poster

How can we help a forgetful learner learn multiple concepts within a limited time frame? While there have been extensive studies in designing optimal schedules for teaching a single concept given a learner's memory model, existing approaches for teaching multiple concepts are typically based on heur…

Cited by 29SourcePDFScholar
2018

Near-Optimal Machine Teaching via Explanatory Teaching Sets

AISTATS 2018poster

Modern applications of machine teaching for humans often involve domain-specific, non- trivial target hypothesis classes. To facilitate understanding of the target hypothesis, it is crucial for the teaching algorithm to use examples which are interpretable to the human learner. In this paper, we pro…

Cited by 0SourcePDFScholar
2018

Teaching Categories to Human Learners With Visual Explanations

CVPR 2018poster

We study the problem of computer-assisted teaching with explanations. Conventional approaches for machine teaching typically only provide feedback at the instance level e.g., the category or label of the instance. However, it is intuitive that clear explanations from a knowledgeable teacher can…

Cited by 86SourcePDFScholar
2018

The INaturalist Species Classification and Detection Dataset

CVPR 2018poster

Existing image classification datasets used in computer vision tend to have a uniform distribution of images across object categories. In contrast, the natural world is heavily imbalanced, as some species are more abundant and easier to photograph than others. To encourage further progress in challe…

2018

Understanding the Role of Adaptivity in Machine Teaching: The Case of Version Space Learners

NeurIPS 2018poster

In real-world applications of education, an effective teacher adaptively chooses the next example to teach based on the learner’s current state. However, most existing work in algorithmic machine teaching focuses on the batch setting, where adaptivity plays no role. In this paper, we study the case…

Cited by 53SourcePDFScholar
2017

Unsupervised Monocular Depth Estimation With Left-Right Consistency

CVPR 2017oral

Learning based methods have shown very promising results for the task of depth estimation in single images. However, most existing approaches treat depth prediction as a supervised regression problem and as a result, require vast quantities of corresponding ground truth depth data for training. Just…

Cited by 3871PDFcodeScholar
2016

Structured Prediction of Unobserved Voxels From a Single Depth Image

CVPR 2016oral

Building a complete 3D model of a scene, given only a single depth image, is underconstrained. To gain a full volumetric model, one needs either multiple views, or a single view together with a library of unambiguous 3D models that will fit the shape of each individual object in the scene. We hypot…

Cited by 206PDFScholar