← Search

Paul Henderson

12 accepted papers

2026

3D-ADAM: A Dataset for 3D Anomaly Detection in Additive Manufacturing

ICRA 2026poster

Surface defects are a primary source of yield loss in manufacturing, yet existing anomaly detection methods often fail in real-world deployment due to limited and unrepresentative datasets. To overcome this, we introduce 3D-ADAM, a 3D Anomaly Detection in Additive Manufacturing dataset, that is the …

2026

Is Training Necessary for Anomaly Detection?

ICML 2026poster

Current state-of-the-art multi-class unsupervised anomaly detection (MUAD) methods rely on training encoder–decoder models to reconstruct anomaly-free features. We first show these approaches have an inherent fidelity–stability dilemma in how they detect anomalies via reconstruction residuals. We th…

Cited by 0SourceScholar
2026

Masked Generative Policy for Robotic Control

ICLR 2026poster

We present Masked Generative Policy (MGP), a novel framework for visuomotor imitation learning. We represent actions as discrete tokens, and train a conditional masked transformer that generates tokens in parallel and then rapidly refines only low-confidence tokens. We further propose two new sampli…

Cited by 1SourcecodeScholar
2026

ReMatch: Boosting Representation through Matching for Multimodal Retrieval

CVPR 2026

We present ReMatch, a framework that leverages the generative strength of MLLMs for multimodal retrieval. Previous approaches treated an MLLM as a simple encoder, ignoring its generative nature, and under-utilising its compositional reasoning and world knowledge. We train the embedding MLLM end-to-e

Cited by 0SourcecodeScholar
2025

Flat'n'Fold: A Diverse Multi-Modal Dataset for Garment Perception and Manipulation

ICRA 2025

We present Flat'n'Fold, a novel large-scale dataset for garment manipulation that addresses critical gaps in existing datasets. Comprising 1,212 human and 887 robot demonstrations of flattening and folding 44 unique garments across 8 categories, Flat'n'Fold surpasses prior datasets in size, scope, a

Cited by 3SourcecodeScholar
2024

Denoising Diffusion via Image-Based Rendering

ICLR 2024poster

Generating 3D scenes is a challenging open problem, which requires synthesizing plausible content that is fully consistent in 3D space. While recent methods such as neural radiance fields excel at view synthesis and 3D reconstruction, they cannot synthesize plausible details in unobserved regions si…

Cited by 11SourcePDFScholar
2024

Understanding and Mitigating Human-Labelling Errors in Supervised Contrastive Learning

ECCV 2024poster

"Human-annotated vision datasets inevitably contain a fraction of human-mislabelled examples. While the detrimental effects of such mislabelling on supervised learning are well-researched, their influence on Supervised Contrastive Learning (SCL) remains largely unexplored. In this paper, we show tha…

Cited by 3SourcePDFScholar
2023

RenderDiffusion: Image Diffusion for 3D Reconstruction, Inpainting and Generation

CVPR 2023poster

Diffusion models currently achieve state-of-the-art performance for both conditional and unconditional image generation. However, so far, image diffusion models do not support tasks required for 3D understanding, such as view-consistent 3D generation or single-view object reconstruction. In this pap…

2022

Unsupervised Causal Generative Understanding of Images

NeurIPS 2022accept

We present a novel framework for unsupervised object-centric 3D scene understanding that generalizes robustly to out-of-distribution images. To achieve this, we design a causal generative model reflecting the physical process by which an image is produced, when a camera captures a scene containing m…

Cited by 10SourcePDFScholar
2020

Unsupervised object-centric video generation and decomposition in 3D

NeurIPS 2020poster

A natural approach to generative modeling of videos is to represent them as a composition of moving objects. Recent works model a set of 2D sprites over a slowly-varying background, but without considering the underlying 3D scene that gives rise to them. We instead propose to model a video as the vi…