← Search

Gerardo Aragon-Camarasa

15 accepted papers

2026

3D-ADAM: A Dataset for 3D Anomaly Detection in Additive Manufacturing

ICRA 2026poster

Surface defects are a primary source of yield loss in manufacturing, yet existing anomaly detection methods often fail in real-world deployment due to limited and unrepresentative datasets. To overcome this, we introduce 3D-ADAM, a 3D Anomaly Detection in Additive Manufacturing dataset, that is the …

2026

Is Training Necessary for Anomaly Detection?

ICML 2026poster

Current state-of-the-art multi-class unsupervised anomaly detection (MUAD) methods rely on training encoder–decoder models to reconstruct anomaly-free features. We first show these approaches have an inherent fidelity–stability dilemma in how they detect anomalies via reconstruction residuals. We th…

Cited by 0SourceScholar
2026

Masked Generative Policy for Robotic Control

ICLR 2026poster

We present Masked Generative Policy (MGP), a novel framework for visuomotor imitation learning. We represent actions as discrete tokens, and train a conditional masked transformer that generates tokens in parallel and then rapidly refines only low-confidence tokens. We further propose two new sampli…

Cited by 1SourcecodeScholar
2025

Can Real-to-Sim Approaches Capture Dynamic Fabric Behavior for Robotic Fabric Manipulation?

IROS 2025

This paper presents a rigorous evaluation of Real-to-Sim parameter estimation approaches for fabric manipulation in robotics. The study systematically assesses three state-of-the-art approaches, namely two differential pipelines and a data-driven approach. We also devise a novel physics-informed neu

Cited by 1SourcecodeScholar
2025

Flat'n'Fold: A Diverse Multi-Modal Dataset for Garment Perception and Manipulation

ICRA 2025

We present Flat'n'Fold, a novel large-scale dataset for garment manipulation that addresses critical gaps in existing datasets. Comprising 1,212 human and 887 robot demonstrations of flattening and folding 44 unique garments across 8 categories, Flat'n'Fold surpasses prior datasets in size, scope, a

Cited by 3SourcecodeScholar
2024

LaCViT: A Label-Aware Contrastive Fine-Tuning Framework for Vision Transformers

ICASSP 2024accepted

Vision Transformers (ViTs) have emerged as popular models in computer vision, demonstrating state-of-the-art performance across various tasks. This success typically follows a two-stage strategy involving pre-training on large-scale datasets using self-supervised signals, such as masked random patch…

Cited by 0SourceScholar
2024

Multiway-Adapter: Adapting Multimodal Large Language Models for Scalable Image-Text Retrieval

ICASSP 2024accepted

As Multimodal Large Language Models (MLLMs) grow in size, adapting them to specialized tasks becomes increasingly challenging due to high computational and memory demands. Indeed, traditional fine-tuning methods are costly, due to the need for extensive, task-specific training. While efficient adapt…

Cited by 0SourceScholar
2024

RoboLLM: Robotic Vision Tasks Grounded on Multimodal Large Language Models

ICRA 2024poster

Robotic vision applications often necessitate a wide range of visual perception tasks, such as object detection, segmentation, and identification. While there have been substantial advances in these individual tasks, integrating specialized models into a unified vision pipeline presents significant…

Cited by 21SourcecodeScholar
2024

TELESIM: A Modular and Plug-and-Play Framework for Robotic Arm Teleoperation using a Digital Twin

ICRA 2024poster

Teleoperating robotic arms can be a challenging task for non-experts, particularly when using complex control devices or interfaces. To address the limitations and challenges of existing teleoperation frameworks, such as cognitive strain, control complexity, robot compatibility, and user evaluation,…

Cited by 7SourceScholar
2024

Understanding and Mitigating Human-Labelling Errors in Supervised Contrastive Learning

ECCV 2024poster

"Human-annotated vision datasets inevitably contain a fraction of human-mislabelled examples. While the detrimental effects of such mislabelling on supervised learning are well-researched, their influence on Supervised Contrastive Learning (SCL) remains largely unexplored. In this paper, we show tha…

Cited by 3SourcePDFScholar
2022

A Continuous Robot Vision Approach for Predicting Shapes and Visually Perceived Weights of Garments

RA-L 2022

We present a continuous perception approach that learns geometric and physical similarities between garments by continuously observing a garment while a robot picks it up from a table. The aim is to capture and encode geometric and physical characteristics of a garment into a manifold where a decisi

Cited by 9SourceScholar
2020

On Simple Reactive Neural Networks for Behaviour-Based Reinforcement Learning

ICRA 2020poster

We present a behaviour-based reinforcement learning approach, inspired by Brook's subsumption architecture, in which simple fully connected networks are trained as reactive behaviours. Our working assumption is that a pick and place robotic task can be simplified by leveraging domain knowledge of a…

Cited by 25SourcecodeScholar
2017

Single-shot clothing category recognition in free-configurations with application to autonomous clothes sorting

IROS 2017poster

This paper proposes a single-shot approach for recognising clothing categories from 2.5D features. We propose two visual features, BSP (B-Spline Patch) and TSD (Topology Spatial Distances) for this task. The local BSP features are encoded by LLC (Locality-constrained Linear Coding) and fused with th…

Cited by 46SourceScholar
2016

Recognising the clothing categories from free-configuration using Gaussian-Process-based interactive perception

ICRA 2016

In this paper, we propose a Gaussian Process-based interactive perception approach for recognising highly-wrinkled clothes. We have integrated this recognition method within a clothes sorting pipeline for the pre-washing stage of an autonomous laundering process. Our approach differs from reported c

Cited by 31SourceScholar
2015

Accurate garment surface analysis using an active stereo robot head with application to dual-arm flattening

ICRA 2015poster

We present a visually guided, dual-arm, industrial robot system that is capable of autonomously flattening garments by means of a novel visual perception pipeline that fully interprets high-quality RGB-D images of a clothing scene based on an active stereo robot head. A segmented clothing range map…

Cited by 133SourceScholar