← Search

Celso M. de Melo

16 accepted papers

2025

Aha! - Predicting What Matters Next: Online Highlight Detection Without Looking Ahead

NeurIPS 2025poster

Real-time understanding of continuous video streams is essential for intelligent agents operating in high-stakes environments, including autonomous vehicles, surveillance drones, and disaster response robots. Yet, most existing video understanding and highlight detection methods assume access to the…

Cited by 0SourceScholar
2025

Bisecle: Binding and Separation in Continual Learning for Video Language Understanding

NeurIPS 2025poster

Frontier vision-language models (VLMs) have made remarkable improvements in video understanding tasks. However, real-world videos typically exist as continuously evolving data streams (e.g., dynamic scenes captured by wearable glasses), necessitating models to continually adapt to shifting data dist…

Cited by 0SourceScholar
2025

ConceptAgent: LLM-Driven Precondition Grounding and Tree Search for Robust Task Planning and Execution

ICRA 2025

Robotic planning and execution in open-world environments is a complex problem due to the vast state spaces and high variability of task embodiment. Recent advances in perception algorithms, combined with Large Language Models (LLMs) for planning, offer promising solutions to these challenges, as th

Cited by 7SourceScholar
2025

Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Mutimodal Models

CVPR 2025highlight

Although large multimodal models (LMMs) have demonstrated remarkable capabilities in visual scene interpretation and reasoning, their capacity for complex and precise 3-dimensional spatial reasoning remains uncertain. Existing benchmarks focus predominantly on 2D spatial understanding and lack a fra…

2025

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models

CVPR 2025highlight

Humans naturally understand 3D spatial relationships, enabling complex reasoning like predicting collisions of vehicles from different directions. Current large multimodal models (LMMs), however, lack of this capability of 3D spatial reasoning. This limitation stems from the scarcity of 3D training…

Cited by 1SourcePDFScholar
2025

SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning

NeurIPS 2025poster

Despite recent advances on multi-modal models, 3D spatial reasoning remains a challenging task for state-of-the-art open-source and proprietary models. Recent studies explore data-driven approaches and achieve enhanced spatial reasoning performance by fine-tuning models on 3D-related visual question…

Cited by 0SourceScholar
2025

Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval

CVPR 2025poster

In this work, we tackle the problem of text-to-video retrieval (T2VR). Inspired by the success of late interaction techniques in text-document, text-image, and text-video retrieval, our approach, Video-ColBERT, introduces a simple and efficient mechanism for fine-grained similarity assessment betwee…

Cited by 0SourcePDFScholar
2024

Unsupervised Video Domain Adaptation with Masked Pre-Training and Collaborative Self-Training

CVPR 2024poster

In this work we tackle the problem of unsupervised domain adaptation (UDA) for video action recognition. Our approach which we call UNITE uses an image teacher model to adapt a video student model to the target domain. UNITE first employs self-supervised pre-training to promote discriminative featur…

2024

ViLCo-Bench: VIdeo Language COntinual learning Benchmark

NeurIPS 2024poster

Video language continual learning involves continuously adapting to information from video and text inputs, enhancing a model’s ability to handle new tasks while retaining prior knowledge. This field is a relatively under-explored area, and establishing appropriate datasets is crucial for facilitati…

2023

AZTR: Aerial Video Action Recognition with Auto Zoom and Temporal Reasoning

ICRA 2023poster

We propose a novel approach for aerial video action recognition. Our method is designed for videos captured using UAVs and can run on edge or mobile devices. We present a learning-based approach that uses customized auto zoom to automatically identify the human target and scale it appropriately. Thi…

Cited by 18SourceScholar
2023

STMT: A Spatial-Temporal Mesh Transformer for MoCap-Based Action Recognition

CVPR 2023poster

We study the problem of human action recognition using motion capture (MoCap) sequences. Unlike existing techniques that take multiple manual steps to derive standardized skeleton representations as model input, we propose a novel Spatial-Temporal Mesh Transformer (STMT) to directly model the mesh s…

2023

Synthetic-to-Real Domain Adaptation for Action Recognition: A Dataset and Baseline Performances

ICRA 2023poster

Human action recognition is a challenging problem, particularly when there is high variability in factors such as subject appearance, backgrounds and viewpoint. While deep neural networks (DNNs) have been shown to perform well on action recognition tasks, they typically require large amounts of high…

Cited by 36SourcecodeScholar
2022

Not Just Streaks: Towards Ground Truth for Single Image Deraining

ECCV 2022poster

"We propose a large-scale dataset of real-world rainy and clean image pairs and a method to remove degradations, induced by rain streaks and rain accumulation, from the image. As there exists no real-world dataset for deraining, current state-of-the-art methods rely on synthetic data and thus are li…

2020

Vision-Based Gesture Recognition in Human-Robot Teams Using Synthetic Data

IROS 2020poster

Building successful collaboration between humans and robots requires efficient, effective, and natural communication. Here we study a RGB-based deep learning approach for controlling robots through gestures (e.g., "follow me"). To address the challenge of collecting high-quality annotated data from…

Cited by 32SourceScholar