← Search

Jose M. F. Moura

9 accepted papers

2025

FedBaF: Federated Learning Aggregation Biased by a Foundation Model

AISTATS 2025poster

Foundation models are now a major focus of leading technology organizations due to their ability to generalize across diverse tasks. Existing approaches for adapting foundation models to new applications often rely on Federated Learning (FL) and disclose the foundation model weights to clients when…

Cited by 0SourceScholar
2018

Adversarial Geometry-Aware Human Motion Prediction

ECCV 2018poster

We explore an approach to forecasting human motion in a few milliseconds given an input 3D skeleton sequence based on a recurrent encoder-decoder framework. Current approaches suffer from the problem of prediction discontinuities and may fail to predict human-like motion in longer time horizons due…

Cited by 317SourcePDFScholar
2018

Few-Shot Human Motion Prediction via Meta-Learning

ECCV 2018poster

Human motion prediction, forecasting human motion in a few milliseconds conditioning on a historical 3D skeleton sequence, is a long-standing problem in computer vision and robotic vision. Existing forecasting algorithms rely on extensive annotated motion capture data and are brittle to novel action…

Cited by 155SourcePDFScholar
2018

Visual Coreference Resolution in Visual Dialog using Neural Module Networks

ECCV 2018poster

Visual dialog entails answering a series of questions grounded in an image, using dialog history as context. In addition to the challenges found in visual question answering (VQA), which can be seen as one-round dialog, visual dialog encompasses several more. We focus on one such problem called ‘vis…

2017

FCN-rLSTM: Deep Spatio-Temporal Neural Networks for Vehicle Counting in City Cameras

ICCV 2017poster

In this paper, we develop deep spatio-temporal neural networks to sequentially count vehicles from low quality videos captured by city cameras (citycams). Citycam videos have low resolution, low frame rate, high occlusion and large perspective, making most existing methods lose their efficacy. To ov…

Cited by 272PDFScholar
2017

Learning Cooperative Visual Dialog Agents With Deep Reinforcement Learning

ICCV 2017oral

We introduce the first goal-driven training for visual question answering and dialog agents. Specifically, we pose a cooperative `image guessing' game between two agents -- Qbot and Abot -- who communicate in natural language dialog so that Qbot can select an unseen image from a lineup of images. We…

Cited by 493PDFcodeScholar
2017

Understanding Traffic Density From Large-Scale Web Camera Data

CVPR 2017poster

Understanding traffic density from large-scale web camera (webcam) videos is a challenging problem because such videos have low spatial and temporal resolution, high occlusion and large perspective. To deeply understand traffic density, we explore both optimization based and deep learning based meth…

Cited by 191PDFcodeScholar
2016

Visual Word2Vec (vis-w2v): Learning Visually Grounded Word Embeddings Using Abstract Scenes

CVPR 2016poster

We propose a model to learn visually grounded word embeddings (vis-w2v) to capture visual notions of semantic relatedness. While word embeddings trained using text have been extremely successful, they cannot uncover notions of semantic relatedness implicit in our visual world. For instance, although…

Cited by 120PDFScholar