← Search

Junhua Mao

7 accepted papers

2023

Pedestrian Crossing Action Recognition and Trajectory Prediction with 3D Human Keypoints

ICRA 2023poster

Accurate understanding and prediction of human behaviors are critical prerequisites for autonomous vehicles, especially in highly dynamic and interactive scenarios such as intersections in dense urban areas. In this work, we aim at identifying crossing pedestrians and predicting their future traject…

Cited by 19SourceScholar
2020

STINet: Spatio-Temporal-Interactive Network for Pedestrian Detection and Trajectory Prediction

CVPR 2020poster

Detecting pedestrians and predicting future trajectories for them are critical tasks for numerous applications, such as autonomous driving. Previous methods either treat the detection and prediction as separate tasks or simply add a trajectory regression head on top of a detector. In this work, we p…

Cited by 81PDFScholar
2016

CNN-RNN: A Unified Framework for Multi-Label Image Classification

CVPR 2016oral

While deep convolutional neural networks (CNNs) have shown a great success in single-label image classification, it is important to note that most real world images contain multiple labels, which could correspond to different objects, scenes, actions and attributes in an image. Traditional approache…

Cited by 1717PDFScholar
2016

Generation and Comprehension of Unambiguous Object Descriptions

CVPR 2016oral

We propose a method that can generate an unambiguous description (known as a referring expression) of a specific object or region in an image, and which can also comprehend or interpret such an expression to infer which object is being described. We show that our method outperforms previous methods…

Cited by 1580PDFcodeScholar
2016

Training and Evaluating Multimodal Word Embeddings with Large-scale Web Annotated Images

NeurIPS 2016poster

In this paper, we focus on training and evaluating effective word embeddings with both text and visual information. More specifically, we introduce a large-scale dataset with 300 million sentences describing over 40 million images crawled and downloaded from publicly available Pins (i.e. an image wi…

Cited by 64SourcePDFScholar
2015

Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question

NeurIPS 2015poster

In this paper, we present the mQA model, which is able to answer questions about the content of an image. The answer can be a sentence, a phrase or a single word. Our model contains four components: a Long Short-Term Memory (LSTM) to extract the question representation, a Convolutional Neural Networ…

Cited by 692SourcePDFScholar
2015

Learning Like a Child: Fast Novel Visual Concept Learning From Sentence Descriptions of Images

ICCV 2015poster

In this paper, we address the task of learning novel visual concepts, and their interactions with other concepts, from a few images with sentence descriptions. Using linguistic context and visual features, our method is able to efficiently hypothesize the semantic meaning of new words and add them t…

Cited by 195PDFScholar