← Search

Jiyang Gao

12 accepted papers

2021

HDMapGen: A Hierarchical Graph Generative Model of High Definition Maps

CVPR 2021poster

High Definition (HD) maps are maps with precise definitions of road lanes with rich semantics of the traffic rules. They are critical for several key stages in an autonomous driving system, including motion forecasting and planning. However, there are only a small amount of real-world road topologie…

Cited by 71PDFScholar
2020

STINet: Spatio-Temporal-Interactive Network for Pedestrian Detection and Trajectory Prediction

CVPR 2020poster

Detecting pedestrians and predicting future trajectories for them are critical tasks for numerous applications, such as autonomous driving. Previous methods either treat the detection and prediction as separate tasks or simply add a trajectory regression head on top of a detector. In this work, we p…

Cited by 81PDFScholar
2020

VectorNet: Encoding HD Maps and Agent Dynamics From Vectorized Representation

CVPR 2020poster

Behavior prediction in dynamic, multi-agent systems is an important problem in the context of self-driving cars, due to the complex representations and interactions of road components, including moving agents (e.g. pedestrians and vehicles) and road context information (e.g. lanes, traffic lights).…

Cited by 1022PDFScholar
2019

End-to-End Multi-View Fusion for 3D Object Detection in LiDAR Point Clouds

CoRL 2019

Recent work on 3D object detection advocates point cloud voxelization in birds-eye view, where objects preserve their physical dimensions and are naturally separable. When represented in this view, however, point clouds are sparse and have highly variable point density, which may cause detectors dif

Cited by 0SourcePDFScholar
2019

NOTE-RCNN: NOise Tolerant Ensemble RCNN for Semi-Supervised Object Detection

ICCV 2019poster

The labeling cost of large number of bounding boxes is one of the main challenges for training modern object detectors. To reduce the dependence on expensive bounding box annotations, we propose a new semi-supervised object detection formulation, in which a few seed box level annotations and a large…

Cited by 123PDFScholar
2018

Motion-Appearance Co-Memory Networks for Video Question Answering

CVPR 2018poster

Video Question Answering (QA) is an important task in understanding video temporal structure. We observe that there are three unique attributes of video QA compared with image QA: (1) it deals with long sequences of images containing richer information not only in quantity but also in variety; (2) m…

Cited by 310SourcePDFScholar
2017

TURN TAP: Temporal Unit Regression Network for Temporal Action Proposals

ICCV 2017poster

We address the problem of Temporal Action Proposal (TAP) generation. This is an important problem, as fast extraction of semantically important (e.g. human actions) segments from untrimmed videos is an important step for large-scale video analysis. To tackle this problem, we propose a novel Temporal…

Cited by 486PDFcodeScholar