← Search

Bertram E Shi

9 accepted papers

2024

OAT: Object-Level Attention Transformer for Gaze Scanpath Prediction

ECCV 2024poster

"Visual search is important in our daily life. The efficient allocation of visual attention is critical to effectively complete visual search tasks. Prior research has predominantly modelled the spatial allocation of visual attention in images at the pixel level, e.g. using a saliency map. However,…

2022

HGCN-GJS: Hierarchical Graph Convolutional Network with Groupwise Joint Sampling for Trajectory Prediction

IROS 2022poster

Pedestrian trajectory prediction is of great importance for downstream tasks, such as autonomous driving and mobile robot navigation. Realistic models of the social interactions within the crowd is crucial for accurate pedestrian trajectory prediction. However, most existing methods do not capture g…

Cited by 16SourceScholar
2021

AVGCN: Trajectory Prediction using Graph Convolutional Networks Guided by Human Attention

ICRA 2021poster

Pedestrian trajectory prediction is a critical yet challenging task especially for crowded scenes. We suggest that introducing an attention mechanism to infer the importance of different neighbors is critical for accurate trajectory prediction in scenes with varying crowd size. In this work, we prop…

Cited by 35SourceScholar
2020

Robot Navigation in Crowds by Graph Convolutional Networks With Attention Learned From Human Gaze

RA-L 2020

Safe and efficient crowd navigation for mobile robot is a crucial yet challenging task. Previous work has shown the power of deep reinforcement learning frameworks to train efficient policies. However, their performance deteriorates when the crowd size grows. We suggest that this can be addressed by

Cited by 144SourceScholar
2019

Gaze Training by Modulated Dropout Improves Imitation Learning

IROS 2019poster

Imitation learning by behavioral cloning is a prevalent method that has achieved some success in vision-based autonomous driving. The basic idea behind behavioral cloning is to have the neural network learn from observing a human expert's behavior. Typically, a convolutional neural network learns to…

Cited by 27SourceScholar
2017

Lattice Long Short-Term Memory for Human Action Recognition

ICCV 2017poster

Human actions captured in video sequences are three-dimensional signals characterizing visual appearance and motion dynamics. To learn action patterns, existing methods adopt Convolutional and/or Recurrent Neural Networks (CNNs and RNNs). CNN based methods are effective in learning spatial appearanc…

Cited by 232PDFScholar
2015

Human Action Recognition Using Factorized Spatio-Temporal Convolutional Networks

ICCV 2015poster

Human actions in video sequences are three-dimensional (3D) spatio-temporal signals characterizing both the visual appearance and motion dynamics of the involved humans and objects. Inspired by the success of convolutional neural networks (CNN) for image classification, recent attempts have been mad…

Cited by 745PDFScholar