← Search

Pradeep Yarlagadda

2 accepted papers

2021

ViNet: Pushing the limits of Visual Modality for Audio-Visual Saliency Prediction

IROS 2021poster

We propose the ViNet architecture for audio-visual saliency prediction. ViNet is a fully convolutional encoder-decoder architecture. The encoder uses visual features from a network trained for action recognition, and the decoder infers a saliency map via trilinear interpolation and 3D convolutions,…

Cited by 100SourcecodeScholar