← Search

Austin Reiter

8 accepted papers

2021

Cross-Modal Retrieval Augmentation for Multi-Modal Classification

EMNLP 2021finding

Recent advances in using retrieval components over external knowledge sources have shown impressive results for a variety of downstream tasks in natural language processing. Here, we explore the use of unstructured external knowledge sources of images and their corresponding captions for improving v…

Cited by 29SourcePDFScholar
2021

Exploring Visual Engagement Signals for Representation Learning

ICCV 2021poster

Visual engagement in social media platforms comprises interactions with photo posts including comments, shares, and likes. In this paper, we leverage such visual engagement clues as supervisory signals for representation learning. However, learning from engagement signals is non-trivial as it is not…

Cited by 14PDFcodeScholar
2021

Intentonomy: A Dataset and Study Towards Human Intent Understanding

CVPR 2021poster

An image is worth a thousand words, conveying information that goes beyond the physical visual content therein. In this paper, we study the intent behind social media images with an aim to analyze how visual information can help the recognition of human intent. Towards this goal, we introduce an int…

Cited by 41PDFcodeScholar
2021

When in Doubt: Improving Classification Performance with Alternating Normalization

EMNLP 2021finding

We introduce Classification with Alternating Normalization (CAN), a non-parametric post-processing step for classification. CAN improves classification accuracy for challenging examples by re-adjusting their predicted class probability distribution using the predicted class distributions of high-con…

2018

A Deep Learning Based Alternative to Beamforming Ultrasound Images

ICASSP 2018accepted

Deep learning methods are capable of performing sophisticated tasks when applied to a myriad of artificial intelligent (AI) research fields. In this paper, we introduce a novel approach to replace the inherently flawed beamforming step during ultrasound image formation by applying deep learning dire…

Cited by 0SourceScholar
2018

Vision-Based Calibration of Dual RCM-Based Robot Arms in Human-Robot Collaborative Minimally Invasive Surgery

RA-L 2018

This letter reports the development of a vision-based calibration method for dual remote center-of-motion (RCM) based robot arms in a human-robot collaborative minimally invasive surgery (MIS) scenario. The method does not require any external tracking sensors and directly uses images captured by th

Cited by 69SourceScholar
2017

Temporal Convolutional Networks for Action Segmentation and Detection

CVPR 2017poster

The ability to identify and temporally segment fine-grained human actions throughout a video is crucial for robotics, surveillance, education, and beyond. Typical approaches decouple this problem by first extracting local spatiotemporal features from video frames and then feeding them into a tempora…

Cited by 2196PDFScholar
2015

Beyond Spatial Pooling: Fine-Grained Representation Learning in Multiple Domains

CVPR 2015poster

Object recognition systems have shown great progress over recent years. However, creating object representations that are robust to changes in viewpoint while capturing local visual details continues to be a challenge. In particular, recent convolutional architectures employ spatial pooling to achie…

Cited by 39SourcePDFScholar