← Search

Shruti Palaskar

10 accepted papers

2026

VLSU: Mapping the Limits of Joint Multimodal Understanding for AI Safety

ICLR 2026poster

Safety evaluation of multimodal foundation models often treats vision and language inputs separately, missing risks from joint interpretation where benign content becomes harmful in combination. Existing approaches also fail to distinguish clearly unsafe content from borderline cases, leading to pro…

Cited by 0SourcecodeScholar
2022

End-to-End Speech Summarization Using Restricted Self-Attention

ICASSP 2022accepted

Speech summarization is typically performed by using a cascade of speech recognition and text summarization models. End-to-end modeling of speech summarization models is challenging due to memory and compute constraints arising from long input audio sequences. Recent work in document summarization h…

Cited by 0SourceScholar
2022

On Advances in Text Generation from Images Beyond Captioning: A Case Study in Self-Rationalization

EMNLP 2022finding

Combining the visual modality with pretrained language models has been surprisingly effective for simple descriptive tasks such as image captioning. More general text generation however remains elusive. We take a step back and ask: How do these models work for more complex generative tasks, i.e. con…

Cited by 3SourcePDFScholar
2021

How2Sign: A Large-Scale Multimodal Dataset for Continuous American Sign Language

CVPR 2021poster

One of the factors that have hindered progress in the areas of sign language recognition, translation, and production is the absence of large annotated datasets. Towards this end, we introduce How2Sign, a multimodal and multiview continuous American Sign Language (ASL) dataset, consisting of a paral…

Cited by 257PDFcodeScholar
2020

ASR Error Correction and Domain Adaptation Using Machine Translation

ICASSP 2020accepted

Off-the-shelf pre-trained Automatic Speech Recognition (ASR) systems are an increasingly viable service for companies of any size building speech-based products. While these ASR systems are trained on large amounts of data, domain mismatch is still an issue for many such parties that want to use thi…

Cited by 0SourceScholar
2019

Learning from Multiview Correlations in Open-domain Videos

ICASSP 2019accepted

An increasing number of datasets contain multiple views, such as video, sound and automatic captions. A basic challenge in representation learning is how to leverage multiple views to learn better representations. This is further complicated by the existence of a latent alignment between views, such…

Cited by 0SourceScholar
2019

Multimodal Grounding for Sequence-to-sequence Speech Recognition

ICASSP 2019accepted

Humans are capable of processing speech by making use of multiple sensory modalities. For example, the environment where a conversation takes place generally provides semantic and/or acoustic context that helps us to resolve ambiguities or to recall named entities. Motivated by this, there have been…

Cited by 0SourceScholar
2018

Linguistic Unit Discovery from Multi-Modal Inputs in Unwritten Languages: Summary of the "Speaking Rosetta" JSALT 2017 Workshop

ICASSP 2018accepted

We summarize the accomplishments of a multi-disciplinary workshop exploring the computational and scientific issues surrounding the discovery of linguistic units (subwords and words) in a language without orthography. We study the replacement of orthographic transcriptions by images and/or translate…

Cited by 0SourceScholar