← Search

Subhashini Venugopalan

17 accepted papers

2026

MAgSeg: Segmentation of Agricultural Landscapes in High-Resolution Satellite Imagery using Multimodal Large Language Models

IJCAI 2026

Agricultural landscape segmentation in the Global South is challenging as it is characterized by fragmented plots, high intra-class variance, and a scarcity of labeled training data. Recent advances in segmentation have been made by Multimodal Large Language Models (MLLMs). However, current approach

Cited by 0Scholar
2025

CURIE: Evaluating LLMs on Multitask Scientific Long-Context Understanding and Reasoning

ICLR 2025poster

Scientific problem-solving involves synthesizing information while applying expert knowledge. We introduce CURIE, a scientific long-Context Understanding, Reasoning, and Information Extraction benchmark to measure the potential of Large Language Models (LLMs) in scientific problem-solving a…

2025

Speech Recognition with LLMs Adapted to Disordered Speech Using Reinforcement Learning

ICASSP 2025accepted

We introduce a large language model (LLM) capable of processing speech inputs and show that tuning it further with reinforcement learning on human preference (RLHF) enables it to adapt better to disordered speech than traditional fine-tuning. Our method replaces low-frequency text tokens in an LLM’s…

Cited by 0SourceScholar
2025

Towards a Single ASR Model That Generalizes to Disordered Speech

ICASSP 2025accepted

This study investigates the impact of integrating a dataset of disordered speech recordings (~1,000 hours) into the fine-tuning of a near state-of-the-art ASR baseline system. Contrary to what one might expect, despite the data being less than 1% of the training data of the ASR system, we find a con…

Cited by 0SourceScholar
2024

Large Language Models As A Proxy For Human Evaluation In Assessing The Comprehensibility Of Disordered Speech Transcription

ICASSP 2024accepted

Automatic Speech Recognition (ASR) systems, despite significant advances in recent years, still have much room for improvement particularly in the recognition of disordered speech. Even so, erroneous transcripts from ASR models can help people with disordered speech be better understood, especially…

Cited by 0SourceScholar
2024

SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers

NeurIPS 2024poster

Seeking answers to questions within long scientific research articles is a crucial area of study that aids readers in quickly addressing their inquiries. However, existing question-answering (QA) datasets based on scientific papers are limited in scale and focus solely on textual content. We introdu…

2023

Is Attention All That NeRF Needs?

ICLR 2023poster

We present Generalizable NeRF Transformer (GNT), a transformer-based architecture that reconstructs Neural Radiance Fields (NeRFs) and learns to render novel views on the fly from source views. While prior works on NeRFs optimize a scene representation by inverting a handcrafted rendering equation,…

2023

Speech Intelligibility Classifiers from 550k Disordered Speech Samples

ICASSP 2023accepted

We developed dysarthric speech intelligibility classifiers on 551,176 disordered speech samples contributed by a diverse set of 468 speakers, with a range of self-reported speaking disorders and rated for their overall intelligibility on a five-point scale. We trained three models following differen…

Cited by 0SourceScholar
2022

Context-Aware Abbreviation Expansion Using Large Language Models

NAACL 2022long

Motivated by the need for accelerating text entry in augmentative and alternative communication (AAC) for people with severe motor impairments, we propose a paradigm in which phrases are abbreviated aggressively as primarily word-initial letters. Our approach is to expand the abbreviations into full…

Cited by 46SourcePDFScholar
2022

Sparse Winning Tickets are Data-Efficient Image Recognizers

NeurIPS 2022accept

Improving the performance of deep networks in data-limited regimes has warranted much attention. In this work, we empirically show that “winning tickets” (small sub-networks) obtained via magnitude pruning based on the lottery ticket hypothesis, apart from being sparse are also effective recognizers…

2021

Guided Integrated Gradients: An Adaptive Path Method for Removing Noise

CVPR 2021poster

Integrated Gradients (IG) is a commonly used feature attribution method for deep neural networks. While IG has many desirable properties, the method often produces spurious/noisy pixel attributions in regions that are not related to the predicted class when applied to visual models. While this has b…

Cited by 137PDFcodeScholar
2021

Scaling Symbolic Methods using Gradients for Neural Model Explanation

ICLR 2021poster

Symbolic techniques based on Satisfiability Modulo Theory (SMT) solvers have been proposed for analyzing and verifying neural network properties, but their usage has been fairly limited owing to their poor scalability with larger networks. In this work, we propose a technique for combining gradient-…

2017

Captioning Images With Diverse Objects

CVPR 2017oral

Recent captioning models are limited in their ability to scale and describe concepts unseen in paired image-text corpora. We propose the Novel Object Captioner (NOC), a deep visual semantic captioning model that can describe a large number of object categories not present in existing image-caption…

Cited by 226PDFScholar
2016

Deep Compositional Captioning: Describing Novel Object Categories Without Paired Training Data

CVPR 2016oral

While recent deep neural network models have achieved promising results on the image captioning task, they rely largely on the availability of corpora with paired image and sentence captions to describe objects in context. In this work, we propose the Deep Compositional Captioner (DCC) to address th…

Cited by 346PDFScholar
2015

Long-Term Recurrent Convolutional Networks for Visual Recognition and Description

CVPR 2015poster

Models comprised of deep convolutional network layers have dominated recent image interpretation tasks; we investigate whether models which are also compositional, or "deep", temporally are effective on tasks involving visual sequences or label sequences. We develop a novel recurrent convolutional a…

Cited by 8345SourcePDFScholar
2015

Sequence to Sequence - Video to Text

ICCV 2015poster

Real-world videos often have complex dynamics; methods for generating open-domain video descriptions should be senstive to temporal structure and allow both input (sequence of frames) and output (sequence of words) of variable length. To approach this problem we propose a novel end-to-end sequence-t…

Cited by 1877PDFcodeScholar