← Search

Sundararajan Srinivasan

7 accepted papers

2025

CriSPO: Multi-Aspect Critique-Suggestion-guided Automatic Prompt Optimization for Text Generation

AAAI 2025technical

Existing automatic prompt engineering methods are typically designed for discriminative tasks, where new task prompts are iteratively refined with limited feedback from a single metric reflecting a single aspect. However, these approaches are suboptimal for generative tasks, which require more nuanc…

2025

Provable Meta-Learning with Low-Rank Adaptations

NeurIPS 2025poster

The power of foundation models (FMs) lies in their capacity to learn highly expressive representations that can be adapted to a broad spectrum of tasks. However, these pretrained models require additional training stages to become effective for downstream applications. In the multi-task setting, pri…

Cited by 0SourceScholar
2025

SEAL: Speaker Error Correction using Acoustic-conditioned Large Language Models

ICASSP 2025accepted

Speaker Diarization (SD) is a crucial component of modern end-to-end ASR pipelines. Traditional SD systems, which are typically audio-based and operate independently of ASR, often introduce speaker errors, particularly during speaker transitions and overlapping speech. Recently, language models incl…

Cited by 0SourceScholar
2024

SpeechGuard: Exploring the Adversarial Robustness of Multi-modal Large Language Models

ACL 2024findings

Integrated Speech and Large Language Models (SLMs) that can follow speech instructions and generate relevant text responses have gained popularity lately. However, the safety and robustness of these models remains largely unclear. In this work, we investigate the potential vulnerabilities of such in…

2023

End-to-End Single-Channel Speaker-Turn Aware Conversational Speech Translation

EMNLP 2023long main

Conventional speech-to-text translation (ST) systems are trained on single-speaker utterances, and they may not generalize to real-life scenarios where the audio contains conversations by multiple speakers. In this paper, we tackle single-channel multi-speaker conversational ST with an end-to-end an…

Cited by 0SourcecodeScholar
2022

Enhancing Contrastive Learning with Temporal Cognizance for Audio-Visual Representation Generation

ICASSP 2022accepted

Audio-visual data allows us to leverage different modalities for downstream tasks. The idea being individual streams can complement each other in the given task, thereby resulting in a model with improved performance. In this work, we present our experimental results on action recognition and video…

Cited by 0SourceScholar
2022

Representation Learning Through Cross-Modal Conditional Teacher-Student Training For Speech Emotion Recognition

ICASSP 2022accepted

Generic pre-trained speech and text representations promise to reduce the need for large labeled datasets on specific speech and language tasks. However, it is not clear how to effectively adapt these representations for speech emotion recognition. Recent public benchmarks show the efficacy of sever…

Cited by 0SourceScholar