← Search

Gouthaman KV

3 accepted papers

2026

A robust PPG foundation model using multimodal physiological supervision

ICML 2026poster

Photoplethysmography (PPG), a non-invasive measure of changes in blood volume, is widely used in both wearable devices and clinical settings. Recent PPG foundation models either use open-source ICU datasets with pretraining paradigms that require high-quality data and thus complicate generalization …

Cited by 0SourceScholar
2025

Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval

ICASSP 2025accepted

Content creators often use music to enhance their videos, from soundtracks in movies to background music in video blogs and social media content. However, identifying the best music for a video can be a difficult and time-consuming task. To address this challenge, we propose a novel framework for au…

Cited by 0SourceScholar
2020

Reducing Language Biases in Visual Question Answering with Visually-Grounded Question Encoder

ECCV 2020poster

Recent studies have shown that current VQA models are heavily biased on the language priors in the train set to answer the question, irrespective of the image. E.g., overwhelmingly answer ""what sport is"" as ""tennis"" or ""what color banana"" as ""yellow."" This behavior restricts them from real-w…

Cited by 94SourcePDFScholar