← Search

Keshav Bhandari

3 accepted papers

2025

Text2midi: Generating Symbolic Music from Captions

AAAI 2025technical

This paper introduces text2midi, an end-to-end model to generate MIDI files from textual descriptions. Leveraging the growing popularity of multimodal generative approaches, text2midi capitalizes on the extensive availability of textual data and the success of large language models (LLMs). Our end-t…

2022

Learning Omnidirectional Flow in 360° Video via Siamese Representation

ECCV 2022poster

"Optical flow estimation in omnidirectional videos faces two significant issues: the lack of benchmark datasets and the challenge of adapting perspective video-based methods to accommodate the omnidirectional nature. This paper proposes the first perceptually natural-synthetic omnidirectional benchm…

2022

VoiceBlock: Privacy through Real-Time Adversarial Attacks with Audio-to-Audio Models

NeurIPS 2022accept

As governments and corporations adopt deep learning systems to collect and analyze user-generated audio data, concerns about security and privacy naturally emerge in areas such as automatic speaker recognition. While audio adversarial examples offer one route to mislead or evade these invasive syste…