← Search

Akash Gupta

5 accepted papers

2026

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

ICML 2026poster

Large Multimodal Models (LMMs) exhibit shortfalls when interpreting images and, by some measures, have poorer spatial cognition than young children or animals. Despite this, they attain high scores on many popular visual benchmarks, with headroom rapidly eroded by surging model progress. To address …

Cited by 0SourceScholar
2024

LLM Task Interference: An Initial Study on the Impact of Task-Switch in Conversational History

EMNLP 2024main

With the recent emergence of powerful instruction-tuned large language models (LLMs), various helpful conversational Artificial Intelligence (AI) systems have been deployed across many applications. When prompted by users, these AI systems successfully perform a wide range of tasks as part of a conv…

2023

MODEFORMER: Modality-Preserving Embedding For Audio-Video Synchronization Using Transformers

ICASSP 2023accepted

Lack of audio-video synchronization is a common problem during television broadcasts and video conferencing, leading to an unsatisfactory viewing experience. A widely accepted paradigm is to create an error detection mechanism that identifies the cases when audio is leading or lagging. We propose Mo…

Cited by 0SourceScholar
2020

Non-Adversarial Video Synthesis With Learned Priors

CVPR 2020poster

Most of the existing works in video synthesis focus on generating videos using adversarial learning. Despite their success, these methods often require input reference frame or fail to generate diverse videos from the given data distribution, with little to no uniformity in the quality of videos tha…

Cited by 24PDFcodeScholar