← Search

Rajiv Ramnath

4 accepted papers

2026

DocVAL: Validated Chain-of-Thought Distillation for Grounded Document VQA

ICML 2026poster

Document visual question answering requires models not only to answer questions correctly, but also to precisely localize answers within complex document layouts. While large vision-language models (VLMs) achieve strong spatial grounding, their inference cost and latency limit real-world deployment;…

Cited by 0SourceScholar
2025

Frequency-Guided Masking for Enhanced Vision Self-Supervised Learning

ICLR 2025poster

We present a novel frequency-based Self-Supervised Learning (SSL) approach that significantly enhances its efficacy for pre-training. Prior work in this direction masks out pre-defined frequencies in the input image and employs a reconstruction loss to pre-train the model. While achieving promising…

2025

TaxaDiffusion: Progressively Trained Diffusion Model for Fine-Grained Species Generation

ICCV 2025poster

We propose TaxaDiffusion, a taxonomy-informed training framework for diffusion models to generate fine-grained animal images with high morphological and identity accuracy. Unlike standard approaches that treat each species as an independent category, TaxaDiffusion incorporates domain knowledge that…

2022

Retrieval Based Response Letter Generation For a Customer Care Setting

NAACL 2022industry

Letter-like communications (such as email) are a major means of customer relationship management within customer-facing organizations. These communications are initiated on a channel by requests from customers and then responded to by the organization on the same channel. For decades, the job has al…

Cited by 2SourcePDFScholar