← Search

Onkar Susladkar

5 accepted papers

2026

Best of Both Worlds: Multimodal Reasoning and Generation via Unified Discrete Flow Matching

ICML 2026poster

We propose UniDFlow, a unified discrete flow-matching framework for multimodal understanding, generation, and editing. It decouples understanding and generation via task-specific low-rank adapters, avoiding objective interference and representation entanglement, while a novel reference-based multimo…

Cited by 0SourceScholar
2026

PyraTok: Language-Aligned Pyramidal Tokenizer for Video Understanding and Generation

CVPR 2026

Discrete video VAEs underpin modern text-to-video generation and video understanding systems, yet existing tokenizers typically learn visual codebooks at a single scale with limited vocabularies and shallow language supervision, leading to poor cross-modal alignment and zero-shot transfer. We introd

Cited by 0SourcecodeScholar
2026

RewardFlow: Generate Images by Optimizing What You Reward

CVPR 2026

RewardFlow is a zero-shot, training-free framework for text-guided image editing and generation based on reward-guided Langevin dynamics. We steer pretrained diffusion and flow-matching models at inference time using a diverse set of differentiable rewards, and control their influence with a prompt-

Cited by 0SourceScholar
2025

ViCTr: Vital Consistency Transfer for Pathology Aware Image Synthesis

ICCV 2025poster

We introduce ViCTr (Vital Consistency Transfer), a framework for advancing medical image synthesis through a principled integration with Rectified Flow trajectories. Unlike traditional approaches, we modify the Tweedie formulation to accommodate linear trajectories within the Rectified Flow framewor…

2023

SLBERT: A Novel Pre-Training Framework for Joint Speech and Language Modeling

ICASSP 2023accepted

We propose SLBERT (Speech and Language pre-training framework for BERT), an end-to-end trainable framework for learning joint representations of speech and language modalities. We enhance the well-known BERT architecture to provide a dual-stream multimodal architecture that processes both speech and…

Cited by 0SourceScholar