← Search

Somali Chaterji

5 accepted papers

2025

Improving Semi-Supervised Semantic Segmentation with Sliced-Wasserstein Feature Alignment and Uniformity

CVPR 2025poster

Semi-supervised semantic segmentation with consistencyregularization capitalizes on unlabeled images to enhancethe accuracy of pixel-level segmentation. Current consistencylearning methods primarily rely on the consistency loss be-tween pseudo-labels and unlabeled images, neglecting the in-formation…

Cited by 0SourcePDFScholar
2025

Learning to Inference Adaptively for Multimodal Large Language Models

ICCV 2025poster

Multimodal Large Language Models (MLLMs) have shown impressive capabilities in visual reasoning, yet come with substantial computational cost, limiting their deployment in resource-constrained settings. Despite recent effort on improving the efficiency of MLLMs, prior solutions fall short in respond…

Cited by 0SourcePDFScholar
2025

SKALD: Learning-Based Shot Assembly for Coherent Multi-Shot Video Creation

ICCV 2025poster

We present SKALD, a multi-shot video assembly method that constructs coherent video sequences from candidate shots with minimal reliance on text. Central to our approach is the Learned Clip Assembly (LCA) score, a learning-based metric that measures temporal and semantic relationships between shots…

Cited by 0SourcePDFScholar
2024

ReCON: Training-Free Acceleration for Text-to-Image Synthesis with Retrieval of Concept Prompt Trajectories

ECCV 2024poster

"Text-to-image diffusion models excel in generating photo-realistic images but are hampered by slow processing times. Training-free retrieval-based acceleration methods, which leverage pre-generated “trajectories,” have been introduced to address this. Yet, these methods often lack diversity and fid…

2022

SmartAdapt: Multi-Branch Object Detection Framework for Videos on Mobiles

CVPR 2022poster

Several recent works seek to create lightweight deep networks for video object detection on mobiles. We observe that many existing detectors, previously deemed computationally costly for mobiles, intrinsically support adaptive inference, and offer a multi-branch object detection framework (MBODF). H…

Cited by 15PDFScholar