← Search

Rishubh Singh

5 accepted papers

2026

MULTIMODALITY AS SUPERVISION: SELF-SUPERVISED SPECIALIZATION TO THE TEST ENVIRONMENT VIA MULTIMODALITY

ICLR 2026poster

The common approach for developing a vision model is generalism, which involves training on a large diverse dataset to cover the varied deployment environments and leads to a model that is expected to solve the problem everywhere. However, many practical applications need to operate in a specific te…

Cited by 0SourcecodeScholar
2024

How Far Can a 1-Pixel Camera Go? Solving Vision Tasks using Photoreceptors and Computationally Designed Visual Morphology

ECCV 2024poster

"A de facto standard approach in solving computer vision tasks is to use a common high-resolution camera and choose its placement on an agent based on human intuition. On the other hand, extremely simple and well-designed visual sensors found throughout nature allow many organisms to exhibit diverse…

Cited by 0SourcePDFScholar
2024

OLAF: A Plug-and-Play Framework for Enhanced Multi-object Multi-part Scene Parsing

ECCV 2024poster

"Multi-object multi-part scene segmentation is a challenging task whose complexity scales exponentially with part granularity and number of scene objects. To address the task, we propose a plug-and-play approach termed OLAF. First, we augment the input (RGB) with channels containing object-based str…

2023

Edges to Shapes to Concepts: Adversarial Augmentation for Robust Vision

CVPR 2023poster

Recent work has shown that deep vision models tend to be overly dependent on low-level or "texture" features, leading to poor generalization. Various data augmentation strategies have been proposed to overcome this so-called texture bias in DNNs. We propose a simple, lightweight adversarial augmenta…

Cited by 7SourcePDFScholar
2022

FLOAT: Factorized Learning of Object Attributes for Improved Multi-Object Multi-Part Scene Parsing

CVPR 2022poster

Multi-object multi-part scene parsing is a challenging task which requires detecting multiple object classes in a scene and segmenting the semantic parts within each object. In this paper, we propose FLOAT, a factorized label space framework for scalable multi-object multi-part parsing. Our framewor…

Cited by 14PDFcodeScholar