← Search

Rishabh Jain

12 accepted papers

2026

Graph Attention-Guided Search for Dense Multi-Agent Pathfinding

AAAI 2026technical

Finding near-optimal solutions for dense multi-agent pathfinding (MAPF) problems in real-time remains challenging even for state-of-the-art planners. To this end, we develop a hybrid framework that integrates a learned heuristic derived from MAGAT, a neural MAPF policy with a graph attention scheme,

Cited by 0SourcePDFScholar
2026

Pairwise is Not Enough: Hypergraph Neural Networks for Multi-Agent Pathfinding

ICLR 2026poster

Multi-Agent Path Finding (MAPF) is a representative multi-agent coordination problem, where multiple agents are required to navigate to their respective goals without collisions. Solving MAPF optimally is known to be NP-hard, leading to the adoption of learning-based approaches to alleviate the onli…

Cited by 0SourcecodeScholar
2025

AesthetiQ: Enhancing Graphic Layout Design via Aesthetic-Aware Preference Alignment of Multi-modal Large Language Models

CVPR 2025poster

Visual layouts are essential in graphic design fields such as advertising, posters, and web interfaces. The application of generative models for content-aware layout generation has recently gained traction. However, these models fail to understand the contextual aesthetic requirements of layout desi…

Cited by 0SourcePDFScholar
2025

Princeton365: A Diverse Dataset with Accurate Camera Pose

ICCV 2025poster

We introduce Princeton365, a large-scale diverse dataset of 365 videos with accurate camera pose. Our dataset bridges the gap between accuracy and data diversity in current SLAM benchmarks by introducing a novel ground truth collection framework that leverages calibration boards and a 360 camera. We…

Cited by 0SourcePDFScholar
2023

Parameter Efficient Local Implicit Image Function Network for Face Segmentation

CVPR 2023poster

Face parsing is defined as the per-pixel labeling of images containing human faces. The labels are defined to identify key facial regions like eyes, lips, nose, hair, etc. In this work, we make use of the structural consistency of the human face to propose a lightweight face-parsing method using a L…

Cited by 10SourcePDFScholar
2023

UMFuse: Unified Multi View Fusion for Human Editing Applications

ICCV 2023poster

Numerous pose-guided human editing methods have been explored by the vision community due to their extensive practical applications. However, most of these methods still use an image-to-image formulation in which a single image is given as input to produce an edited image as output. This objective b…

Cited by 1PDFScholar
2023

VGFlow: Visibility Guided Flow Network for Human Reposing

CVPR 2023poster

The task of human reposing involves generating a realistic image of a model standing in an arbitrary conceivable pose. There are multiple difficulties in generating perceptually accurate images and existing methods suffers from limitations in preserving texture, maintaining pattern coherence, respec…

Cited by 7SourcePDFScholar
2022

Extending Logic Explained Networks to Text Classification

EMNLP 2022main

Recently, Logic Explained Networks (LENs) have been proposed as explainable-by-design neural models providing logic explanations for their predictions.However, these models have only been applied to vision and tabular data, and they mostly favour the generation of global explanations, while local on…

Cited by 15SourcePDFScholar
2021

ZFlow: Gated Appearance Flow-Based Virtual Try-On With 3D Priors

ICCV 2021poster

Image-based virtual try-on involves synthesizing perceptually convincing images of a model wearing a particular garment and has garnered significant research interest due to its immense practical applicability. Recent methods involve a two-stage process: i) warping of the garment to align with the m…

Cited by 77PDFScholar
2020

Dialog without Dialog Data: Learning Visual Dialog Agents from VQA Data

NeurIPS 2020poster

Can we develop visually grounded dialog agents that can efficiently adapt to new tasks without forgetting how to talk to people? Such agents could leverage a larger variety of existing data to generalize to a new task, minimizing expensive data collection and annotation. In this work, we study a set…

2019

nocaps: novel object captioning at scale

ICCV 2019poster

Image captioning models have achieved impressive results on datasets containing limited visual concepts and large amounts of paired image-caption training data. However, if these models are to ever function in the wild, a much larger variety of visual concepts must be learned, ideally from less supe…

Cited by 420PDFcodeScholar