← Search

Binod Bhattarai

11 accepted papers

2026

HandVQA: Diagnosing and Improving Fine-Grained Spatial Reasoning about Hands in Vision-Language Models

CVPR 2026

Understanding the fine-grained articulation of human hands is critical in high-stakes settings such as robot-assisted surgery, chip manufacturing, and AR/VR-based human-AI interaction. Despite achieving near-human performance on general vision-language benchmarks, current vision-language models (VLM

Cited by 0SourceScholar
2024

Prompt Augmentation for Self-supervised Text-guided Image Manipulation

CVPR 2024poster

Text-guided image editing finds applications in various creative and practical fields. While recent studies in image generation have advanced the field they often struggle with the dual challenges of coherent image transformation and context preservation. In response our work introduces prompt augme…

Cited by 2SourcePDFScholar
2023

Joint Training of Hierarchical GANs and Semantic Segmentation for Expression Translation

ICASSP 2023accepted

Manipulating images by changing only specific attributes has been a long-standing research problem. Existing methods that rely solely on a global generator often suffer from changing unwanted attributes along with the desired attributes. Although hierarchical networks consisting of global and local…

Cited by 0SourceScholar
2023

Why Is the Winner the Best?

CVPR 2023poster

International benchmarking competitions have become fundamental for the comparative performance assessment of image analysis methods. However, little attention has been given to investigating what can be learnt from these competitions. Do they really generate scientific progress? What are common and…

Cited by 29SourcePDFScholar
2020

Auglabel: Exploiting Word Representations to Augment Labels for Face Attribute Classification

ICASSP 2020accepted

Augmenting data in image space (eg. flipping, cropping etc) and activation space (eg. dropout) are being widely used to regularise deep neural networks and have been successfully applied on several computer vision tasks. Unlike previous works, which are mostly focused on doing augmentation in the af…

Cited by 0SourceScholar
2018

Semi-supervised Adversarial Learning to Generate Photorealistic Face Images of New Identities from 3D Morphable Model

ECCV 2018poster

We propose a novel end-to-end semi-supervised adversarial framework to generate photorealistic face images of new identities with a wide range of expressions, poses, and illuminations conditioned by synthetic images sampled from a 3D morphable model. Previous adversarial style-transfer methods eithe…

2016

A joint learning approach for cross domain age estimation

ICASSP 2016accepted

We propose a novel joint learning method for cross domain age estimation, a domain adaptation problem. The proposed method learns a low dimensional projection along with a re-gressor, in the projection space, in a joint framework. The projection aligns the features from two different domains, i.e. s…

Cited by 0SourceScholar
2016

CP-mtML: Coupled Projection Multi-Task Metric Learning for Large Scale Face Retrieval

CVPR 2016poster

We propose a novel Coupled Projection multi-task Met- ric Learning (CP-mtML) method for large scale face re- trieval. In contrast to previous works which were limited to low dimensional features and small datasets, the proposed method scales to large datasets with high dimensional face descriptors.…

Cited by 61PDFcodeScholar