← Search

Yanbin Hao

22 accepted papers

2026

Accelerating Controllable Generation via Hybrid-grained Cache

AAAI 2026technical

Controllable generative models have been widely used to improve the realism of synthetic visual content. However, such models must handle control conditions and content generation computational requirements, resulting in generally low generation efficiency. To address this issue, we propose a Hybrid

Cited by 0SourcePDFScholar
2026

Glimpse: Geometry Learning of Multi-scale Structural Priors for 3D Pose Estimation

ICML 2026poster

Monocular 3D human pose estimation is fundamentally challenged by severe occlusion and inherent depth ambiguity. To address this, we propose Glimpse, a framework that learns robust 3D poses by explicitly modeling anatomical geometry from a single image. We recast the problem as geometry learning of …

Cited by 0SourceScholar
2026

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning

ICASSP 2026poster

Multimodal Large Language Models (MLLMs) perform well in single-image visual grounding but struggle with real-world tasks that demand cross-image reasoning and multi-modal instructions. To address this, we adopt a reinforcement learning (RL) based post-training strategy for MLLMs in multi-image grou…

Cited by 0SourcePDFScholar
2026

Res-Bench: Benchmarking the Robustness of Multimodal Large Language Models to Dynamic Resolution Input

AAAI 2026technical

Multimodal Large Language Models (MLLMs) increasingly support dynamic image resolutions. However, current evaluation paradigms primarily assess semantic performance, overlooking the critical question of resolution robustness - whether performance remains stable across varying input resolutions. To a

Cited by 0SourcePDFScholar
2026

SNS-Grasp: Semantic-guided Noise Scaling for Grasp Generation

AAAI 2026technical

While diffusion models show promise for intent-based grasp generation, their isotropic noise schedules struggle with joint-specific sensitivity and task-aware variability. This limitation leads to grasps with suboptimal semantic alignment or physical feasibility. To address this challenge, we propos

Cited by 0SourcePDFScholar
2026

SPEED: Scalable, Precise, and Efficient Concept Erasure for Diffusion Models

ICLR 2026poster

Erasing concepts from large-scale text-to-image (T2I) diffusion models has become increasingly crucial due to the growing concerns over copyright infringement, offensive content, and privacy violations. In scalable applications, fine-tuning-based methods are time-consuming to precisely erase multipl…

Cited by 0SourcecodeScholar
2025

A Sanity Check for AI-generated Image Detection

ICLR 2025poster

With the rapid development of generative models, discerning AI-generated content has evoked increasing attention from both industry and academia. In this paper, we conduct a sanity check on whether the task of AI-generated image detection has been solved. To start with, we present Chameleon dataset,…

2025

Hand1000: Generating Realistic Hands from Text with Only 1,000 Images

AAAI 2025technical

Text-to-image generation models have achieved remarkable advancements in recent years, aiming to produce realistic images from textual descriptions. However, these models often struggle with generating anatomically accurate representations of human hands. The resulting images frequently exhibit issu…

Cited by 4SourcePDFScholar
2025

Improving Open-vocabulary Video Visual Relation Detection with Decomposed Prompt Learning and Relation Adjustment

ICASSP 2025accepted

Open-vocabulary video visual relation detection (VidVRD) expands the scope of detecting object relations in videos to include unseen categories. It marks considerable advancement in recognizing novel relations solely by training on a base set, thus extending the frontiers of automated video understa…

Cited by 0SourceScholar
2025

Mixture of Multimodal Adapters for Sentiment Analysis

NAACL 2025long

Pre-trained language model (PLM) have achieved great success in text sentiment analysis. However, in practical applications, sentiment is not only conveyed through language but also hidden in other modalities. Therefore, multimodal sentiment analysis (MSA) has attracted increasing research interest.…

2025

Precise, Fast, and Low-cost Concept Erasure in Value Space: Orthogonal Complement Matters

CVPR 2025poster

The success of text-to-image generation enabled by diffusion models has imposed an urgent need to erase unwanted concepts, e.g., copyrighted, offensive, and unsafe ones, from the pre-trained models in a precise, timely, and low-cost manner. The twofold demand of concept erasure requires a precise re…

2025

RAGG: Retrieval-Augmented Grasp Generation Model

AAAI 2025technical

Intent-based grasp generation inherently involves challenges such as manipulation ambiguity and modality gaps. To address these, we propose a novel Retrieval-Augmented Grasp Generation model (RAGG). Our key insight is that when humans manipulate new objects, they initially mimic the interaction patt…

Cited by 0SourcePDFScholar
2024

3D-GOI: 3D GAN Omni-Inversion for Multifaceted and Multi-object Editing

ECCV 2024poster

"The current GAN inversion methods typically can only edit the appearance and shape of a single object and background while overlooking spatial information. In this work, we propose a 3D editing framework, to enable multifaceted editing of affine information (scale, translation, and rotation) on mul…

2024

Boosting Few-Shot Learning via Attentive Feature Regularization

AAAI 2024technical

Few-shot learning (FSL) based on manifold regularization aims to improve the recognition capacity of novel objects with limited training samples by mixing two samples from different categories with a blending factor. However, this mixing operation weakens the feature representation due to the linear…

Cited by 11SourcePDFScholar
2024

Enhance Image Classification via Inter-Class Image Mixup with Diffusion Model

CVPR 2024poster

Text-to-image (T2I) generative models have recently emerged as a powerful tool enabling the creation of photo-realistic images and giving rise to a multitude of applications. However the effective integration of T2I models into fundamental image classification tasks remains an open question. A preva…

2024

Enhancing Zero-Shot Vision Models by Label-Free Prompt Distribution Learning and Bias Correcting

NeurIPS 2024spotlight

Vision-language models, such as CLIP, have shown impressive generalization capacities when using appropriate text descriptions. While optimizing prompts on downstream labeled data has proven effective in improving performance, these methods entail labor costs for annotations and are limited by their…

2024

PointTFA: Training-Free Clustering Adaption for Large 3D Point Cloud Models

IJCAI 2024poster

The success of contrastive learning models like CLIP, known for aligning 2D image-text pairs, has inspired the development of triplet alignment for Large 3D Point Cloud Models (3D-PCM). Examples like ULIP integrate images, text, and point clouds into a unified semantic space. However, despite showin…

2023

Bi-Directional Distribution Alignment for Transductive Zero-Shot Learning

CVPR 2023poster

It is well-known that zero-shot learning (ZSL) can suffer severely from the problem of domain shift, where the true and learned data distributions for the unseen classes do not match. Although transductive ZSL (TZSL) attempts to improve this by allowing the use of unlabelled examples from the unseen…

2021

Aggregated Multi-GANs for Controlled 3D Human Motion Prediction

AAAI 2021technical

Human motion prediction from historical pose sequence is at the core of many applications in machine intelligence. However, in current state-of-the-art methods, the predicted future motion is confined within the same activity. One can neither generate predictions that differ from the current activit…

2021

Motion Prediction Using Trajectory Cues

ICCV 2021poster

Predicting human motion from a historical pose sequence is at the core of many applications in computer vision. Current state-of-the-art methods concentrate on learning motion contexts in the pose space, however, the high dimensionality and complex nature of human pose invoke inherent difficulties i…

Cited by 64PDFcodeScholar
2019

R2GAN: Cross-Modal Recipe Retrieval With Generative Adversarial Network

CVPR 2019poster

Representing procedure text such as recipe for crossmodal retrieval is inherently a difficult problem, not mentioning to generate image from recipe for visualization. This paper studies a new version of GAN, named Recipe Retrieval Generative Adversarial Network (R2GAN), to explore the feasibility of…

Cited by 159PDFcodeScholar