← Search

Hua Yang

22 accepted papers

2026

Adapter Shield: A Unified Framework with Built-in Authentication for Preventing Unauthorized Zero-Shot Image-to-Image Generation

CVPR 2026

With the rapid progress in diffusion models, image synthesis has advanced to the stage of zero-shot image-to-image generation, where high-fidelity replication of facial identities or artistic styles can be achieved using just one portrait or artwork, without modifying any model weights. Although the

Cited by 0SourceScholar
2026

Lightweight Adaptive Topological Layout and Semantic Mapping in Vision-and-Language Navigation on Websites

AAAI 2026technical

Vision-and-Language navigation on websites requires agents to navigate target webpages and answer questions based on human instructions. Current web agents primarily leverage Large Language Models (LLMs) for semantic understanding and reasoning, but still suffer from limited navigation performance a

Cited by 0SourcePDFScholar
2026

MM-HELIX: Boosting Multimodal Long-Chain Reflective Reasoning with Holistic Platform and Adaptive Hybrid Policy Optimization

ICLR 2026poster

While current Multimodal Large Language Models (MLLMs) have demonstrated proficiency in reasoning tasks such as mathematics and logic, their capacity for long-chain reflective reasoning, a prerequisite for solving complex real-world problems, remains largely underexplored. In this work, we first co…

Cited by 0SourcecodeScholar
2026

OneLIP: Unlocking and Improving Long-Text Representations of CLIP via One-Stage Adaptation

AAAI 2026technical

Contrastive Language-Image Pretraining (CLIP) has demonstrated impressive generalization on vision-language tasks by aligning images and short texts. However, its inherent 77-token length limits the capacity of capturing complex semantics in long captions. Existing long-text adaptations for CLIP typ

Cited by 0SourcePDFScholar
2026

P-GenRM: Personalized Generative Reward Model with Test-time User-based Scaling

ICLR 2026oral

Personalized alignment of large language models seeks to adapt responses to individual user preferences, typically via reinforcement learning. A key challenge is obtaining accurate, user-specific reward signals in open-ended scenarios. Existing personalized reward models face two persistent limitati…

Cited by 0SourcecodeScholar
2025

3D Vision-tactile Reconstruction from Infrared and Visible Images for Robotic Fine-grained Tactile Perception

IROS 2025

To achieve human-like haptic perception in anthropomorphic grippers, the compliant sensing surfaces of vision tactile sensor (VTS) must evolve from conventional planar configurations to biomimetically curved topographies with continuous surface gradients. However, planar VTSs have challenges when ex

Cited by 1SourceScholar
2025

Discovering Clone Negatives via Adaptive Contrastive Learning for Image-Text Matching

ICLR 2025poster

In this paper, we identify a common yet challenging issue in image-text matching, i.e., clone negatives: negative image-text pairs that semantically resemble positive pairs, leading to ambiguous and sub-optimal matching outcomes. To tackle this issue, we propose Adaptive Contrastive Learning (AdaCL)…

Cited by 1SourcePDFScholar
2025

Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing

NeurIPS 2025oral

Large Multi-modality Models (LMMs) have made significant progress in visual understanding and generation, but they still face challenges in General Visual Editing, particularly in following complex instructions, preserving appearance consistency, and supporting flexible input formats. To study this…

Cited by 0SourcecodeScholar
2025

Hydrodynamics Regularization in Reinforcement Learning for Navigating Crowded Scenarios

RA-L 2025

The navigation task in dense crowds is a key research problem in real-world scenarios. It requires an agent to avoid collisions in dynamic environments and reach the agent's destination, ensuring high accuracy and efficiency in its decisions. Existing methods typically treat pedestrians as rigid bod

Cited by 0SourceScholar
2025

MTIL: Encoding Full History With Mamba for Temporal Imitation Learning

RA-L 2025

Standard imitation learning (IL) methods have achieved considerable success in robotics, yet often rely on the Markov assumption, which falters in long-horizon tasks where history is crucial for resolving perceptual ambiguity. This limitation stems not only from a conceptual gap but also from a fund

Cited by 6SourcecodeScholar
2025

OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference

ACL 2025long

Recent advancements in open-source multi-modal large language models (MLLMs) have primarily focused on enhancing foundational capabilities, leaving a significant gap in human preference alignment. This paper introduces OmniAlign-V, a comprehensive dataset of 200K high-quality training samples featur…

2024

FC-GNN: Recovering Reliable and Accurate Correspondences from Interferences

CVPR 2024poster

Finding correspondences between images is essential for many computer vision tasks and sparse matching pipelines have been popular for decades. However matching noise within and between images along with inconsistent keypoint detection frequently degrades the matching performance. We review these pr…

2024

GelRoller: A Rolling Vision-based Tactile Sensor for Large Surface Reconstruction Using Self-Supervised Photometric Stereo Method

ICRA 2024poster

Accurate perception of the surrounding environment stands as a primary objective for robots. Through tactile interaction, vision-based tactile sensors provide the capability to capture high-resolution and multi-modal surface information of objects, thereby facilitating robots in achieving more dexte…

Cited by 2SourcecodeScholar
2023

Co-training with High-Confidence Pseudo Labels for Semi-supervised Medical Image Segmentation

IJCAI 2023poster

Consistency regularization and pseudo labeling-based semi-supervised methods perform co-training using the pseudo labels from multi-view inputs. However, such co-training models tend to converge early to a consensus, degenerating to the self-training ones, and produce low-confidence pseudo labels fr…

2022

Dynamic Spatial Propagation Network for Depth Completion

AAAI 2022technical

Image-guided depth completion aims to generate dense depth maps with sparse depth measurements and corresponding RGB images. Currently, spatial propagation networks (SPNs) are the most popular affinity-based methods in depth completion, but they still suffer from the representation limitation of the…

Cited by 137SourcePDFScholar
2018

Online Multi-Object Tracking with Dual Matching Attention Networks

ECCV 2018poster

In this paper, we propose an online Multi-Object Tracking (MOT) approach which integrates the merits of single object tracking and data association methods in a unified framework to handle noisy detections and frequent interactions between targets. Specifically, for applying single object tracking i…

Cited by 455SourcePDFScholar
2016

Joint instance and feature importance re-weighting for person reidentification

ICASSP 2016accepted

Person reidentification refers to the task of recognizing the same person under different non-overlapping camera views. Presently, person reidentification based on metric learning is proved to be effective among various techniques, which exploits the labeled data to learn a subspace that maximizes t…

Cited by 0SourceScholar