← Search

Han Yang

36 accepted papers

2026

Authority Backdoor: A Certifiable Backdoor Mechanism for Authoring DNNs

AAAI 2026technical

Deep Neural Networks (DNNs), as valuable intellectual property, face unauthorized use. Existing protections, such as digital watermarking, are largely passive; they provide only post-hoc ownership verification and cannot actively prevent the illicit use of a stolen model. This work proposes a proact

Cited by 0SourcePDFScholar
2026

Fine-Grained Classification for Depth Estimation From Monocular Microscopy for Robotic Micromanipulation of Motile Cells

RA-L 2026

Manipulation of motile cells is crucial for biological research and clinical applications. However, obtaining Z-axis visual feedback under monocular microscopy remains a challenge for robotic micromanipulation. Traditional depth-from-focus and depth-from-defocus methods fail to handle motile cells d

Cited by 0SourceScholar
2026

PhyScensis: Physics-Augmented LLM Agents for Complex Physical Scene Arrangement

ICLR 2026poster

Automatically generating interactive 3D environments is crucial for scaling up robotic data collection in simulation. While prior work has primarily focused on 3D asset placement, it often overlooks the physical relationships between objects (e.g., contact, support, balance, and containment), which…

Cited by 0SourceScholar
2026

Robotic Cell Manipulation at the Solid-Liquid Interface for Cryopreservation

ICRA 2026poster

Automating cell manipulation at a solid-liquid interface is a critical challenge for biomedical applications such as embryo cryopreservation. Unlike manipulation in a full liquid medium, the cell-substrate contact creates a significant static friction force that is not readily measurable with curren…

Cited by 0Scholar
2025

$\text{S}^2$Q-VDiT: Accurate Quantized Video Diffusion Transformer with Salient Data and Sparse Token Distillation

NeurIPS 2025poster

Diffusion transformers have emerged as the mainstream paradigm for video generation models. However, the use of up to billions of parameters incurs significant computational costs. Quantization offers a promising solution by reducing memory usage and accelerating inference. Nonetheless, we observe t…

Cited by 0SourcecodeScholar
2025

3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning

CVPR 2025poster

Constructing compact and informative 3D scene representations is essential for effective embodied exploration and reasoning, especially in complex environments over extended periods. Existing representations, such as object-centric 3D scene graphs, oversimplify spatial relationships by modeling scen…

Cited by 1SourcePDFScholar
2025

Continuous Convolution for Automated Measurement of Sperm Flagella

ICRA 2025

Quantifying sperm flagellar beating behavior (e.g., beating amplitude, frequency, and wavelength) plays a crucial role in biological research, clinical diagnostics, and the design of sperm-inspired microrobots. However, existing computational methods struggle to accurately and efficiently analyze th

Cited by 0SourcecodeScholar
2025

E2Former: An Efficient and Equivariant Transformer with Linear-Scaling Tensor Products

NeurIPS 2025spotlight

Equivariant Graph Neural Networks (EGNNs) have demonstrated significant success in modeling microscale systems, including those in chemistry, biology and materials science. However, EGNNs face substantial computational challenges due to the high cost of constructing edge features via spherical tenso…

Cited by 0SourceScholar
2025

Efficient and Scalable Density Functional Theory Hamiltonian Prediction through Adaptive Sparsity

ICML 2025poster

Hamiltonian matrix prediction is pivotal in computational chemistry, serving as the foundation for determining a wide range of molecular properties. While SE(3) equivariant graph neural networks have achieved remarkable success in this domain, their substantial computational cost—driven by high-orde…

2025

Enhancing the Scalability and Applicability of Kohn-Sham Hamiltonians for Molecular Systems

ICLR 2025spotlight

Density Functional Theory (DFT) is a pivotal method within quantum chemistry and materials science, with its core involving the construction and solution of the Kohn-Sham Hamiltonian. Despite its importance, the application of DFT is frequently limited by the substantial computational resources requ…

Cited by 0SourcePDFScholar
2025

Feature Disentangling Dual-stream Network for User Bias Alleviation in Social Media Prediction

ICASSP 2025accepted

Social media popularity prediction is increasingly crucial for optimizing user engagement and guiding content recommendation systems. However, existing methods suffer from an excessive reliance on user information, which disproportionately influences predictions and leads to the neglect of content d…

Cited by 0SourceScholar
2025

HSRDiff: A Hierarchical Self-Regulation Diffusion Model for Stochastic Semantic Segmentation

AAAI 2025technical

In safety-critical domains such as medical diagnostics and autonomous driving, single-image evidence is sometimes insufficient to reflect the inherent ambiguity of vision problems. Therefore, multiple plausible assumptions that match the image semantics may be needed to reflect the actual distributi…

2025

Multi-Teacher Knowledge Distillation with Reinforcement Learning for Visual Recognition

AAAI 2025technical

Multi-teacher Knowledge Distillation (KD) transfers diverse knowledge from a teacher pool to a student network. The core problem of multi-teacher KD is how to balance distillation strengths among various teachers. Most existing methods often develop weighting strategies from an individual perspectiv…

2025

Multi-party Collaborative Attention Control for Image Customization

CVPR 2025poster

The rapid development of diffusion models has fueled a growing demand for customized image generation. However, current customization methods face several limitations: 1) typically accept either image or text conditions alone; 2) customization in complex visual scenarios often leads to subject leaka…

2025

SLC${2}$-SLAM: Semantic-Guided Loop Closure Using Shared Latent Code for NeRF SLAM

RA-L 2025

Targeting the notorious cumulative drift errors in NeRF SLAM, we propose a Semantic-guided Loop Closure using Shared Latent Code, dubbed SLC<inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$^{2}$</tex-math></inline-f

Cited by 6SourceScholar
2025

UniMuMo: Unified Text, Music, and Motion Generation

AAAI 2025technical

We introduce UniMuMo, a unified multimodal model capable of taking arbitrary text, music, and motion data as input conditions to generate outputs across all three modalities. To address the lack of time-synchronized data, we align unpaired music and motion data based on rhythmic patterns to leverage…

2024

CLIP-KD: An Empirical Study of CLIP Model Distillation

CVPR 2024poster

Contrastive Language-Image Pre-training (CLIP) has become a promising language-supervised visual pre-training framework. This paper aims to distill small CLIP models supervised by a large teacher CLIP model. We propose several distillation strategies including relation feature gradient and contrasti…

2024

GNNCert: Deterministic Certification of Graph Neural Networks against Adversarial Perturbations

ICLR 2024oral

Graph classification, which aims to predict a label for a graph, has many real-world applications such as malware detection, fraud detection, and healthcare. However, many studies show an attacker could carefully perturb the structure and/or node features in a graph such that a graph classifier misc…

Cited by 10SourcePDFScholar
2024

Long-Short-Range Message-Passing: A Physics-Informed Framework to Capture Non-Local Interaction for Scalable Molecular Dynamics Simulation

ICLR 2024poster

Computational simulation of chemical and biological systems using *ab initio* molecular dynamics has been a challenge over decades. Researchers have attempted to address the problem with machine learning and fragmentation-based methods. However, the two approaches fail to give a satisfactory descrip…

2024

Training-free Multi-objective Diffusion Model for 3D Molecule Generation

ICLR 2024poster

Searching for novel and diverse molecular candidates is a critical undertaking in drug and material discovery. Existing approaches have successfully adapted the diffusion model, the most effective generative model in image generation, to create 1D SMILES strings, 2D chemical graphs, or 3D molecular…

Cited by 10SourcePDFScholar
2024

Weakly-Supervised Depth Completion during Robotic Micromanipulation from a Monocular Microscopic Image

ICRA 2024poster

Obtaining three-dimensional information, especially the z-axis depth information, is crucial for robotic micromanipulation. Due to the unavailability of depth sensors such as lidars in micromanipulation setups, traditional depth acquisition methods such as depth from focus or depth from defocus dire…

Cited by 0SourceScholar
2023

Pareto Invariant Risk Minimization: Towards Mitigating the Optimization Dilemma in Out-of-Distribution Generalization

ICLR 2023poster

Recently, there has been a growing surge of interest in enabling machine learning systems to generalize well to Out-of-Distribution (OOD) data. Most efforts are devoted to advancing optimization objectives that regularize models to capture the underlying invariance; however, there often are compromi…

2023

SILT: Shadow-Aware Iterative Label Tuning for Learning to Detect Shadows from Noisy Labels

ICCV 2023poster

Existing shadow detection datasets often contain missing or mislabeled shadows, which can hinder the performance of deep learning models trained directly on such data. To address this issue, we propose SILT, the Shadow-aware Iterative Label Tuning framework, which explicitly considers noise in shado…

Cited by 19PDFcodeScholar
2022

A Personable Robot: Meta-Analysis of Robot Personality and Human Acceptance

RA-L 2022

For robots to be of use to humans, they first must be accepted. One important variable that might impact this acceptance is a robot’s personality. To date, results of studies examining robot personality have produced mixed results. One method of making sense of these results is meta-analysis. Theref

Cited by 22SourceScholar
2022

Exact Shape Correspondence via 2D graph convolution

NeurIPS 2022accept

For exact 3D shape correspondence (matching or alignment), i.e., the task of matching each point on a shape to its exact corresponding point on the other shape (or to be more specific, matching at geodesic error 0), most existing methods do not perform well due to two main problems. First, on nearly…

Cited by 6SourcePDFScholar
2022

Learning Causally Invariant Representations for Out-of-Distribution Generalization on Graphs

NeurIPS 2022accept

Despite recent success in using the invariance principle for out-of-distribution (OOD) generalization on Euclidean data (e.g., images), studies on graph data are still limited. Different from images, the complex nature of graphs poses unique challenges to adopting the invariance principle. In partic…

2022

Understanding and Improving Graph Injection Attack by Promoting Unnoticeability

ICLR 2022poster

Recently Graph Injection Attack (GIA) emerges as a practical attack scenario on Graph Neural Networks (GNNs), where the adversary can merely inject few malicious nodes instead of modifying existing nodes or edges, i.e., Graph Modification Attack (GMA). Although GIA has achieved promising results, li…

2021

Disentangled Cycle Consistency for Highly-Realistic Virtual Try-On

CVPR 2021poster

Image virtual try-on replaces the clothes on a person image with a desired in-shop clothes image. It is challenging because the person and the in-shop clothes are unpaired. Existing methods formulate virtual try-on as either in-painting or cycle consistency. Both of these two formulations encourage…

Cited by 132PDFcodeScholar
2020

CPR-GCN: Conditional Partial-Residual Graph Convolutional Network in Automated Anatomical Labeling of Coronary Arteries

CVPR 2020oral

Automated anatomical labeling plays a vital role in coronary artery disease diagnosing procedure. The main challenge in this problem is the large individual variability inherited in human anatomy. Existing methods usually rely on the position information and the prior knowledge of the topology of th…

Cited by 56PDFScholar
2020

Towards Photo-Realistic Virtual Try-On by Adaptively Generating-Preserving Image Content

CVPR 2020poster

Image visual try-on aims at transferring a target clothes image onto a reference person, and has become a hot topic in recent years. Prior arts usually focus on preserving the character of a clothes image (e.g. texture, logo, embroidery) when warping it to arbitrary human pose. However, it remains a…

Cited by 337PDFScholar