← Search

Zhen Han

28 accepted papers

2026

Eliminate Distance Differences Induced by Backdoor Attacks: Layer-Selective Training and Clipping to Mask Backdoor Models

CVPR 2026

Federated learning (FL) enables a central server to collaboratively train a global model with multiple clients while preserving data privacy. However, the distributed nature of FL makes the paradigm vulnerable to backdoor attacks, as proved by numerous recent studies. Although existing studies impro

Cited by 0SourceScholar
2026

FRBAT: Conditionally-Visible Physical Backdoor Attack via Fluorescence

AAAI 2026technical

Deep neural networks are increasingly vulnerable to physically deployable backdoor attacks, which manipulate real-world objects to induce targeted model failures. However, current physical backdoor attacks predominantly rely on perpetually visible triggers appended to target objects. These methods i

Cited by 0SourcePDFScholar
2026

Rethinking Surgical Smoke: A Smoke-Type-Aware Laparoscopic Video Desmoking Method and Dataset

AAAI 2026technical

Electrocautery or lasers will inevitably generate surgical smoke, which hinders the visual guidance of laparoscopic videos for surgical procedures. The surgical smoke can be classified into different types based on its motion patterns, leading to distinctive spatio-temporal characteristics across sm

Cited by 0SourcePDFScholar
2026

Tutor-Student Reinforcement Learning: A Dynamic Curriculum for Robust Deepfake Detection

CVPR 2026

Standard supervised training for deepfake detection treats all samples with uniform importance, which can be suboptimal for learning robust and generalizable features. In this work, we propose a novel Tutor-Student Reinforcement Learning (TSRL) framework to dynamically optimize the training curricul

Cited by 0SourcecodeScholar
2026

Wan-Weaver: Interleaved Multi-modal Generation via Decoupled Training

CVPR 2026

Recent unified models have made unprecedented progress in both understanding and generation. However, while most of them accept multi-modal inputs, they typically produce only single-modality outputs. This challenge of producing interleaved content is mainly due to training data scarcity and the dif

Cited by 0SourceScholar
2026

WebArbiter: A Generative Reasoning Process Reward Model for Web Agents

ICLR 2026poster

Web agents hold great potential for automating complex computer tasks, yet their interactions involve long horizons, multi-step decisions, and actions that can be irreversible. In such settings, outcome-based supervision is sparse and delayed, often rewarding incorrect trajectories and failing to su…

Cited by 0SourceScholar
2025

ACE: All-round Creator and Editor Following Instructions via Diffusion Transformer

ICLR 2025poster

Diffusion models have emerged as a powerful generative technology and have been found to be applicable in various scenarios. Most existing foundational diffusion models are primarily designed for text-guided visual generation and do not support multi-modal conditions, which are essential for many vi…

Cited by 10SourcePDFScholar
2025

BYOKG-RAG: Multi-Strategy Graph Retrieval for Knowledge Graph Question Answering

EMNLP 2025

Knowledge graph question answering (KGQA) presents significant challenges due to the structural and semantic variations across input graphs. Existing works rely on Large Language Model (LLM) agents for graph traversal and retrieval; an approach that is sensitive to traversal initialization, as it is

2025

HybGRAG: Hybrid Retrieval-Augmented Generation on Textual and Relational Knowledge Bases

ACL 2025long

Given a semi-structured knowledge base (SKB), where text documents are interconnected by relations, how can we effectively retrieve relevant information to answer user questions?Retrieval-Augmented Generation (RAG) retrieves documents to assist large language models (LLMs) in question answering; whi…

2025

ICE-Bench: A Unified and Comprehensive Benchmark for Image Creating and Editing

ICCV 2025poster

Image generation has witnessed significant advancements in the past few years. However, evaluating the performance of image generation models remains a formidable challenge. In this paper, we propose ICE-Bench, a unified and comprehensive benchmark designed to rigorously assess image generation mode…

2025

Link-based Contrastive Learning for One-Shot Unsupervised Domain Adaptation

CVPR 2025poster

Unsupervised domain adaptation (UDA) aims to learn discriminative features from a labeled source domain by supervised learning and to transfer the knowledge to an unlabeled target domain via distribution alignment. However, in some real-world scenarios, e.g., public safety or access control, it's di…

Cited by 0SourcePDFScholar
2025

VACE: All-in-One Video Creation and Editing

ICCV 2025poster

Diffusion Transformer has demonstrated powerful capability and scalability in generating high-quality images and videos. Further pursuing the unification of generation and editing tasks has yielded significant progress in the domain of image content creation. However, due to the intrinsic demands fo…

Cited by 0SourcePDFScholar
2025

WebPilot: A Versatile and Autonomous Multi-Agent System for Web Task Execution with Strategic Exploration

AAAI 2025technical

LLM-based autonomous agents often fail to execute complex web tasks that require dynamic interaction, largely due to the inherent uncertainty and complexity of these environments. Existing LLM-based web agents typically rely on rigid, expert-designed policies specific to certain states and actions,…

Cited by 21SourcePDFScholar
2024

SCEdit: Efficient and Controllable Image Diffusion Generation via Skip Connection Editing

CVPR 2024highlight

Image diffusion models have been utilized in various tasks such as text-to-image generation and controllable image synthesis. Recent research has introduced tuning methods that make subtle adjustments to the original models yielding promising results in specific adaptations of foundational generativ…

2024

Torque Ripple Reduction in Quasi-Direct Drive Motors Through Angle-Based Repetitive Learning Observer and Model Predictive Torque Controller

IROS 2024poster

Torque ripple reduction in quasi-direct drive (QDD) motors is crucial in their robotic applications for dynamic locomotion and dexterous manipulation. In this paper, we present a novel approach for reducing torque ripples of QDD motors, which integrates an angle-based repetitive learning observer (A…

Cited by 0SourceScholar
2024

Visual Question Decomposition on Multimodal Large Language Models

EMNLP 2024finding

Question decomposition has emerged as an effective strategy for prompting Large Language Models (LLMs) to answer complex questions. However, while existing methods primarily focus on unimodal language models, the question decomposition capability of Multimodal Large Language Models (MLLMs) has yet t…

Cited by 0SourcePDFScholar
2023

Benchmarking Robustness of Adaptation Methods on Pre-trained Vision-Language Models

NeurIPS 2023poster

Various adaptation methods, such as LoRA, prompts, and adapters, have been proposed to enhance the performance of pre-trained vision-language models in specific domains. As test samples in real-world applications usually differ from adaptation data, the robustness of these adaptation methods against…

2023

ECOLA: Enhancing Temporal Knowledge Embeddings with Contextualized Language Representations

ACL 2023findings

Since conventional knowledge embedding models cannot take full advantage of the abundant textual information, there have been extensive research efforts in enhancing knowledge embedding using texts. However, existing enhancement approaches cannot apply to temporal knowledge graphs (tKGs), which cont…

2022

Multi-Hop Open-Domain Question Answering over Structured and Unstructured Knowledge

NAACL 2022findings

Open-domain question answering systems need to answer question of our interests with structured and unstructured information. However, existing approaches only select one source to generate answer or only conduct reasoning on structured information. In this paper, we pro- pose a Document-Entity Hete…

Cited by 23SourcePDFScholar
2021

Explainable Subgraph Reasoning for Forecasting on Temporal Knowledge Graphs

ICLR 2021poster

Modeling time-evolving knowledge graphs (KGs) has recently gained increasing interest. Here, graph representation learning has become the dominant paradigm for link prediction on temporal KGs. However, the embedding-based approaches largely operate in a black-box fashion, lacking the ability to inte…

Cited by 229SourcePDFScholar
2021

Learning Neural Ordinary Equations for Forecasting Future Links on Temporal Knowledge Graphs

EMNLP 2021main

There has been an increasing interest in inferring future links on temporal knowledge graphs (KG). While links on temporal KGs vary continuously over time, the existing approaches model the temporal KGs in discrete state spaces. To this end, we propose a novel continuum model by extending the idea o…

2021

Time-dependent Entity Embedding is not All You Need: A Re-evaluation of Temporal Knowledge Graph Completion Models under a Unified Framework

EMNLP 2021main

Various temporal knowledge graph (KG) completion models have been proposed in the recent literature. The models usually contain two parts, a temporal embedding layer and a score function derived from existing static KG modeling approaches. Since the approaches differ along several dimensions, includ…

2021

TimeTraveler: Reinforcement Learning for Temporal Knowledge Graph Forecasting

EMNLP 2021main

Temporal knowledge graph (TKG) reasoning is a crucial task that has gained increasing research interest in recent years. Most existing methods focus on reasoning at past timestamps to complete the missing facts, and there are only a few works of reasoning on known TKGs to forecast future facts. Comp…

2021

When Face Recognition Meets Occlusion: A New Benchmark

ICASSP 2021accepted

The existing face recognition datasets usually lack occlusion samples, which hinders the development of face recognition. Especially during the COVID-19 coronavirus epidemic, wearing a mask has become an effective means of preventing the virus spread. Traditional CNN-based face recognition models tr…

Cited by 0SourceScholar
2020

Image Super-Resolution Using Residual Global Context Network

ICASSP 2020accepted

Recent studies have showed that convolutional neural networks (CNN) can effectively improve the performance of single image super-resolution (SR). However, previous methods rarely considered long-range dependencies between pixels and channel-wise interdependencies at the same time. They ignores the…

Cited by 0SourceScholar
2019

Rain Streak Removal via Multi-scale Mixture Exponential Power Model

ICASSP 2019accepted

Rain streaks severely hamper the visible performance of the outdoor surveillance videos, which becomes an attractive issue in recent computer vision research. Existing methods usually encode rain streaks into Gaussian Mixture Model (GM-M). However, the limited number of Gaussian components in the GM…

Cited by 0SourceScholar
2017

A joint learning based Face Super Resolution approach via contextual topological structure

ICASSP 2017accepted

Face Super Resolution(FSR) is to infer High Resolution(HR) facial images from given Low Resolution(LR) ones with the assistance of LR and HR training pairs. Among existing methods, local patch based methods are superior in visual and objective quality than global based methods. These local patch bas…

Cited by 0SourceScholar