← Search

Han Fang

43 accepted papers

2026

ASAP: Exploiting the Satisficing Generalization Edge in Neural Combinatorial Optimization

ICML 2026poster

Deep Reinforcement Learning (DRL) has emerged as a promising approach for solving Combinatorial Optimization (CO) problems, such as the 3D Bin Packing Problem (3D-BPP), Traveling Salesman Problem (TSP), or Vehicle Routing Problem (VRP), but these neural solvers often exhibit brittleness when facing …

Cited by 0SourceScholar
2026

ASIR: Steganography for Diffusion Models via Antipodal Sampling and Iterative Recovery

ICML 2026poster

Messages embedded in diffusion generation noise suffer from severe attenuation due to denoising and VAE decoding, creating a persistent capacity–robustness trade-off. Identifying that extraction accuracy strictly correlates with the distance between candidate hypothesis images, we propose ASIR, a tr…

Cited by 0SourceScholar
2026

Adaptive Evidential Learning for Temporal-Semantic Robustness in Moment Retrieval

AAAI 2026technical

In the domain of moment retrieval, accurately identifying temporal segments within videos based on natural language queries remains challenging. Traditional methods often employ pre-trained models that struggle with fine-grained information and deterministic reasoning, leading to difficulties in ali

Cited by 0SourcePDFScholar
2026

FARI: Robust One-Step Inversion for Watermarking in Diffusion Models

ICLR 2026poster

Inversion-based watermarking is a promising approach to authenticate diffusion-generated images, yet practical use is bottlenecked by inversion that is both slow and error-prone. While the primary challenge in the watermarking setting is robustness against external distortions, existing approaches o…

Cited by 0SourceScholar
2026

Generalized Parallel Scaling with Interdependent Generations

ICLR 2026poster

Parallel LLM inference scaling involves sampling a set of $N>1$ responses for a single input prompt. However, these $N$ parallel responses tend to be generated independently from each other, partitioning compute resources and leaving potentially useful information in one generation untapped by other…

Cited by 0SourceScholar
2026

GlyphShield: Document Watermarking for the Physical World via Vector Typeface Synthesis

AAAI 2026technical

Document protection has become a critical issue for preventing unauthorized copying, distribution, and tampering. Document encryption is a proven solution, but it is not resistant to attacks from the physical world such as screenshots, printing and photographing. A common document protection techniq

Cited by 0SourcePDFScholar
2026

OMP: One-step Meanflow Policy with Directional Alignment

ICML 2026poster

Robot manipulation has increasingly adopted data-driven generative policy frameworks, yet the field faces a persistent trade-off: diffusion models suffer from high inference latency, while flow-based methods often require complex architectural constraints. Although in image generation domain, the Me…

Cited by 0SourceScholar
2026

Revisiting Coding-Based Approaches to Overcome the Curse of Dimensionality in Learning-Based Watermarking

ICML 2026poster

Deep learning–based watermarking has substantially improved robustness to real-world noise, but its performance degrades as the payload dimension increases. In contrast, coding-based methods such as quantization index modulation (QIM) do not suffer from this curse of dimensionality, although they ar…

Cited by 0SourceScholar
2026

Sim-to-Real: An Unsupervised Noise Layer for Screen-Camera Watermarking Robustness

AAAI 2026technical

Unauthorized screen capturing and dissemination pose severe security threats such as data leakage and information theft. Several studies propose robust watermarking methods to track the copyright of Screen-Camera (SC) images, facilitating post-hoc certification against infringement. These techniques

Cited by 0SourcePDFScholar
2025

AD2T: Adversarial Distortion Domain Translation for Robust Watermarking against Non-differentiable Distortions

ICASSP 2025accepted

Deep watermarking models optimize robustness by incorporating distortions between the encoder and decoder. To tackle non-differentiable distortions, current methods only train the decoder with distorted images, which breaks the joint optimization of the encoder-decoder, resulting in suboptimal perfo…

Cited by 0SourceScholar
2025

CoSDA: Enhancing the Robustness of Inversion-based Generative Image Watermarking Framework

AAAI 2025technical

Generative image watermarking inserts secret watermarks into generated images and plays an important role in tracing the usages of generative models. For watermarking of diffusion models, inversion-based framework emerges as an effective approach. Such framework employs a robust mechanism to embed…

Cited by 0SourcePDFScholar
2025

DeKeyNLU: Enhancing Natural Language to SQL Generation through Task Decomposition and Keyword Extraction

EMNLP 2025

Natural Language to SQL (NL2SQL) provides a new model-centric paradigm that simplifies database access for non-technical users by converting natural language queries into SQL commands. Recent advancements, particularly those integrating Retrieval-Augmented Generation (RAG) and Chain-of-Thought (CoT)

2025

END^2: Robust Dual-Decoder Watermarking Framework Against Non-Differentiable Distortions

AAAI 2025technical

DNN-based watermarking methods have rapidly advanced, with the ``Encoder-Noise Layer-Decoder'' (END) framework being the most widely used. To ensure end-to-end training, the noise layer in the framework must be differentiable. However, real-world distortions are often non-differentiable, leading to…

Cited by 0SourcePDFScholar
2025

FASTER: Face Attribute Sliders with Semantic Rewards

ICASSP 2025accepted

Large-scale text-to-image generative models have demonstrated remarkable success in generating diverse and high-quality faces. However, current methods for face editing often unintentionally modify facial features that are intended to be preserved. Multi-step denoising methods necessitate storing mu…

Cited by 0SourceScholar
2025

Geometric Alignment and Prior Modulation for View-Guided Point Cloud Completion on Unseen Categories

ICCV 2025poster

View-Guided Point Cloud Completion (VG-PCC) aims to reconstruct complete point clouds from partial inputs by referencing single-view images. While existing VG-PCC models perform well on in-class predictions, they exhibit significant performance drops when generalizing to unseen categories. We identi…

Cited by 0SourcePDFScholar
2025

Improving Model Factuality with Fine-grained Critique-based Evaluator

ACL 2025long

Factuality evaluation aims to detect factual errors produced by language models (LMs) and hence guide the development of more factual models. Towards this goal, we train a factuality evaluator, FenCE, that provides LM generators with claim-level factuality feedback. In particular, we train FenCE to…

2025

Learning Auxiliary Tasks Improves Reference-Free Hallucination Detection in Open-Domain Long-Form Generation

ACL 2025short

Hallucination, the generation of factually incorrect information, remains a significant challenge for large language models (LLMs), especially in open-domain long-form generation. Existing approaches for detecting hallucination in long-form tasks either focus on limited domains or rely heavily on ex…

Cited by 0SourcePDFScholar
2025

Lie Detector: Unified Backdoor Detection via Cross-Examination Framework

NeurIPS 2025poster

Institutions with limited data and computing resources often outsource model training to third-party providers in a semi-honest setting, assuming adherence to prescribed training protocols with pre-defined learning paradigm (e.g., supervised or semi-supervised learning). However, this practice can i…

Cited by 0SourceScholar
2025

ROAR: Reducing Inversion Error in Generative Image Watermarking

ICCV 2025poster

Generative image watermarking enables the proactive detection and traceability of generated images. Among existing methods, inversion-based frameworks achieve highly conceal ed watermark embedding by injecting watermarks into the latent representation before the diffusion process. The robustness of…

Cited by 0SourcePDFScholar
2025

RoPaSS: Robust Watermarking for Partial Screen-Shooting Scenarios

AAAI 2025technical

Screen-shooting robust watermarking is an effective means of preventing screen content leakage from unauthorized camera shooting, as it can trace the leaked source through the watermark extraction thereby providing an effective deterrent. However, current screen-shooting resilient watermarking schem…

Cited by 0SourcePDFScholar
2025

SynTag: Enhancing the Geometric Robustness of Inversion-based Generative Image Watermarking

ICCV 2025poster

Robustness is significant for generative image watermarking, typically achieved by injecting distortion-invariant watermark features. The leading paradigm, i.e., inversion-based framework, excels against non-geometric distortions but struggles with geometric ones. To address this, we propose SynTag,…

Cited by 0SourcePDFScholar
2025

T2SMark: Balancing Robustness and Diversity in Noise-as-Watermark for Diffusion Models

NeurIPS 2025poster

Diffusion models have advanced rapidly in recent years, producing high-fidelity images while raising concerns about intellectual property protection and the misuse of generative AI. Image watermarking for diffusion models, particularly Noise-as-Watermark (NaW) methods, encode watermark as specific s…

Cited by 0SourceScholar
2025

TAG-WM: Tamper-Aware Generative Image Watermarking via Diffusion Inversion Sensitivity

ICCV 2025poster

AI-generated content (AIGC) enables efficient visual creation but raises copyright and authenticity risks. As a common technique for integrity verification and source tracing, digital image watermarking is regarded as a potential solution to above issues. However, the widespread adoption and advanci…

Cited by 0SourcePDFScholar
2025

Think Smarter not Harder: Adaptive Reasoning with Inference Aware Optimization

ICML 2025poster

Solving mathematics problems has been an intriguing capability of large language models, and many efforts have been made to improve reasoning by extending reasoning length, such as through self-correction and extensive long chain-of-thoughts. While promising in problem-solving, advanced long reasoni…

Cited by 7SourcePDFScholar
2025

Tracking-Aware Deformation Field Estimation for Non-rigid 3D Reconstruction in Robotic Surgeries

IROS 2025

Minimally invasive procedures have been advanced rapidly by the robotic laparoscopic surgery. The latter greatly assists surgeons in sophisticated and precise operations with reduced invasiveness. Nevertheless, it is still safety critical to be aware of even the least tissue deformation during instr

Cited by 1SourcecodeScholar
2025

Trusted Unified Feature-Neighborhood Dynamics for Multi-View Classification

AAAI 2025technical

Multi-view classification (MVC) faces inherent challenges due to domain gaps and inconsistencies across different views, often resulting in uncertainties during the fusion process. While Evidential Deep Learning (EDL) has been effective in addressing view uncertainty, existing methods predominantly…

2025

Ultra-high Resolution Watermarking Framework Resistant to Extreme Cropping and Scaling

NeurIPS 2025poster

Recent developments in DNN-based image watermarking techniques have achieved impressive results in protecting digital content. However, most existing methods are constrained to low-resolution images as they need to encode the entire image, leading to prohibitive memory and computational costs when a…

Cited by 0SourceScholar
2025

ViCo: A Multitask Video-enhanced and Cognition-preserving Modality Alignment Training Framework

ICASSP 2025accepted

The rapid development of multimodal large language models (MLLMs) has brought significant breakthroughs to this field. However, current MLLMs typically rely on vision instruction tuning based on large language models (LLMs) to endow them with multimodal capabilities, which may lead to low video util…

Cited by 0SourceScholar
2024

Effective Long-Context Scaling of Foundation Models

NAACL 2024long

We present an effective recipe to train strong long-context LLMs that are capable of utilizing massive context windows of up to 32,000 tokens. Our models are built through continual pretraining from Llama 2 checkpoints with longer text sequences and on a dataset where long texts are upsampled. We pe…

Cited by 231SourcePDFScholar
2024

Gaussian Shading: Provable Performance-Lossless Image Watermarking for Diffusion Models

CVPR 2024poster

Ethical concerns surrounding copyright protection and inappropriate content generation pose challenges for the practical implementation of diffusion models. One effective solution involves watermarking the generated images. However existing methods often compromise the model performance or require a…

2024

Generating Stereophonic Music with Single-Stage Language Models

ICASSP 2024accepted

The recent success of audio language models (LMs) has revolutionized the field of neural music generation. Among all audio LM approaches, MusicGen has demonstrated the success of a single-stage LMs based music generation framework, without needing to train multiple LMs. Despite its promising perform…

Cited by 0SourceScholar
2024

INViT: A Generalizable Routing Problem Solver with Invariant Nested View Transformer

ICML 2024poster

Recently, deep reinforcement learning has shown promising results for learning fast heuristics to solve routing problems. Meanwhile, most of the solvers suffer from generalizing to an unseen distribution or distributions with different scales. To address this issue, we propose a novel architecture,…

2024

MuST: Robust Image Watermarking for Multi-Source Tracing

AAAI 2024technical

In recent years, with the popularity of social media applications, massive digital images are available online, which brings great convenience to image recreation. However, the use of unauthorized image materials in multi-source composite images is still inadequately regulated, which may cause signi…

2024

Representation Deficiency in Masked Language Modeling

ICLR 2024poster

Masked Language Modeling (MLM) has been one of the most prominent approaches for pretraining bidirectional text encoders due to its simplicity and effectiveness. One notable concern about MLM is that the special $\texttt{[MASK]}$ symbol causes a discrepancy between pretraining data and downstream da…

2023

AutoStegaFont: Synthesizing Vector Fonts for Hiding Information in Documents

AAAI 2023technical

Hiding information in text documents has been a hot topic recently, with the most typical schemes of utilizing fonts. By constructing several fonts with similar appearances, information can be effectively represented and embedded in documents. However, due to the unstructured characteristic, font ve…

Cited by 3SourcePDFScholar
2023

DeAR: A Deep-Learning-Based Audio Re-recording Resilient Watermarking

AAAI 2023technical

Audio watermarking is widely used for leaking source tracing. The robustness of the watermark determines the traceability of the algorithm. With the development of digital technology, audio re-recording (AR) has become an efficient and covert means to steal secrets. AR process could drastically dest…

Cited by 43SourcePDFScholar
2023

Flow-Based Robust Watermarking with Invertible Noise Layer for Black-Box Distortions

AAAI 2023technical

Deep learning-based digital watermarking frameworks have been widely studied recently. Most existing methods adopt an ``encoder-noise layer-decoder''-based architecture where the embedding and extraction processes are accomplished separately by the encoder and the decoder. However, one potential dra…

2023

Tracing the Origin of Adversarial Attack for Forensic Investigation and Deterrence

ICCV 2023poster

Deep neural networks are vulnerable to adversarial attacks. In this paper, we take the role of investigators who want to trace the attack and identify the source, that is, the particular model which the adversarial examples are generated from. Techniques derived would aid forensic investigation of a…

Cited by 5PDFcodeScholar
2022

Speech Pattern Based Black-Box Model Watermarking for Automatic Speech Recognition

ICASSP 2022accepted

As an effective method for intellectual property (IP) protection, model watermarking technology has been applied on a wide variety of deep neural networks (DNN), including speech classification models. However, how to design a black-box watermarking scheme for automatic speech recognition (ASR) mode…

Cited by 0SourceScholar
2021

Adaptive Re-Balancing Network with Gate Mechanism for Long-Tailed Visual Question Answering

ICASSP 2021accepted

Visual Question Answering (VQA) is a challenging task which requires a fine-grained semantic understanding of visual and textual contents. Existing works focus on better modality representations. However, these methods give little consideration to the long-tailed data distribution in common VQA data…

Cited by 0SourceScholar
2020

Generate to Adapt: Resolution Adaption Network for Surveillance Face Recognition

ECCV 2020poster

Although deep learning techniques have largely improved face recognition, unconstrained surveillance face recognition (FR) is still an unsolved challenge, due to the limited training data and the gap of domain distribution. Previous methods mostly match low-resolution and high-resolution faces in di…

Cited by 27SourcePDFScholar
2019

DUP-Net: Denoiser and Upsampler Network for 3D Adversarial Point Clouds Defense

ICCV 2019poster

Neural networks are vulnerable to adversarial examples, which poses a threat to their application in security sensitive systems. We propose a Denoiser and UPsampler Network (DUP-Net) structure as defenses for 3D adversarial point cloud classification, where the two modules reconstruct surface smooth…

Cited by 202PDFcodeScholar