← Search

Qian Wang

76 accepted papers

2026

A Training-Free Framework for High-Fidelity Appearance Transfer via Diffusion Transformers

ICASSP 2026poster

Diffusion Transformers (DiTs) excel at generation, but their global self-attention makes controllable, reference-image-based editing a distinct challenge. Unlike U-Nets, naively injecting local appearance into a DiT can disrupt its holistic scene structure. We address this by proposing the first tra…

Cited by 0SourcePDFScholar
2026

Beyond Sequential Tools: A Unified VLM Agent System for Photographic Post-Processing via Dynamic Multi-Expert Fusion

CVPR 2026

Real-world image restoration is challenged by complex, coupled degradations. Existing "all-in-one" models often lack generalization, while agentic systems suffer from inefficient sequential tool invocation. We propose a VLM-guided one-shot framework for universal photographic post-processing. Our sy

Cited by 0SourcecodeScholar
2026

Divide and Conquer: Reliable Multi-View Evidential Learning for Deepfake Detection

ICML 2026poster

With the evolution of generative models, deepfakes have achieved near-perfect semantic realism, leaving forensic traces only in subtle structural anomalies. However, existing single-view paradigms often fail to generalize, as dominant semantic features overwhelm subtle artifact cues within entangled…

Cited by 0SourceScholar
2026

LLM DNA: Tracing Model Evolution via Functional Representations

ICLR 2026oral

The explosive growth of large language models (LLMs) has created a vast but opaque landscape: millions of models exist, yet their evolutionary relationships through fine-tuning, distillation, or adaptation are often undocumented or unclear, complicating LLM management. Existing methods are limited b…

Cited by 0SourcecodeScholar
2026

Subsecond 3D Mesh Generation for Robot Manipulation

ICRA 2026poster

3D meshes are a fundamental representation widely used in computer science and engineering. In robotics, they are particularly valuable because they capture objects in a form that aligns directly with how robots interact with the physical world, enabling core capabilities such as predicting stable g…

2026

Tell as You Want: Customizing Image Narrative with Knowledge and Thoughts

AAAI 2026technical

With the advancement of vision-language models, image captioning has made significant progress, leading to the generation of more accurate and detailed descriptions. Current image captioning primarily focuses on describing the apparent visual characteristics, which are easily observed by most humans

Cited by 0SourcePDFScholar
2026

Turning Stale Gradients into Stable Gradients: Coherent Coordinate Descent with Implicit Landscape Smoothing for Lightweight Zeroth-Order Optimization

ICML 2026poster

Zeroth-Order (ZO) optimization is pivotal for scenarios where backpropagation is unavailable, such as memory-constrained on-device learning and black-box optimization. However, existing methods face a stark trade-off: they are either sample-inefficient (e.g., standard finite differences) or suffer f…

Cited by 0SourceScholar
2026

Two-Time-Scale Composite Learning Online Identification and Control for Compliant-Joint Robots

ICRA 2026poster

SP-based synthesis yields two-time-scale control that allows compliant-joint robots to achieve high-quality tracking at low implementation cost. Composite learning enables exact online identification and control of robots without the stringent condition known as persistent excitation (PE). However, …

Cited by 0Scholar
2025

AlignedGen: Aligning Style Across Generated Images

NeurIPS 2025poster

Diffusion-based generative models struggle to maintain high style consistency across generated images via text description. Although several style-aligned image generation methods have been proposed to address this issue, they exhibit suboptimal performance and are primarily built upon the U-Net arc…

Cited by 0SourcecodeScholar
2025

Anywhere: A Multi-Agent Framework for User-Guided, Reliable, and Diverse Foreground-Conditioned Image Generation

AAAI 2025technical

Recent advancements in image-conditioned image generation have demonstrated substantial progress. However, foreground-conditioned image generation remains underexplored, encountering challenges such as compromised object integrity, foreground-background inconsistencies, limited diversity, and reduce…

Cited by 0SourcePDFScholar
2025

BEAST: Efficient Tokenization of B-Splines Encoded Action Sequences for Imitation Learning

NeurIPS 2025poster

We present the B-spline Encoded Action Sequence Tokenizer (BEAST), a novel action tokenizer that encodes action sequences into compact discrete or continuous tokens using B-splines. In contrast to existing action tokenizers based on vector quantization or byte pair encoding, BEAST requires no separ…

Cited by 0SourceScholar
2025

DCIM-AVSR: Efficient Audio-Visual Speech Recognition via Dual Conformer Interaction Module

ICASSP 2025accepted

Speech recognition is the technology that enables machines to interpret and process human speech, converting spoken language into text or commands. This technology is essential for applications such as virtual assistants, transcription services, and communication tools. The Audio-Visual Speech Recog…

Cited by 0SourceScholar
2025

Detecting Adversarial Data Using Perturbation Forgery

CVPR 2025poster

As a defense strategy against adversarial attacks, adversarial detection aims to identify and filter out adversarial data from the data flow based on discrepancies in distribution and noise patterns between natural and adversarial data. Although previous detection methods achieve high performance in…

2025

Differential-Flatness-Based Tracking Control for Tractor-Trailers in Reversing Maneuvers

IROS 2025

In this paper, we propose a differential-flatness-based controller (DFBC) for precise trajectory tracking of tractor-trailers, particularly during reversing maneuvers, which are challenging due to unstable equilibrium points. The proposed controller leverages the differential flatness property of tr

Cited by 0SourceScholar
2025

EditCLIP: Representation Learning for Image Editing

ICCV 2025poster

We introduce EditCLIP, a novel representation-learning approach for image editing. Our method learns a unified representation of edits by jointly encoding an input image and its edited counterpart, effectively capturing their transformation. To evaluate its effectiveness, we employ EditCLIP to solve…

2025

FreeControl: Efficient, Training-Free Structural Control via One-Step Attention Extraction

NeurIPS 2025poster

Controlling the spatial and semantic structure of diffusion-generated images remains a challenge. Existing methods like ControlNet rely on handcrafted condition maps and retraining, limiting flexibility and generalization. Inversion-based approaches offer stronger alignment but incur high inference…

Cited by 0SourceScholar
2025

From Zero to Detail: Deconstructing Ultra-High-Definition Image Restoration from Progressive Spectral Perspective

CVPR 2025poster

Ultra-high-definition (UHD) image restoration faces significant challenges due to its high resolution, complex content, and intricate details. To cope with these challenges, we analyze the restoration process in depth through a progressive spectral perspective, and deconstruct the complex UHD restor…

2025

Hierarchical Frequency Tagging Probe (HFTP): A Unified Approach to Investigate Syntactic Structure Representations in Large Language Models and the Human Brain

NeurIPS 2025poster

Large Language Models (LLMs) demonstrate human-level or even superior language abilities, effectively modeling syntactic structures, yet the specific computational units responsible remain unclear. A key question is whether LLM behavioral capabilities stem from mechanisms akin to those in the human…

Cited by 0SourcecodeScholar
2025

MITracker: Multi-View Integration for Visual Object Tracking

CVPR 2025highlight

Multi-view object tracking (MVOT) offers promising solutions to challenges such as occlusion and target loss, which are common in traditional single-view tracking. However, progress has been limited by the lack of comprehensive multi-view datasets and effective cross-view integration methods. To ove…

Cited by 0SourcePDFScholar
2025

MUC: Mixture of Uncalibrated Cameras for Robust 3D Human Body Reconstruction

AAAI 2025technical

Multiple cameras can provide comprehensive multi-view video coverage of a person. Fusing this multi-view data is crucial for tasks like behavioral analysis, although it traditionally requires camera calibration—a process that is often complex. Moreover, previous studies have overlooked the challenge…

2025

MegaAgent: A Large-Scale Autonomous LLM-based Multi-Agent System Without Predefined SOPs

ACL 2025finding

LLM-based multi-agent systems (MAS) have shown promise in tackling complex tasks. However, existing solutions often suffer from limited agent coordination and heavy reliance on predefined Standard Operating Procedures (SOPs), which demand extensive human input. To address these limitations, we propo…

2025

NovelQA: Benchmarking Question Answering on Documents Exceeding 200K Tokens

ICLR 2025poster

Recent advancements in Large Language Models (LLMs) have pushed the boundaries of natural language processing, especially in long-context understanding. However, the evaluation of these models' long-context abilities remains a challenge due to the limitations of current benchmarks. To address this g…

2025

One-Shot Learning for Pose-Guided Person Image Synthesis in the Wild

ICASSP 2025accepted

Current Pose-Guided Person Image Synthesis (PGPIS) methods depend heavily on large amounts of labeled triplet data to train the generator in a supervised manner. However, they often falter when applied to in-the-wild samples, primarily due to the distribution gap between the training datasets and re…

Cited by 4SourceScholar
2025

PointMapPolicy: Structured Point Cloud Processing for Multi-Modal Imitation Learning

NeurIPS 2025poster

Robotic manipulation systems benefit from complementary sensing modalities, where each provides unique environmental information. Point clouds capture detailed geometric structure, while RGB images provide rich semantic context. Current point cloud methods struggle to capture fine-grained detail, es…

Cited by 0SourcecodeScholar
2025

RAGD: Regional-Aware Diffusion Model for Text-to-Image Generation

ICCV 2025poster

Regional prompting, or compositional generation, which enables fine-grained spatial control, has gained increasing attention for its practicality in real-world applications. However, previous methods either introduce additional trainable modules, thus only applicable to specific models, or manipulat…

2025

Revisiting Adversarial Patch Defenses on Object Detectors: Unified Evaluation, Large-Scale Dataset, and New Insights

ICCV 2025poster

Developing reliable defenses against patch attacks on object detectors has attracted increasing interest. However, we identify that existing defense evaluations lack a unified and comprehensive framework, resulting in inconsistent and incomplete assessments of current methods. To address this issue,…

2025

Solid-SQL: Enhanced Schema-linking based In-context Learning for Robust Text-to-SQL

COLING 2025main

Recently, large language models (LLMs) have significantly improved the performance of text-to-SQL systems. Nevertheless, many state-of-the-art (SOTA) approaches have overlooked the critical aspect of system robustness. Our experiments reveal that while LLM-driven methods excel on standard datasets,…

2025

Spatially-variant Blur Degradation Model Based on Depth Estimation

ICASSP 2025accepted

It is well known that the number of aligned images in the single image super-resolution (SISR) models training is limited. Synthesizing data is an effective way to address this issue. However, many degradation models only consider using spatially-invariant blur kernels to blur high-resolution (HR) i…

Cited by 0SourceScholar
2025

The Invisible Hand: Unveiling Provider Bias in Large Language Models for Code Generation

ACL 2025long

Large Language Models (LLMs) have emerged as the new recommendation engines, surpassing traditional methods in both capability and scope, particularly in code generation. In this paper, we reveal a novel **provider bias** in LLMs: without explicit directives, these models show systematic preferences…

2025

VidSeg: Training-free Video Semantic Segmentation based on Diffusion Models

CVPR 2025poster

We introduce the first training-free approach for Video Semantic Segmentation (VSS) based on pre-trained diffusion models. A growing research direction attempts to employ diffusion models to perform downstream vision tasks by exploiting their deep understanding of image semantics. Yet, the majority…

Cited by 0SourcePDFScholar
2024

360DVD: Controllable Panorama Video Generation with 360-Degree Video Diffusion Model

CVPR 2024poster

Panorama video recently attracts more interest in both study and application courtesy of its immersive experience. Due to the expensive cost of capturing 360-degree panoramic videos generating desirable panorama videos by prompts is urgently required. Lately the emerging text-to-video (T2V) diffusio…

2024

Context-Driven Index Trimming: A Data Quality Perspective to Enhancing Precision of RALMs

EMNLP 2024finding

Retrieval-Augmented Large Language Models(RALMs) have made significant strides in enhancing the accuracy of generated responses. However, existing research often overlooks the data quality issues within retrieval results, often caused by inaccurate existing vector-distance-based retrieval methods. W…

2024

CryptoTrade: A Reflective LLM-based Agent to Guide Zero-shot Cryptocurrency Trading

EMNLP 2024main

The utilization of Large Language Models (LLMs) in financial trading has primarily been concentrated within the stock market, aiding in economic and financial decisions. Yet, the unique opportunities presented by the cryptocurrency market, noted for its on-chain data’s transparency and the critical…

2024

Delocate: Detection and Localization for Deepfake Videos with Randomly-Located Tampered Traces

IJCAI 2024poster

Deepfake videos are becoming increasingly realistic, showing few tampering traces on facial areas that vary between frames. Consequently, existing Deepfake detection methods struggle to detect unknown domain Deepfake videos while accurately locating the tampered region. To address this limitation,…

2024

DiaGBT: An Explainable and Evolvable Robot Control Framework using Dialogue Generative Behavior Trees

IROS 2024poster

Manipulating robots using natural language is the preferred way for non-technical specialists. The challenge lies in reliability and adaptability especially when robots operate in unstructured surroundings. In this paper, we propose a novel framework called Dialogue Generative Behavior Trees (DiaGBT…

Cited by 0SourceScholar
2024

Dynamic Budget Throttling in Repeated Second-Price Auctions

AAAI 2024technical

In today's online advertising markets, a crucial requirement for an advertiser is to control her total expenditure within a time horizon under some budget. Among various budget control methods, throttling has emerged as a popular choice, managing an advertiser's total expenditure by selecting only…

Cited by 10SourcePDFScholar
2024

EX-Graph: A Pioneering Dataset Bridging Ethereum and X

ICLR 2024poster

While numerous public blockchain datasets are available, their utility is constrained by an exclusive focus on blockchain data. This constraint limits the incorporation of relevant social network data into blockchain analysis, thereby diminishing the breadth and depth of insight that can be derived.…

2024

Every Node Is Different: Dynamically Fusing Self-Supervised Tasks for Attributed Graph Clustering

AAAI 2024technical

Attributed graph clustering is an unsupervised task that partitions nodes into different groups. Self-supervised learning (SSL) shows great potential in handling this task, and some recent studies simultaneously learn multiple SSL tasks to further boost performance. Currently, different SSL tasks ar…

2024

From Text Segmentation to Enhanced Representation Learning: A Novel Approach to Multi-Label Classification for Long Texts

EMNLP 2024finding

Multi-label text classification (MLTC) is an important task in the field of natural language processing. Most existing models rely on high-quality text representations provided by pre-trained language models (PLMs). They hence face the challenge of input length limitation caused by PLMs, when dealin…

2024

MaIL: Improving Imitation Learning with Selective State Space Models

CoRL 2024poster

This work introduces Mamba Imitation Learning (MaIL), a novel imitation learning (IL) architecture that offers a computationally efficient alternative to state-of-the-art (SoTA) Transformer policies. Transformer-based policies have achieved remarkable results due to their ability in handling human-r…

Cited by 7SourceScholar
2024

Mining Gaze for Contrastive Learning toward Computer-Assisted Diagnosis

AAAI 2024technical

Obtaining large-scale radiology reports can be difficult for medical images due to ethical concerns, limiting the effectiveness of contrastive pre-training in the medical image domain and underscoring the need for alternative methods. In this paper, we propose eye-tracking as an alternative to text…

2024

Multi-Chain Graphs of Graphs: A New Approach to Analyzing Blockchain Datasets

NeurIPS 2024poster

Machine learning applied to blockchain graphs offers significant opportunities for enhanced data analysis and applications. However, the potential of this field is constrained by the lack of a large-scale, cross-chain dataset that includes hierarchical graph-level data. To address this issue, we pre…

2024

Multiscale Matching Driven by Cross-Modal Similarity Consistency for Audio-Text Retrieval

ICASSP 2024accepted

Audio-text retrieval (ATR), which retrieves a relevant caption given an audio clip (A2T) and vice versa (T2A), has recently attracted much research attention. Existing methods typically aggregate information from each modality into a single vector for matching, but this sacrifices local details and…

Cited by 0SourceScholar
2024

ROBIN: Robust and Invisible Watermarks for Diffusion Models with Adversarial Optimization

NeurIPS 2024poster

Watermarking generative content serves as a vital tool for authentication, ownership protection, and mitigation of potential misuse. Existing watermarking methods face the challenge of balancing robustness and concealment. They empirically inject a watermark that is both invisible and robust and pas…

2024

Revisiting Adversarial Training Under Long-Tailed Distributions

CVPR 2024poster

Deep neural networks are vulnerable to adversarial attacks leading to erroneous outputs. Adversarial training has been recognized as one of the most effective methods to counter such attacks. However existing adversarial training techniques have predominantly been evaluated on balanced datasets wher…

2023

Coordinated Dynamic Bidding in Repeated Second-Price Auctions with Budgets

ICML 2023poster

In online ad markets, a rising number of advertisers are employing bidding agencies to participate in ad auctions. These agencies are specialized in designing online algorithms and bidding on behalf of their clients. Typically, an agency usually has information on multiple advertisers, so she can po…

Cited by 6SourcePDFScholar
2023

Detecting Adversarial Faces Using Only Real Face Self-Perturbations

IJCAI 2023poster

Adversarial attacks aim to disturb the functionality of a target system by adding specific noise to the input samples, bringing potential threats to security and robustness when applied to facial recognition systems. Although existing defense techniques achieve high accuracy in detecting some specif…

2023

Downstream-agnostic Adversarial Examples

ICCV 2023poster

Self-supervised learning usually uses a large amount of unlabeled data to pre-train an encoder which can be used as a general-purpose feature extractor, such that downstream users only need to perform fine-tuning operations to enjoy the benefit of "big model". Despite this promising prospect, the se…

Cited by 30PDFcodeScholar
2023

FC-TrackNet: Fast Convergence Net for 6D Pose Tracking in Synthetic Domains

AAAI 2023technical

In this work, we propose a fast convergence track net, or FC-TrackNet, based on a synthetic data-driven approach to maintaining long-term 6D pose tracking. Comparison experiments are performed on two different datasets, The results demonstrate that our approach can achieve a consistent tracking freq…

Cited by 1SourcePDFScholar
2023

Implicit Identity Driven Deepfake Face Swapping Detection

CVPR 2023poster

In this paper, we consider the face swapping detection from the perspective of face identity. Face swapping aims to replace the target face with the source face and generate the fake face that the human cannot distinguish between real and fake. We argue that the fake face contains the explicit ident…

Cited by 135SourcePDFScholar
2023

Learning to Bid in Repeated First-Price Auctions with Budgets

ICML 2023poster

Budget management strategies in repeated auctions have received growing attention in online advertising markets. However, previous work on budget management in online bidding mainly focused on second-price auctions. The rapid shift from second-price auctions to first-price auctions for online ads in…

Cited by 20SourcePDFScholar
2023

Panoptic Compositional Feature Field for Editable Scene Rendering With Network-Inferred Labels via Metric Learning

CVPR 2023poster

Despite neural implicit representations demonstrating impressive high-quality view synthesis capacity, decomposing such representations into objects for instance-level editing is still challenging. Recent works learn object-compositional representations supervised by ground truth instance annotation…

Cited by 7SourcePDFScholar
2023

Revisiting Adversarial Robustness Distillation from the Perspective of Robust Fairness

NeurIPS 2023poster

Adversarial Robustness Distillation (ARD) aims to transfer the robustness of large teacher models to small student models, facilitating the attainment of robust performance on resource-limited devices. However, existing research on ARD primarily focuses on the overall robustness of student models, o…

2023

When Does Aggregating Multiple Skills with Multi-Task Learning Work? A Case Study in Financial NLP

ACL 2023long

Multi-task learning (MTL) aims at achieving a better model by leveraging data and knowledge from multiple tasks. However, MTL does not always work – sometimes negative transfer occurs between tasks, especially when aggregating loosely related skills, leaving it an open question when MTL works. Previ…

2022

A Few Seconds Can Change Everything: Fast Decision-based Attacks against DNNs

IJCAI 2022poster

Previous researches have demonstrated deep learning models' vulnerabilities to decision-based adversarial attacks, which craft adversarial examples based solely on information from output decisions (top-1 labels). However, existing decision-based attacks have two major limitations, i.e., expensive q…

2022

Addressing Asymmetry in Multilingual Neural Machine Translation with Fuzzy Task Clustering

COLING 2022main

Multilingual neural machine translation (NMT) enables positive knowledge transfer among multiple translation tasks with a shared underlying model, but a unified multilingual model usually suffers from capacity bottleneck when tens or hundreds of languages are involved. A possible solution is to clus…

2022

How to Distribute Data across Tasks for Meta-Learning?

AAAI 2022technical

Meta-learning models transfer the knowledge acquired from previous tasks to quickly learn new ones. They are trained on benchmarks with a fixed number of data points per task. This number is usually arbitrary and it is unknown how it affects performance at testing. Since labelling of data is expensi…

Cited by 7SourcePDFScholar
2022

Reliability Exploration with Self-Ensemble Learning for Domain Adaptive Person Re-identification

AAAI 2022technical

Person re-identifcation (Re-ID) based on unsupervised domain adaptation (UDA) aims to transfer the pre-trained model from one labeled source domain to an unlabeled target domain. Existing methods tackle this problem by using clustering methods to generate pseudo labels. However, pseudo labels produc…

Cited by 46SourcePDFScholar
2021

CARTL: Cooperative Adversarially-Robust Transfer Learning

ICML 2021oral

Transfer learning eases the burden of training a well-performed model from scratch, especially when training data is scarce and computation power is limited. In deep learning, a typical strategy for transfer learning is to freeze the early layers of a pre-trained model and fine-tune the rest of its…

2021

InverseNet: Augmenting Model Extraction Attacks with Training Data Inversion

IJCAI 2021poster

Cloud service providers, including Google, Amazon, and Alibaba, have now launched machine-learning-as-a-service (MLaaS) platforms, allowing clients to access sophisticated cloud-based machine learning models via APIs. Unfortunately, however, the commercial value of these models makes them alluring t…

Cited by 56SourcePDFScholar
2021

Recent Advances in Adversarial Training for Adversarial Robustness

IJCAI 2021poster

Adversarial training is one of the most effective approaches for deep learning models to defend against adversarial examples. Unlike other defense strategies, adversarial training aims to enhance the robustness of models intrinsically. During the past few years, adversarial training has been studi…

Cited by 596SourcePDFScholar
2021

Synchronous Interactive Decoding for Multilingual Neural Machine Translation

AAAI 2021technical

To simultaneously translate a source language into multiple different target languages is one of the most common scenarios of multilingual translation. However, existing methods cannot make full use of translation model information during decoding, such as intra-lingual and inter-lingual future info…

2020

A Spatiotemporal Volumetric Interpolation Network for 4D Dynamic Medical Image

CVPR 2020poster

Dynamic medical images are often limited in its application due to the large radiation doses and longer image scanning and reconstruction times. Existing methods attempt to reduce the volume samples in the dynamic sequence by interpolating the volumes between the acquired samples. However, these met…

Cited by 34PDFcodeScholar
2019

advPattern: Physical-World Attacks on Deep Person Re-Identification via Adversarially Transformable Patterns

ICCV 2019poster

Person re-identification (re-ID) is the task of matching person images across camera views, which plays an important role in surveillance and security applications. Inspired by great progress of deep learning, deep re-ID models began to be popular and gained state-of-the-art performance. However, re…

Cited by 61PDFcodeScholar