← Search

Lanjun Wang

21 accepted papers

2026

A Training-Free Framework for High-Fidelity Appearance Transfer via Diffusion Transformers

ICASSP 2026poster

Diffusion Transformers (DiTs) excel at generation, but their global self-attention makes controllable, reference-image-based editing a distinct challenge. Unlike U-Nets, naively injecting local appearance into a DiT can disrupt its holistic scene structure. We address this by proposing the first tra…

Cited by 0SourcePDFScholar
2026

Key Decision-Makers in Multi-Agent Debates: Who Holds the Power?

AAAI 2026technical

Recent studies on LLM agent scaling have highlighted the potential of Multi-Agent Debate (MAD) to enhance reasoning abilities. However, the critical aspect of role allocation strategies remains underexplored. In this study, we demonstrate that allocating roles with differing viewpoints to specific p

Cited by 0SourcePDFScholar
2026

Layer-wise Instance Binding for Regional and Occlusion Control in Text-to-Image Diffusion Transformers

CVPR 2026

Region-instructed layout control in text-to-image generation is highly practical, yet existing methods suffer from limitations: (i) training-based approaches inherit data bias and often degrade image quality, and (ii) current techniques struggle with occlusion order, limiting real-world usability. T

Cited by 0SourcecodeScholar
2026

Reason2Attack: Jailbreaking Text-to-Image Models via LLM Reasoning

AAAI 2026technical

Text-to-Image (T2I) models typically deploy safety mechanisms to prevent the generation of sensitive images. Unfortunately, recent jailbreaking attack methods manually design instructions for the LLM to generate adversarial prompts, which effectively exposing safety vulnerabilities of T2I models.

Cited by 0SourcePDFScholar
2026

T2I-RiskyPrompt: A Benchmark for Safety Evaluation, Attack, and Defense on Text-to-Image Model

AAAI 2026technical

Using risky text prompts, such as pornography and violent prompts, to test the safety of text-to-image (T2I) models is a critical task. However, existing risky prompt datasets are limited in three key areas: 1) limited risky categories, 2) coarse-grained annotation, and 3) low effectiveness. To addr

Cited by 0SourcePDFScholar
2026

Thinking as Society: Multi-Social-Agent Self-Distillation for Multimodal Misinformation Detection

ICLR 2026poster

Multimodal Misinformation Detection (MMD) in realistic, mixed-sourced scenarios must incorporate robust reasoning capabilities to handle the social complexity and diverse types of forgeries. While MLLM-based agents are increasingly used for MMD task due to their powerful reasoning abilities, they su…

Cited by 0SourceScholar
2026

When Planning Fails Despite Correct Execution: On Epistemic Calibration for LLM-Based Multi-Agent Systems

ICML 2026poster

LLM-based multi-agent systems can fail even when planned actions are executed correctly because agents may misjudge their knowledge when evaluating plan feasibility, a phenomenon we term epistemic miscalibration in planning. Unlike execution errors, epistemic miscalibration is latent during planning…

Cited by 0SourceScholar
2025

Aesthetic Perception Prompting for Interpretable Image Aesthetics Assessment with MLLMs

ICASSP 2025accepted

Image Aesthetic Assessment (IAA) aims to rate the aesthetic quality of images and has many practical applications. However, existing methods typically rely on limited annotated data for training, leading to two key issues: 1) score-only predictions lack interpretability, making it hard for users to…

Cited by 0SourceScholar
2025

Anywhere: A Multi-Agent Framework for User-Guided, Reliable, and Diverse Foreground-Conditioned Image Generation

AAAI 2025technical

Recent advancements in image-conditioned image generation have demonstrated substantial progress. However, foreground-conditioned image generation remains underexplored, encountering challenges such as compromised object integrity, foreground-background inconsistencies, limited diversity, and reduce…

Cited by 0SourcePDFScholar
2025

RoCo-Sim: Enhancing Roadside Collaborative Perception through Foreground Simulation

ICCV 2025poster

Roadside Collaborative Perception refers to a system where multiple roadside units collaborate to pool their perceptual data, assisting vehicles in enhancing their environmental awareness. Existing roadside perception methods concentrate on model design but overlook data issues like calibration erro…

2025

TRCE: Towards Reliable Malicious Concept Erasure in Text-to-Image Diffusion Models

ICCV 2025poster

Recent advances in text-to-image diffusion models enable photorealistic image generation, but they also risk producing malicious content, such as NSFW images. To mitigate risk, concept erasure methods are studied to facilitate the model to unlearn specific concepts. However, current studies struggle…

2024

AnyScene: Customized Image Synthesis with Composited Foreground

CVPR 2024poster

Recent advancements in text-to-image technology have significantly advanced the field of image customization. Among various applications the task of customizing diverse scenes for user-specified composited elements holds great application value but has not been extensively explored. Addressing this…

Cited by 1SourcePDFScholar
2022

Cosine Model Watermarking against Ensemble Distillation

AAAI 2022technical

Many model watermarking methods have been developed to prevent valuable deployed commercial models from being stealthily stolen by model distillations. However, watermarks produced by most existing model watermarking methods can be easily evaded by ensemble distillation, because averaging the outpu…

Cited by 26SourcePDFScholar
2022

Human Guided Exploitation of Interpretable Attention Patterns in Summarization and Topic Segmentation

EMNLP 2022main

The multi-head self-attention mechanism of the transformer model has been thoroughly investigated recently. In one vein of study, researchers are interested in understanding why and how transformers work. In another vein, researchers propose new attention augmentation methods to make transformers mo…

2021

Finding Representative Interpretations on Convolutional Neural Networks

ICCV 2021poster

Interpreting the decision logic behind effective deep convolutional neural networks (CNN) on images complements the success of deep learning models. However, the existing methods can only interpret some specific decision logic on individual or a small number of images. To facilitate human understand…

Cited by 11PDFScholar
2021

Personalized Cross-Silo Federated Learning on Non-IID Data

AAAI 2021technical

Non-IID data present a tough challenge for federated learning. In this paper, we explore a novel idea of facilitating pairwise collaborations between clients with similar data. We propose FedAMP, a new method employing federated attentive message passing to facilitate similar clients to collaborate…

Cited by 744SourcePDFScholar
2021

Robust Counterfactual Explanations on Graph Neural Networks

NeurIPS 2021poster

Massive deployment of Graph Neural Networks (GNNs) in high-stake applications generates a strong demand for explanations that are robust to noise and align well with human intuition. Most existing methods generate explanations by identifying a subgraph of an input graph that has a strong correlation…

Cited by 139SourcePDFScholar
2021

T3-Vis: visual analytic for Training and fine-Tuning Transformers in NLP

EMNLP 2021system demonstrations

Transformers are the dominant architecture in NLP, but their training and fine-tuning is still very challenging. In this paper, we present the design and implementation of a visual analytic framework for assisting researchers in such process, by providing them with valuable insights about the model’…

2020

A High Precision Pipeline for Financial Knowledge Graph Construction

COLING 2020main

Motivated by applications such as question answering, fact checking, and data integration, there is significant interest in constructing knowledge graphs by extracting information from unstructured information sources, particularly text documents. Knowledge graphs have emerged as a standard for stru…

Cited by 48SourcePDFScholar