← Search

Yuan Yuan

79 accepted papers

2026

BLM-Guard: Explainable Multimodal Ad Moderation with Chain-of-Thought and Policy-Aligned Rewards

AAAI 2026technical

Short-video platforms now host vast multimodal ads whose deceptive visuals, speech and subtitles demand finer-grained, policy-driven moderation than community safety filters. We present BLM-Guard, a content-audit framework for commercial ads that fuses Chain-of-Thought reasoning with rule-based poli

Cited by 0SourcePDFScholar
2026

Beyond Model Ranking: Predictability-Aligned Evaluation for Time Series Forecasting

ICML 2026poster

In the era of increasingly complex AI models for time series forecasting, progress is often measured by marginal improvements on benchmark leaderboards. However, standard evaluations rely on aggregate metrics (e.g., MSE) that conflate model capability with the intrinsic difficulty of the evaluated i…

Cited by 0SourceScholar
2026

ConInstruct: Evaluating Large Language Models on Conflict Detection and Resolution in Instructions

AAAI 2026technical

Instruction-following is a critical capability of Large Language Models (LLMs). While existing works primarily focus on assessing how well LLMs adhere to user instructions, they often overlook scenarios where instructions contain conflicting constraints—a common occurrence in complex prompts. The be

Cited by 0SourcePDFScholar
2026

DGTF: Cross-Domain Decentralized Graph Learning with Topology-Aware Knowledge Fusion

AAAI 2026technical

Cross-Domain Decentralized Graph Learning (CD-DGL) is a promising paradigm that enables efficient, privacy-preserving collaboration among multiple parties to unlock the value of cross-domain graph data. However, it faces two fundamental challenges. First, inconsistent label spaces across domains dri

Cited by 0SourcePDFScholar
2026

DynaSchedBench: Calibrated Dynamic Scheduling Benchmarks and Observability Paradox in LLM-based Scheduling Agents

ICML 2026poster

Progress in neural combinatorial optimization for Dynamic Flexible Job Shop Scheduling Problem (DFJSP) is currently hindered by a methodological tension: static benchmarks encourage benchmark overfitting, while uncalibrated generators obscure algorithmic difficulty with stochastic noise. To resolve …

Cited by 0SourceScholar
2026

Forgetting Whenever You Want: A Decentralized Continual Learning Framework with On-Demand Unlearning

ICML 2026poster

Decentralized class continual learning refers to a paradigm where distributed clients continuously acquire new classes while retaining previously learned information without relying on a central server. With increasing emphasis on privacy preservation, there is a growing need for on-demand unlearnin…

Cited by 0SourceScholar
2026

FreeGaussian: Annotation-free Control of Articulated Objects via 3D Gaussian Splats with Flow Derivatives

AAAI 2026technical

Reconstructing controllable Gaussian splats for articulated objects from monocular video is especially challenging due to its inherently insufficient constraints. Existing methods address this by relying on dense masks and manually defined control signals, limiting their real-world applications. In

Cited by 0SourcePDFScholar
2026

Generative Adaptation of Dynamics to Environmental Shifts via Weight-space Diffusion

ICML 2026poster

Data-driven dynamics prediction often fails under environmental shifts, while traditional fine-tuning remains computationally prohibitive for hardware-constrained or data-scarce applications. We propose DynaDiff, a generative meta-learning framework that transitions the paradigm from gradient-based …

Cited by 0SourceScholar
2026

Harmonizing Real-Time Constraints and Long-Horizon Reasoning: An Asynchronous Agentic Framework for Dynamic Scheduling

IJCAI 2026

The Dynamic Flexible Job Shop Scheduling Problem (DFJSP) necessitates a trade-off between instant reaction to stochastic disturbances and global optimization of production goals. Conventional priority rules are insufficiently flexible to handle complex disruptions, whereas learning-based approaches

Cited by 0Scholar
2026

LoRAGen: Structure-Aware Weight Space Learning for LoRA Generation

ICLR 2026poster

The widespread adoption of Low-Rank Adaptation (LoRA) for efficient fine-tuning of large language models has created demand for scalable parameter generation methods that can synthesize adaptation weights directly from task descriptions, avoiding costly task-specific training. We present LoRAGen, a…

Cited by 0SourcecodeScholar
2026

VQ-Transplant: Efficient VQ-Module Integration for Pre-trained Visual Tokenizers

ICLR 2026poster

Vector Quantization (VQ) underpins modern discrete visual tokenization. However, training quantization modules for state-of-the-art VQ-based models requires significant computational resources which, in practice, all but prevents the development of novel, cutting-edge VQ techniques under resource co…

Cited by 0SourceScholar
2025

AdDriftBench: A Benchmark for Detecting Data Drift and Label Drift in Short Video Advertising

EMNLP 2025

With the commercialization of short video platforms (SVPs), the demand for compliance auditing of advertising content has grown rapidly. The rise of large vision-language models (VLMs) offers new opportunities for automating ad content moderation. However, short video advertising scenarios present u

Cited by 0SourcePDFScholar
2025

Adv-CPG: A Customized Portrait Generation Framework with Facial Adversarial Attacks

CVPR 2025poster

Recent Customized Portrait Generation (CPG) methods, taking a facial image and a textual prompt as inputs, have attracted substantial attention. Although these methods generate high-fidelity portraits, they fail to prevent the generated portraits from being tracked and misused by malicious face reco…

2025

ControlText: Unlocking Controllable Fonts in Multilingual Text Rendering without Font Annotations

EMNLP 2025

This work demonstrates that diffusion models can achieve font-controllable multilingual text rendering using just raw images without font label annotations. Visual text rendering remains a significant challenge. While recent methods condition diffusion on glyphs, it is impossible to retrieve exact f

2025

Diffusion Transformers as Open-World Spatiotemporal Foundation Models

NeurIPS 2025poster

The urban environment is characterized by complex spatio-temporal dynamics arising from diverse human activities and interactions. Effectively modeling these dynamics is essential for understanding and optimizing urban systems. In this work, we introduce UrbanDiT, a foundation model for open-world u…

Cited by 0SourcecodeScholar
2025

Edge Contrastive Learning: An Augmentation-Free Graph Contrastive Learning Model

AAAI 2025technical

Graph contrastive learning (GCL) aims to learn representations from unlabeled graph data in a self-supervised manner and has developed rapidly in recent years. However, edge-level contrasts are not well explored by most existing GCL methods. Most studies in GCL only regard edges as auxiliary informa…

2025

Enhancing Low-Rank Adaptation with Recoverability-Based Reinforcement Pruning for Object Counting

AAAI 2025technical

Object counting is crucial for understanding the distribution of objects in different scenarios. Recently, many object counting networks have been designed to be more complex to achieve marginal improvements, leading to excessive time spent on model design. With the development of large models (LMs)…

Cited by 0SourcePDFScholar
2025

GoRA: Gradient-driven Adaptive Low Rank Adaptation

NeurIPS 2025poster

Low-Rank Adaptation (LoRA) is a crucial method for efficiently fine-tuning large language models (LLMs), with its effectiveness influenced by two key factors: rank selection and weight initialization. While numerous LoRA variants have been proposed to improve performance by addressing one of these a…

Cited by 0SourcecodeScholar
2025

How Distributed Collaboration Influences the Diffusion Model Training? A Theoretical Perspective

ICML 2025poster

This paper examines the theoretical performance of distributed diffusion models in environments where computational resources and data availability vary significantly among workers. Traditional models centered on single-worker scenarios fall short in such distributed settings, particularly when some…

Cited by 0SourcePDFScholar
2025

Multilingual Retrieval Augmented Generation for Culturally-Sensitive Tasks: A Benchmark for Cross-lingual Robustness

ACL 2025finding

The paradigm of retrieval-augmented generated (RAG) helps mitigate hallucinations of large language models (LLMs). However, RAG also introduces biases contained within the retrieved documents. These biases can be amplified in scenarios which are multilingual and culturally-sensitive, such as territo…

Cited by 0SourcePDFScholar
2025

NightHaze: Nighttime Image Dehazing via Self-Prior Learning

AAAI 2025technical

Masked autoencoder (MAE) shows that severe augmentation during training produces robust representations for high-level tasks. This paper brings the MAE-like framework to nighttime image enhancement, demonstrating that severe augmentation during training produces strong network priors that are resili…

Cited by 5SourcePDFScholar
2025

PDUDT: Provable Decentralized Unlearning under Dynamic Topologies

ICML 2025poster

This paper investigates decentralized unlearning, aiming to eliminate the impact of a specific client on the whole decentralized system. However, decentralized communication characterizations pose new challenges for effective unlearning: the indirect connections make it difficult to trace the specif…

Cited by 0SourcePDFScholar
2025

Reinforcement Learning with Adaptive Reward Modeling for Expensive-to-Evaluate Systems

ICML 2025poster

Training reinforcement learning (RL) agents requires extensive trials and errors, which becomes prohibitively time-consuming in systems with costly reward evaluations. To address this challenge, we propose adaptive reward modeling (AdaReMo) which accelerates RL training by decomposing the complicate…

2025

SPEA: Large-Scale Entity Alignment via Self-Partitioning

ICASSP 2025accepted

The task of entity alignment (EA) seeks to identify corresponding entities across different knowledge graphs (KGs). However, in large-scale KG alignment tasks, the complexity of the problem renders traditional entity structure representation methods, designed for small-scale KGs, ineffective. Partit…

Cited by 0SourceScholar
2025

Semantic Segmentation on Raindrop Degraded Images Using Two-Stage Dual Teacher-Student Learning

AAAI 2025technical

Existing semantic segmentation methods face challenges when processing input images degraded by raindrops on the lens or windshield. Unlike other adverse conditions such as fog and nighttime, which degrade visual quality, raindrops not only impair visual appearances but also introduce misleading occ…

Cited by 0SourcePDFScholar
2025

Towards Rationality in Language and Multimodal Agents: A Survey

NAACL 2025long

This work discusses how to build more rational language and multimodal agents and what criteria define rationality in intelligent systems.Rationality is the quality of being guided by reason, characterized by decision-making that aligns with evidence and logical principles. It plays a crucial role i…

2025

TurnaboutLLM: A Deductive Reasoning Benchmark from Detective Games

EMNLP 2025

This paper introduces TurnaboutLLM, a novel framework and dataset for evaluating the deductive reasoning abilities of Large Language Models (LLMs) by leveraging the interactive gameplay of detective games Ace Attorney and Danganronpa. The framework tasks LLMs with identifying contradictions between

Cited by 0SourcePDFScholar
2024

DeS3: Adaptive Attention-Driven Self and Soft Shadow Removal Using ViT Similarity

AAAI 2024technical

Removing soft and self shadows that lack clear boundaries from a single image is still challenging. Self shadows are shadows that are cast on the object itself. Most existing methods rely on binary shadow masks, without considering the ambiguous boundaries of soft and self shadows. In this paper, we…

2024

Dual-Rain: Video Rain Removal using Assertive and Gentle Teachers

ECCV 2024poster

"Existing video deraining methods addressing both rain accumulation and rain streaks rely on synthetic data for training as clear ground-truths are unavailable. Hence, they struggle to handle real-world rain videos due to domain gaps. In this paper, we present Dual-Rain, a novel video deraining meth…

Cited by 4SourcePDFScholar
2024

End-to-End Video Semantic Segmentation in Adverse Weather using Fusion Blocks and Temporal-Spatial Teacher-Student Learning

NeurIPS 2024poster

Adverse weather conditions can significantly degrade the video frames, causing existing video semantic segmentation methods to produce erroneous predictions. In this work, we target adverse weather conditions and introduce an end-to-end domain adaptation strategy that leverages a fusion block, tempo…

Cited by 1SourcePDFScholar
2024

HEAP: Unsupervised Object Discovery and Localization with Contrastive Grouping

AAAI 2024technical

Unsupervised object discovery and localization aims to detect or segment objects in an image without any supervision. Recent efforts have demonstrated a notable potential to identify salient foreground objects by utilizing self-supervised transformer features. However, their scopes only build upon p…

Cited by 3SourcePDFScholar
2024

Improving Factual Error Correction by Learning to Inject Factual Errors

AAAI 2024technical

Factual error correction (FEC) aims to revise factual errors in false claims with minimal editing, making them faithful to the provided evidence. This task is crucial for alleviating the hallucination problem encountered by large language models. Given the lack of paired data (i.e., false claims and…

2024

MetaISP: Efficient RAW-to-sRGB Mappings with Merely 1M Parameters

IJCAI 2024poster

State-of-the-art deep ISP models alleviate the dilemma of limited generalization capabilities across heterogeneous inputs by increasing the size and complexity of the network, which inevitably leads to considerable growth in parameter counts and FLOPs. To address this challenge, this paper presents…

Cited by 0SourcePDFScholar
2024

NightRain: Nighttime Video Deraining via Adaptive-Rain-Removal and Adaptive-Correction

AAAI 2024technical

Existing deep-learning-based methods for nighttime video deraining rely on synthetic data due to the absence of real-world paired data. However, the intricacies of the real world, particularly with the presence of light effects and low-light regions affected by noise, create significant domain gaps,…

Cited by 12SourcePDFScholar
2024

Semantic Segmentation in Multiple Adverse Weather Conditions with Domain Knowledge Retention

AAAI 2024technical

Semantic segmentation's performance is often compromised when applied to unlabeled adverse weather conditions. Unsupervised domain adaptation is a potential approach to enhancing the model's adaptability and robustness to adverse weather. However, existing methods encounter difficulties when sequent…

Cited by 4SourcePDFScholar
2024

Spatio-Temporal Few-Shot Learning via Diffusive Neural Network Generation

ICLR 2024poster

Spatio-temporal modeling is foundational for smart city applications, yet it is often hindered by data scarcity in many cities and regions. To bridge this gap, we propose a novel generative pre-training framework, GPD, for spatio-temporal few-shot learning with urban knowledge transfer. Unlike conve…

2023

PivotFEC: Enhancing Few-shot Factual Error Correction with a Pivot Task Approach using Large Language Models

EMNLP 2023long findings

Factual Error Correction (FEC) aims to rectify false claims by making minimal revisions to align them more accurately with supporting evidence. However, the lack of datasets containing false claims and their corresponding corrections has impeded progress in this field. Existing distantly supervised…

Cited by 0SourceScholar
2023

Weakly-Supervised Scene-Specific Crowd Counting Using Real-Synthetic Hybrid Data

ICASSP 2023accepted

Due to the domain gap between the public large-scale datasets and actual scenes, the crowd counting models trained on the common datasets have a significant performance degradation when applying in practical applications. To address the above issue, one of the solution is to label additional data fr…

Cited by 0SourceScholar
2022

BiP-Net: Bidirectional Perspective Strategy Based Arbitrary-Shaped Text Detection Network

ICASSP 2022accepted

Detecting irregular-shaped text instances is the main challenge for text detection. Existing approaches can be roughly divided into top-down and bottom-up perspective methods. The former encodes text contours into unified units, which always fails to fit highly curved text contours. The latter repre…

Cited by 0SourceScholar
2022

EMGC²F: Efficient Multi-view Graph Clustering with Comprehensive Fusion

IJCAI 2022poster

This paper proposes an Efficient Multi-view Graph Clustering with Comprehensive Fusion (EMGC²F) model and a corresponding efficient optimization algorithm to address multi-view graph clustering tasks effectively and efficiently. Compared to existing works, our proposals have the following highlights…

Cited by 14SourcePDFScholar
2022

Self-Supervised Transformers for Unsupervised Object Discovery Using Normalized Cut

CVPR 2022poster

Transformers trained with self-supervision using self-distillation loss (DINO) have been shown to produce attention maps that highlight salient foreground objects. In this paper, we show a graph-based method that uses the self-supervised transformer features to discover an object from an image. Visu…

Cited by 194PDFScholar
2022

Targeted Supervised Contrastive Learning for Long-Tailed Recognition

CVPR 2022poster

Real-world data often exhibits long tail distributions with heavy class imbalance, where the majority classes can dominate the training process and alter the decision boundaries of the minority classes. Recently, researchers have investigated the potential of supervised contrastive learning for long…

Cited by 249PDFcodeScholar
2021

A Knowledge-Based Fast Motion Planning Method Through Online Environmental Feature Learning

ICRA 2021poster

The sampling-based partial motion planning algorithm has come into widespread application in dynamic mobile robot navigation due to its low calculation costs and excellent performance in avoiding obstacles. However, when confronted with complicated scenarios, the motion planning algorithms are easil…

Cited by 10SourceScholar
2020

Efficient Dynamic Scene Deblurring Using Spatially Variant Deconvolution Network With Optical Flow Guided Training

CVPR 2020poster

In order to remove the non-uniform blur of images captured from dynamic scenes, many deep learning based methods design deep networks for large receptive fields and strong fitting capabilities, or use multi-scale strategy to deblur image on different scales gradually. Restricted by the fixed structu…

Cited by 121PDFScholar
2020

Enhanced Non-Local Cascading Network with Attention Mechanism for Hyperspectral Image Denoising

ICASSP 2020accepted

Because of the complexity of imaging environment, hyper-spectral remote sensing images (HSIs) often suffer from different kinds of noise. Despite the success in natural image denoising, most of the existing CNN-based HSIs denoising methods still suffer from the problem of inadequate noise suppressio…

Cited by 0SourceScholar
2020

Learning Longterm Representations for Person Re-Identification Using Radio Signals

CVPR 2020poster

Person Re-Identification (ReID) aims to recognize a person-of-interest across different places and times. Existing ReID methods rely on images or videos collected using RGB cameras. They extract appearance features like clothes, shoes, hair, etc. Such features, however, can change drastically from o…

Cited by 124PDFScholar
2020

Unsupervised Semantic Aggregation and Deformable Template Matching for Semi-Supervised Learning

NeurIPS 2020poster

Unlabeled data learning has attracted considerable attention recently. However, it is still elusive to extract the expected high-level semantic feature with mere unsupervised learning. In the meantime, semi-supervised learning (SSL) demonstrates a promising future in leveraging few samples. In this…

2019

MARGINALIZED AVERAGE ATTENTIONAL NETWORK FOR WEAKLY-SUPERVISED LEARNING

ICLR 2019poster

In weakly-supervised temporal action localization, previous works have failed to locate dense and integral regions for each entire action due to the overestimation of the most salient regions. To alleviate this issue, we propose a marginalized average attentional network (MAAN) to suppress the domin…

Cited by 107SourcePDFScholar
2017

Embedding structured contour and location prior in siamesed fully convolutional networks for road detection

ICRA 2017poster

Road detection from the perspective of moving vehicles is a challenging issue in autonomous driving. Recently, many deep learning methods spring up for this task because they can extract high-level local features to find road regions from raw RGB data, such as Convolutional Neural Networks (CNN) and…

Cited by 307SourceScholar
2017

Temporal Dynamic Graph LSTM for Action-Driven Video Object Detection

ICCV 2017poster

In this paper, we investigate a weakly-supervised object detection framework. Most existing frameworks focus on using static images to learn object detectors. However, these detectors often fail to generalize to videos because of the existing domain shift. Therefore, we investigate learning these de…

Cited by 108PDFcodeScholar