← Search

yan yang

45 accepted papers

2026

ALERT: Adversarial Learning Enhanced Stability-aware Routing Transformer for Adaptive Depression Detection

AAAI 2026technical

Detecting depression through social media is a complex task, as noisy user-generated content creates significant interference between persistent depressive patterns and transient emotional expressions. Two main challenges arise: First, negative mood indicators are not exclusive to depressed individu

Cited by 1SourcePDFScholar
2026

Multi-Agent Collaboration for PrSTL Specifications With Temporal Collective Counting Operators

RA-L 2026

We address the collaborative path planning problem for multi-agent systems with heterogeneous capabilities, subject to uncertainty and operating under complex task specifications. Conventional Probabilistic Signal Temporal Logic (PrSTL) frameworks exhibit significant limitations in describing multi-

Cited by 0SourceScholar
2026

Multi-Agent Collaboration for PrSTL Specifications with Temporal Collective Counting Operators

ICRA 2026poster

We address the collaborative path planning problem for multi-agent systems with heterogeneous capabilities, subject to uncertainty and operating under complex task specifications. Conventional Probabilistic Signal Temporal Logic (PrSTL) frameworks exhibit significant limitations in describing multi-…

Cited by 0SourceScholar
2026

Sparse4DGS: 4D Gaussian Splatting for Sparse-Frame Dynamic Scene Reconstruction

AAAI 2026technical

Dynamic Gaussian Splatting approaches have achieved remarkable performance for 4D scene reconstruction. However, these approaches rely on dense-frame video sequences for photorealistic reconstruction. In real-world scenarios, due to equipment constraints, sometimes only sparse frames are accessible.

Cited by 0SourcePDFScholar
2026

Zero-Shot Image Denoising via Hybrid Prior-Guided Pseudo Sample Generation

CVPR 2026

Zero-shot image denoising has gained prominence in recent years, as it inherently relies on the intrinsic priors of images rather than learning from external data. Nevertheless, most existing methods either fail to fully exploit global priors, or do not properly preserve the fine-grained details gov

Cited by 0SourceScholar
2025

Bilevel Reinforcement Learning via the Development of Hyper-gradient without Lower-Level Convexity

AISTATS 2025poster

Bilevel reinforcement learning (RL), which features intertwined two-level problems, has attracted growing interest recently. The inherent non-convexity of the lower-level RL problem is, however, to be an impediment to developing bilevel optimization methods. By employing the fixed point equation ass…

Cited by 0SourceScholar
2025

Exploring Temporal Event Cues for Dense Video Captioning in Cyclic Co-Learning

AAAI 2025technical

Dense video captioning aims to detect and describe all events in untrimmed videos. This paper presents a dense video captioning network called Multi-Concept Cyclic Learning (MCCL), which aims to: (1) detect multiple concepts at the frame level and leverage these concepts to provide temporal event cu…

Cited by 0SourcePDFScholar
2025

Growing a Twig to Accelerate Large Vision-Language Models

ICCV 2025poster

Large vision-language models (VLMs) have demonstrated remarkable capabilities in open-world multimodal understanding, yet their high computational overheads pose great challenges for practical deployment. Some recent works have proposed methods to accelerate VLMs by pruning redundant visual tokens g…

2025

ImPart: Importance-Aware Delta-Sparsification for Improved Model Compression and Merging in LLMs

ACL 2025long

With the proliferation of task-specific large language models, delta compression has emerged as a method to mitigate the resource challenges of deploying numerous such models by effectively compressing the delta model parameters. Previous delta-sparsification methods either remove parameters randoml…

2025

Long-Range Multi-Scale Fusion for Efficient Single Image Super-Resolution

ICASSP 2025accepted

Improving the performance of single image super-resolution (SISR) via extending the effective receptive field (ERF) of the model has become an admired paradigm in the field due to the universal self-similarity prior of natural images. However, it cannot fully explore model capability by solely incre…

Cited by 0SourceScholar
2025

Pi-SQL: Enhancing Text-to-SQL with Fine-Grained Guidance from Pivot Programming Languages

EMNLP 2025

Text-to-SQL transforms the user queries from natural language to executable SQL programs, enabling non-experts to interact with complex databases. Existing prompt-based methods craft meticulous text guidelines and examples to facilitate SQL generation, but their accuracy is hindered by the large sem

Cited by 0SourcePDFScholar
2025

ProBench: Judging Multimodal Foundation Models on Open-ended Multi-domain Expert Tasks

ACL 2025finding

Solving expert-level multimodal tasks is a key milestone in general intelligence. As the capabilities of multimodal large language models (MLLMs) continue to evolve, evaluation of frontier multimodal intelligence becomes necessary yet challenging. In this work, we introduce ProBench, a benchmark of…

Cited by 0SourcePDFScholar
2025

ProTOD: Proactive Task-oriented Dialogue System Based on Large Language Model

COLING 2025main

Large Language Model (LLM)-based Task-Oriented Dialogue (TOD) systems show promising performance in helping users achieve specific goals in a zero-shot setting. However, existing systems engage with users in a reactive manner, relying on a basic single-query mechanism with the knowledge base and emp…

2025

SeqAR: Jailbreak LLMs with Sequential Auto-Generated Characters

NAACL 2025long

The widespread applications of large language models (LLMs) have brought about concerns regarding their potential misuse. Although aligned with human preference data before release, LLMs remain vulnerable to various malicious attacks. In this paper, we adopt a red-teaming strategy to enhance LLM saf…

2025

Storyboard-guided Alignment for Fine-grained Video Action Recognition

NeurIPS 2025poster

Fine-grained video action recognition can be formulated as a video–text matching problem. Previous approaches primarily rely on global video semantics to consolidate video embeddings, often leading to misaligned video–text pairs due to inaccurate atomic-level action understanding. This inaccuracy ar…

Cited by 0SourceScholar
2025

Towards Prospective Medical Image Reconstruction via Knowledge-Informed Dynamic Optimal Transport

NeurIPS 2025poster

Medical image reconstruction from measurement data is a vital but challenging inverse problem. Deep learning approaches have achieved promising results, but often requires paired measurement and high-quality images, which is typically simulated through a forward model, i.e., retrospective reconstruc…

Cited by 0SourcecodeScholar
2024

Distract Large Language Models for Automatic Jailbreak Attack

EMNLP 2024main

Extensive efforts have been made before the public release of Large language models (LLMs) to align their behaviors with human values. However, even meticulously aligned LLMs remain vulnerable to malicious manipulations such as jailbreaking, leading to unintended behaviors. In this work, we propose…

2024

Dynamic Replay Training for Class-Incremental Learning

ICASSP 2024accepted

Replay-based methods for Class-Incremental Learning (CIL) typically employ new classes and a limited subset of old classes stored in memory to facilitate the model training. However, these methods often lead to class imbalance and catastrophic forgetting, where the model forgets previously learned t…

Cited by 0SourceScholar
2024

Event-based Few-shot Fine-grained Human Action Recognition

IROS 2024poster

Few-shot fine-grained human (FGH) action recognition is crucial in the context of human-robot interaction within open-set real-world environments. Existing works mainly focus on features extracted from RGB frames. However, their performances are drastically impacted in challenging scenarios, such as…

Cited by 1SourceScholar
2024

GTMGC: Using Graph Transformer to Predict Molecule’s Ground-State Conformation

ICLR 2024spotlight

The ground-state conformation of a molecule is often decisive for its properties. However, experimental or computational methods, such as density functional theory (DFT), are time-consuming and labor-intensive for obtaining this conformation. Deep learning (DL) based molecular representation learnin…

Cited by 5SourcePDFScholar
2024

HeterGCL: Graph Contrastive Learning Framework on Heterophilic Graph

IJCAI 2024poster

Graph Contrastive Learning (GCL) has attracted significant research attention due to its self-supervised ability to learn robust node representations. Unfortunately, most methods primarily focus on homophilic graphs, rendering them less effective for heterophilic graphs. In addition, the complexity…

2024

LDP: Language-driven Dual-Pixel Image Defocus Deblurring Network

CVPR 2024poster

Recovering sharp images from dual-pixel (DP) pairs with disparity-dependent blur is a challenging task. Existing blur map-based deblurring methods have demonstrated promising results. In this paper we propose to the best of our knowledge the first framework to introduce the contrastive language-imag…

Cited by 12SourcePDFScholar
2024

Plan, Generate and Complicate: Improving Low-resource Dialogue State Tracking via Easy-to-Difficult Zero-shot Data Augmentation

ACL 2024findings

Data augmentation methods have been a promising direction to improve the performance of small models for low-resource dialogue state tracking. However, traditional methods rely on pre-defined user goals and neglect the importance of data complexity in this task. In this paper, we propose EDZ-DA, an…

2023

K3DN: Disparity-Aware Kernel Estimation for Dual-Pixel Defocus Deblurring

CVPR 2023poster

The dual-pixel (DP) sensor captures a two-view image pair in a single snapshot by splitting each pixel in half. The disparity occurs in defocus blurred regions between the two views of the DP pair, while the in-focus sharp regions have zero disparity. This motivates us to propose a K3DN framework fo…

Cited by 12SourcePDFScholar
2023

Learning a Simple Low-Light Image Enhancer From Paired Low-Light Instances

CVPR 2023poster

Low-light Image Enhancement (LIE) aims at improving contrast and restoring details for images captured in low-light conditions. Most of the previous LIE algorithms adjust illumination using a single input image with several handcrafted priors. Those solutions, however, often fail in revealing image…

2023

Only a Few Classes Confusing: Pixel-Wise Candidate Labels Disambiguation for Foggy Scene Understanding

AAAI 2023technical

Not all semantics become confusing when deploying a semantic segmentation model for real-world scene understanding of adverse weather. The true semantics of most pixels have a high likelihood of appearing in the few top classes according to confidence ranking. In this paper, we replace the one-hot p…

Cited by 9SourcePDFScholar
2022

Understanding Gender Bias in Knowledge Base Embeddings

ACL 2022long

Knowledge base (KB) embeddings have been shown to contain gender biases. In this paper, we study two questions regarding these biases: how to quantify them, and how to trace their origins in KB? Specifically, first, we develop two novel bias measures respectively for a group of person entities and a…

Cited by 10SourcePDFScholar
2022

Unsupervised Underwater Image Restoration: From a Homology Perspective

AAAI 2022technical

Underwater images suffer from degradation due to light scattering and absorption. It remains challenging to restore such degraded images using deep neural networks since real-world paired data is scarcely available while synthetic paired data cannot approximate real-world data perfectly. In this pap…

2021

Adversarial Reweighting for Partial Domain Adaptation

NeurIPS 2021poster

Partial domain adaptation (PDA) has gained much attention due to its practical setting. The current PDA methods usually adapt the feature extractor by aligning the target and reweighted source domain distributions. In this paper, we experimentally find that the feature adaptation by the reweighted d…

2021

KERS: A Knowledge-Enhanced Framework for Recommendation Dialog Systems with Multiple Subgoals

EMNLP 2021finding

Recommendation dialogs require the system to build a social bond with users to gain trust and develop affinity in order to increase the chance of a successful recommendation. It is beneficial to divide up, such conversations with multiple subgoals (such as social chat, question answering, recommenda…

2020

Bayes-enhanced Lifelong Attention Networks for Sentiment Classification

COLING 2020main

The classic deep learning paradigm learns a model from the training data of a single task and the learned model is also tested on the same task. This paper studies the problem of learning a sequence of tasks (sentiment classification tasks in our case). After each sentiment classification task is le…

Cited by 8SourcePDFScholar
2017

A distributed approach to automated manufacturing systems with complex structures using Petri nets

ICRA 2017poster

One of the major challenges from both a theoretical and practical perspectives, for effectively establishing unattended operation of automated manufacturing systems (AMSs), is to resolve the deadlock. In the existing methods on deadlock problem, most of them are focused on the models with either fle…

Cited by 3SourceScholar
2016

Distributed supervisor synthesis for automated manufacturing systems with flexible routes and assembly operations using Petri nets

ICRA 2016

Automated manufacturing systems (AMSs) are developing rapidly with increasingly sophisticated operations and complex topologies. In this paper, we propose a new kind of AMS structure, namely, AESMs, with processes expressed by flexible routes and embedded by marked graph blocks. Flexible routes and

Cited by 2SourceScholar
2015

Supervisor design and simplification for Automated Manufacturing Systems using colored Petri nets

ICRA 2015poster

Colored Petri nets are widely used to model Automated Manufacturing Systems thanks to their compactness to describe complex networked systems. Compared to general Petri nets, they allow many folding techniques so as to condense the system model. With them, many control synthesis problems are reduced…

Cited by 8SourceScholar
2015

Supervisors and their simplification in automated manufacturing systems via Petri nets

ICRA 2015poster

Most contemporary manufacturing systems appear as complex event-driven automation facilities. Supervisor synthesis and simplification are fundamental in automated manufacturing systems (AMSs). From design and implementation standpoints, it is preferable to decrease supervisor scales so as to mitigat…

Cited by 2SourceScholar