← Search

Dongyu Zhang

24 accepted papers

2026

BrainCGT: A Brain Graph Transformer for Modeling Causal Connectivity in Neurological Disorder Diagnosis

IJCAI 2026

Brain connectivity analysis is a fundamental tool for identifying biomarkers and understanding of neurological disorders. Most existing approaches employ graph transformers over undirected functional connectivity networks, which are typically estimated using correlation statistics. Although effectiv

Cited by 0Scholar
2026

DEIG: Detail-Enhanced Instance Generation with Fine-Grained Semantic Control

AAAI 2026technical

Multi-Instance Generation has advanced significantly in spatial placement and attribute binding. However, existing approaches still face challenges in fine-grained semantic understanding, particularly when dealing with complex textual descriptions.To overcome these limitations, we propose DEIG, a no

Cited by 0SourcePDFScholar
2026

Learning Heterogeneous Degradation Representation for Real-World Super-Resolution

ICLR 2026poster

Real-World Super-Resolution (RWSR) aims to reconstruct high-resolution images from low-resolution inputs captured under complex, real-life conditions, where diverse distortions result in significant degradation heterogeneity. Many methods rely on degradation representations, yet they struggle with t…

Cited by 0SourceScholar
2026

TAPE: Task-Adaptive Prototype Evolution in Audio-Language Models for Fully Few-shot Class-incremental Audio Classification

CVPR 2026

Fully Few-shot Class-incremental Audio Classification (FFCAC) is challenging since the training samples are limited both in the incremental sessions and in the base session. Existing few-shot learning methods suffer from catastrophic forgetting and overfitting when applied to FFCAC.Pre-trained Audio

Cited by 0SourcecodeScholar
2025

Bold Claims or Self-Doubt? Factuality Hallucination Type Detection via Belief State

EMNLP 2025

Large language models are prone to generating hallucination that deviates from factual information. Existing studies mainly focus on detecting the presence of hallucinations but lack a systematic classification approach, which hinders deeper exploration of their characteristics. To address this, we

Cited by 0SourcePDFScholar
2025

Cultural Bias Matters: A Cross-Cultural Benchmark Dataset and Sentiment-Enriched Model for Understanding Multimodal Metaphors

ACL 2025long

Metaphors are pervasive in communication, making them crucial for natural language processing (NLP). Previous research on automatic metaphor processing predominantly relies on training data consisting of English samples, which often reflect Western European or North American biases. This cultural sk…

2025

LLDB: Efficient Low-Light Image Enhancement with Difffusion Bridge

ICASSP 2025accepted

This paper investigates a low-light image enhancement method based on the diffusion bridge framework. Currently, low-light image enhancement tasks still face challenges in noise reduction and detail restoration, and existing diffusion model methods are time-consuming and have unstable diffusion proc…

Cited by 0SourceScholar
2025

ROSE: A Reward-Oriented Data Selection Framework for LLM Task-Specific Instruction Tuning

EMNLP 2025

Instruction tuning has underscored the significant potential of large language models (LLMs) in producing more human controllable and effective outputs in various domains. In this work, we focus on the data selection problem for task-specific instruction tuning of LLMs. Prevailing methods primarily

2025

Rhetorical Device-Aware Sarcasm Detection with Counterfactual Data Augmentation

ACL 2025finding

Sarcasm is a complex form of sentiment expression widely used in human daily life. Previous work primarily defines sarcasm as a form of verbal irony, which covers only a subset of real-world sarcastic expressions. However, sarcasm serves multifaceted functions and manifests itself through various rh…

2025

SpeechHGT: A Multimodal Hypergraph Transformer for Speech-Based Early Alzheimer’s Disease Detection

IJCAI 2025

Early detection of Alzheimer's disease (AD) through spontaneous speech analysis represents a promising, non-invasive diagnostic approach. Existing methods predominantly rely on fusion-based multimodal deep learning, effectively integrating linguistic and acoustic features. However, these methods ina

2025

SymGaussian: Occluded Human Rendering with Multi-scale Symmetry Feature from Monocular Video

ICASSP 2025accepted

The growing demand for high-quality 3D human rendering in real-world applications highlights significant challenges. These challenges are particularly evident in dealing with occlusion in monocular video. Previous methods often rely on controlled datasets and overlook the inherent symmetry of the hu…

Cited by 0SourceScholar
2024

EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoE

AAAI 2024technical

Building scalable vision-language models to learn from diverse, multimodal data remains an open challenge. In this paper, we introduce an Efficient Vision-languagE foundation model, namely EVE, which is one unified multimodal Transformer pre-trained solely by one unified pre-training task. Specifica…

Cited by 11SourcePDFScholar
2024

Image Restoration Through Generalized Ornstein-Uhlenbeck Bridge

ICML 2024poster

Diffusion models exhibit powerful generative capabilities enabling noise mapping to data via reverse stochastic differential equations. However, in image restoration, the focus is on the mapping relationship from low-quality to high-quality images. Regarding this issue, we introduce the Generalized…

2024

Take Its Essence, Discard Its Dross! Debiasing for Toxic Language Detection via Counterfactual Causal Effect

COLING 2024main

Researchers have attempted to mitigate lexical bias in toxic language detection (TLD). However, existing methods fail to disentangle the “useful” and “misleading” impact of lexical bias on model decisions. Therefore, they do not effectively exploit the positive effects of the bias and lead to a degr…

2024

Towards Comprehensive Detection of Chinese Harmful Memes

NeurIPS 2024poster

Harmful memes have proliferated on the Chinese Internet, while research on detecting Chinese harmful memes significantly lags behind due to the absence of reliable datasets and effective detectors. To this end, we present the comprehensive detection of Chinese harmful memes. We introduce ToxiCN MM,…

2023

MultiCMET: A Novel Chinese Benchmark for Understanding Multimodal Metaphor

EMNLP 2023long findings

Metaphor is a pervasive aspect of human communication, and its presence in multimodal forms has become more prominent with the progress of mass media. However, there is limited research on multimodal metaphor resources beyond the English language. Furthermore, the existing work in natural language p…

Cited by 0SourceScholar
2022

Enhancing Prototypical Few-Shot Learning By Leveraging The Local-Level Strategy

ICASSP 2022accepted

Aiming at recognizing the samples from novel categories with few reference samples, few-shot learning (FSL) is a challenging problem. We found that the existing works often build their few-shot model based on the image-level feature by mixing all local-level features, which leads to the discriminati…

Cited by 0SourceScholar
2021

MultiMET: A Multimodal Dataset for Metaphor Understanding

ACL 2021long

Metaphor involves not only a linguistic phenomenon, but also a cognitive phenomenon structuring human thought, which makes understanding it challenging. As a means of cognition, metaphor is rendered by more than texts alone, and multimodal information in which vision/audio content is integrated with…

Cited by 56SourcePDFScholar
2020

Transferable, Controllable, and Inconspicuous Adversarial Attacks on Person Re-identification With Deep Mis-Ranking

CVPR 2020oral

The success of DNNs has driven the extensive applications of person re-identification (ReID) into a new era. However, whether ReID inherits the vulnerability of DNNs remains unexplored. To examine the robustness of ReID systems is rather important because the insecurity of ReID systems may cause sev…

Cited by 106PDFcodeScholar
2018

Towards Human-Machine Cooperation: Self-Supervised Sample Mining for Object Detection

CVPR 2018poster

Though quite challenging, leveraging large-scale unlabeled or partially labeled images in a cost-effective way has increasingly attracted interests for its great importance to computer vision. To tackle this problem, many Active Learning (AL) methods have been developed. However, these methods mainl…

Cited by 136SourcePDFScholar
2017

Look Into Person: Self-Supervised Structure-Sensitive Learning and a New Benchmark for Human Parsing

CVPR 2017poster

Human parsing has recently attracted a lot of research interests due to its huge application potentials. However existing datasets have limited number of images and annotations, and lack the variety of human appearances and the coverage of challenging cases in unconstrained environment. In this pape…

Cited by 621PDFcodeScholar