← Search

Wen Zhang

72 accepted papers

2026

PosterIQ: A Design Perspective Benchmark for Poster Understanding and Generation

CVPR 2026

We present PosterIQ, a design-driven benchmark for poster understanding and generation, annotated across composition structure, typographic hierarchy, and semantic intent. It includes 7,765 image-annotation instances and 822 generation prompts spanning real, professional, and synthetic cases. To bri

Cited by 0SourcecodeScholar
2026

Self-Correction Distillation for Structured Data Question Answering

AAAI 2026technical

Structured data question answering (QA), including table QA, Knowledge Graph (KG) QA, and temporal KG QA, is a pivotal research area. Advances in large language models (LLMs) have driven significant progress in unified structural QA frameworks like TrustUQA. However, these frameworks face

Cited by 0SourcePDFScholar
2026

Supervise Less, See More: Training-free Nuclear Instance Segmentation with Prototype-Guided Prompting

ICML 2026poster

Accurate nuclear instance segmentation is a pivotal task in computational pathology, supporting data-driven clinical insights and facilitating downstream translational applications. While large vision foundation models have shown promise for zero-shot biomedical segmentation, most existing approache…

Cited by 0SourceScholar
2026

UniHR: Hierarchical Representation Learning for Unified Knowledge Graph Link Prediction

AAAI 2026technical

Real-world knowledge graphs (KGs) contain not only standard triple-based facts, but also more complex, heterogeneous types of facts, such as hyper-relational facts with auxiliary key-value pairs, temporal facts with additional timestamps, and nested facts that imply relationships between facts. Thes

Cited by 0SourcePDFScholar
2026

Unveiling the Landscape of Clinical Depression Assessment: From Behavioral Signatures to Psychiatric Reasoning

AAAI 2026technical

Depression is a widespread mental disorder that affects millions worldwide. While automated depression assessment shows promise, most studies rely on limited or non-clinically validated data, and often prioritize complex model design over real-world effectiveness. In this paper, we aim to unveil the

Cited by 0SourcePDFScholar
2025

Beyond Completion: A Foundation Model for General Knowledge Graph Reasoning

ACL 2025finding

In natural language processing (NLP) and computer vision (CV), the successful application of foundation models across diverse tasks has demonstrated their remarkable potential. However, despite the rich structural and textual information embedded in knowledge graphs (KGs), existing research of found…

2025

Enrich-on-Graph: Query-Graph Alignment for Complex Reasoning with LLM Enriching

EMNLP 2025

Large Language Models (LLMs) exhibit strong reasoning capabilities in complex tasks. However, they still struggle with hallucinations and factual errors in knowledge-intensive scenarios like knowledge graph question answering (KGQA). We attribute this to the semantic gap between structured knowledge

2025

Have We Designed Generalizable Structural Knowledge Promptings? Systematic Evaluation and Rethinking

ACL 2025long

Large language models (LLMs) have demonstrated exceptional performance in text generation within current NLP research. However, the lack of factual accuracy is still a dark cloud hanging over the LLM skyscraper. Structural knowledge prompting (SKP) is a prominent paradigm to integrate external knowl…

2025

Integrating Drug Substructures and Longitudinal Electronic Health Records for Personalized Drug Recommendation

NeurIPS 2025poster

Drug recommendation systems aim to identify optimal drug combinations for patient care, balancing therapeutic efficacy and safety. Advances in large-scale longitudinal EHRs have enabled learning-based approaches that leverage patient histories such as diagnoses, procedures, and previously prescribed…

Cited by 0SourceScholar
2025

K-ON: Stacking Knowledge on the Head Layer of Large Language Model

AAAI 2025technical

Recent advancements in large language models (LLMs) have significantly improved various natural language processing (NLP) tasks. Typically, LLMs are trained to predict the next token, aligning well with many NLP tasks. However, in knowledge graph (KG) scenarios, entities are the fundamental units an…

Cited by 0SourcePDFScholar
2025

Knowledge Graph Pooling and Unpooling for Concept Abstraction

COLING 2025main

Knowledge graph embedding (KGE) aims to embed entities and relations as vectors in a continuous space and has proven to be effective for KG tasks. Recently, graph neural networks (GNN) based KGEs gain much attention due to their strong capability of encoding complex graph structures. However, most G…

Cited by 0SourcePDFScholar
2025

MAGI: Multi-Agent Guided Interview for Psychiatric Assessment

ACL 2025finding

Automating structured clinical interviews could revolutionize mental healthcare accessibility, yet existing large language models (LLMs) approaches fail to align with psychiatric diagnostic protocols. We present MAGI, the first framework that transforms the gold-standard Mini International Neuropsyc…

Cited by 0SourcePDFScholar
2025

Multiple Heads are Better than One: Mixture of Modality Knowledge Experts for Entity Representation Learning

ICLR 2025poster

Learning high-quality multi-modal entity representations is an important goal of multi-modal knowledge graph (MMKG) representation learning, which can en- hance reasoning tasks within the MMKGs, such as MMKG completion (MMKGC). The main challenge is to collaboratively model the structural informatio…

2025

Noise-powered Multi-modal Knowledge Graph Representation Framework

COLING 2025main

The rise of Multi-modal Pre-training highlights the necessity for a unified Multi-Modal Knowledge Graph (MMKG) representation learning framework. Such a framework is essential for embedding structured knowledge into multi-modal Large Language Models effectively, alleviating issues like knowledge mis…

2025

PKAG-DDI: Pairwise Knowledge-Augmented Language Model for Drug-Drug Interaction Event Text Generation

ACL 2025long

Drug-drug interactions (DDIs) arise when multiple drugs are administered concurrently. Accurately predicting the specific mechanisms underlying DDIs (named DDI events or DDIEs) is critical for the safe clinical use of drugs. DDIEs are typically represented as textual descriptions. However, most comp…

2025

RTQA : Recursive Thinking for Complex Temporal Knowledge Graph Question Answering with Large Language Models

EMNLP 2025

Current temporal knowledge graph question answering (TKGQA) methods primarily focus on implicit temporal constraints, lacking the capability to handle more complex temporal queries, and struggle with limited reasoning abilities and error propagation in decomposition frameworks. We propose RTQA, a no

2025

Radiation and Directivity Analysis of a Vibrating Dome-Shaped Radiator Mounted on an Infinite Baffle

ICASSP 2025accepted

Accurate modeling and analysis of a radiator mounted on an infinite baffle are crucial for understanding its acoustic radiation characteristics. This paper investigates the radiation behavior of a convex dome-shaped radiator in such a condition, showing that, under the far-field approximation, the p…

Cited by 0SourceScholar
2025

ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning

NeurIPS 2025poster

Large Language Models (LLMs) have shown remarkable capabilities in reasoning, exemplified by the success of OpenAI-o1 and DeepSeek-R1. However, integrating reasoning with external search processes remains challenging, especially for complex multi-hop questions requiring multiple retrieval steps. We…

Cited by 0SourceScholar
2025

SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs

EMNLP 2025

Although large language models (LLMs) have made significant progress in understanding Structured Knowledge (SK) like KG and Table, existing evaluations for SK understanding are non-rigorous (i.e., lacking evaluations of specific capabilities) and focus on a single type of SK. Therefore, we aim to pr

2025

Tokenization, Fusion, and Augmentation: Towards Fine-grained Multi-modal Entity Representation

AAAI 2025technical

Multi-modal knowledge graph completion (MMKGC) aims to discover unobserved knowledge from given multi-modal knowledge graphs (MMKG), collaboratively leveraging structural information from the triples and multi-modal information of the entities to overcome the inherent incompleteness. Existing MMKGC…

2025

TrustUQA: A Trustful Framework for Unified Structured Data Question Answering

AAAI 2025technical

Natural language question answering (QA) over structured data sources such as tables and knowledge graphs have been widely investigated, especially with Large Language Models (LLMs) in recent years. The main solutions include question to formal query parsing and retrieval-based answer generation. Ho…

2024

A Multi-Modal Contrastive Diffusion Model for Therapeutic Peptide Generation

AAAI 2024technical

Therapeutic peptides represent a unique class of pharmaceutical agents crucial for the treatment of human diseases. Recently, deep generative models have exhibited remarkable potential for generating therapeutic peptides, but they only utilize sequence or structure information alone, which hinders t…

2024

Heterogeneous Causal Metapath Graph Neural Network for Gene-Microbe-Disease Association Prediction

IJCAI 2024poster

The recent focus on microbes in human medicine highlights their potential role in the genetic framework of diseases. To decode the complex interactions among genes, microbes, and diseases, computational predictions of gene-microbe-disease (GMD) associations are crucial. Existing methods primarily ad…

2024

Improving PTM Site Prediction by Coupling of Multi-Granularity Structure and Multi-Scale Sequence Representation

AAAI 2024technical

Protein post-translational modification (PTM) site prediction is a fundamental task in bioinformatics. Several computational methods have been developed to predict PTM sites. However, existing methods ignore the structure information and merely utilize protein sequences. Furthermore, designing a mor…

2024

Improving Paratope and Epitope Prediction by Multi-Modal Contrastive Learning and Interaction Informativeness Estimation

IJCAI 2024poster

Accurately predicting antibody-antigen binding residues, i.e., paratopes and epitopes, is crucial in antibody design. However, existing methods solely focus on uni-modal data (either sequence or structure), disregarding the complementary information present in multi-modal data, and most methods pred…

2024

Knowledgeable Preference Alignment for LLMs in Domain-specific Question Answering

ACL 2024findings

Deploying large language models (LLMs) to real scenarios for domain-specific question answering (QA) is a key thrust for LLM applications, which poses numerous challenges, especially in ensuring that responses are both accommodating to user requirements and appropriately leveraging domain-specific k…

2024

Learning to Plan for Retrieval-Augmented Large Language Models from Knowledge Graphs

EMNLP 2024finding

Improving the performance of large language models (LLMs) in complex question-answering (QA) scenarios has always been a research focal point. Recent studies have attempted to enhance LLMs’ performance by combining step-wise planning with external retrieval. While effective for advanced models like…

2024

MKGL: Mastery of a Three-Word Language

NeurIPS 2024spotlight

Large language models (LLMs) have significantly advanced performance across a spectrum of natural language processing (NLP) tasks. Yet, their application to knowledge graphs (KGs), which describe facts in the form of triplets and allow minimal hallucinations, remains an underexplored frontier. In th…

Cited by 1SourcePDFScholar
2024

Prompt-fused Framework for Inductive Logical Query Answering

COLING 2024main

Answering logical queries on knowledge graphs (KG) poses a significant challenge for machine reasoning. The primary obstacle in this task stems from the inherent incompleteness of KGs. Existing research has predominantly focused on addressing the issue of missing edges in KGs, thereby neglecting ano…

2024

Revisit and Outstrip Entity Alignment: A Perspective of Generative Models

ICLR 2024poster

Recent embedding-based methods have achieved great successes in exploiting entity alignment from knowledge graph (KG) embeddings of multiple modalities. In this paper, we study embedding-based entity alignment (EEA) from a perspective of generative models. We show that EEA shares similarities with t…

2024

Stereophonic Music Source Separation with Spatially-Informed Bridging Band-Split Network

ICASSP 2024accepted

Stereophonic music source separation (MSS) is a problem of extracting individual source tracks, e.g. bass, drums, vocals, from a stereo music recording. Deep neural network (DNN) based MSS systems have demonstrated great promise though spatial panning cues and time-frequency spectral structures in s…

Cited by 0SourceScholar
2024

Structure-CLIP: Towards Scene Graph Knowledge to Enhance Multi-Modal Structured Representations

AAAI 2024technical

Large-scale vision-language pre-training has achieved significant performance in multi-modal understanding and generation tasks. However, existing methods often perform poorly on image-text matching tasks that require structured representations, i.e., representations of objects, attributes, and rela…

2024

Unleashing the Power of Imbalanced Modality Information for Multi-modal Knowledge Graph Completion

COLING 2024main

Multi-modal knowledge graph completion (MMKGC) aims to predict the missing triples in the multi-modal knowledge graphs by incorporating structural, visual, and textual information of entities into the discriminant models. The information from different modalities will work together to measure the tr…

2024

ZeroDDI: A Zero-Shot Drug-Drug Interaction Event Prediction Method with Semantic Enhanced Learning and Dual-modal Uniform Alignment

IJCAI 2024poster

Drug-drug interactions (DDIs) can result in various pharmacological changes, which can be categorized into different classes known as DDI events (DDIEs). In recent years, previously unobserved/unseen DDIEs have been emerging, posing a new classification task when unseen classes have no labelled inst…

2023

Analogical Inference Enhanced Knowledge Graph Embedding

AAAI 2023technical

Knowledge graph embedding (KGE), which maps entities and relations in a knowledge graph into continuous vector spaces, has achieved great success in predicting missing links in knowledge graphs. However, knowledge graphs often contain incomplete triples that are difficult to inductively infer by KGE…

2023

DUET: Cross-Modal Semantic Grounding for Contrastive Zero-Shot Learning

AAAI 2023technical

Zero-shot learning (ZSL) aims to predict unseen classes whose samples have never appeared during training. One of the most effective and widely used semantic information for zero-shot image classification are attributes which are annotations for class-level visual characteristics. However, the curre…

2023

Entity-Agnostic Representation Learning for Parameter-Efficient Knowledge Graph Embedding

AAAI 2023technical

We propose an entity-agnostic representation learning method for handling the problem of inefficient parameter storage costs brought by embedding knowledge graphs. Conventional knowledge graph embedding methods map elements in a knowledge graph, including entities and relations, into continuous vect…

2023

Exploring All-In-One Knowledge Distillation Framework for Neural Machine Translation

EMNLP 2023long main

Conventional knowledge distillation(KD) approaches are commonly employed to compress neural machine translation(NMT) models. However, they only obtain one lightweight student each time. Consequently, we have to conduct KD multiple times when different students are required at the same time, which co…

Cited by 0SourcecodeScholar
2023

Exploring Better Text Image Translation with Multimodal Codebook

ACL 2023long

Text image translation (TIT) aims to translate the source texts embedded in the image to target translations, which has a wide range of applications and thus has important research value. However, current studies on TIT are confronted with two main bottlenecks: 1) this task lacks a publicly availabl…

2023

Generalizing to Unseen Elements: A Survey on Knowledge Extrapolation for Knowledge Graphs

IJCAI 2023poster

Knowledge graphs (KGs) have become valuable knowledge resources in various applications, and knowledge graph embedding (KGE) methods have garnered increasing attention in recent years. However, conventional KGE methods still face challenges when it comes to handling unseen entities or relations duri…

Cited by 26SourcePDFScholar
2023

Joint Training and Decoding for Multilingual End-to-End Simultaneous Speech Translation

ICASSP 2023accepted

Recent studies on end-to-end speech translation(ST) have facilitated the exploration of multilingual end-to-end ST and end-to-end simultaneous ST. In this paper, we investigate end-to-end simultaneous speech translation in a one-to-many multilingual setting which is closer to applications in real sc…

Cited by 0SourceScholar
2023

Multi-Relational Contrastive Learning Graph Neural Network for Drug-Drug Interaction Event Prediction

AAAI 2023technical

Drug-drug interactions (DDIs) could lead to various unexpected adverse consequences, so-called DDI events. Predicting DDI events can reduce the potential risk of combinatorial therapy and improve the safety of medication use, and has attracted much attention in the deep learning community. Recently,…

2023

Multi-view Contrastive Learning Hypergraph Neural Network for Drug-Microbe-Disease Association Prediction

IJCAI 2023poster

Identifying the potential associations among drugs, microbes and diseases is of great significance in exploring the pathogenesis and improving precision medicine. There are plenty of computational methods for pair-wise association prediction, such as drug-microbe and microbe-disease associations, bu…

2023

Rethinking the Reasonability of the Test Set for Simultaneous Machine Translation

ICASSP 2023accepted

Simultaneous machine translation (SimulMT) models start translation before the end of the source sentence, making the translation monotonically aligned with the source sentence. However, the general full-sentence translation test set is acquired by offline translation of the entire source sentence,…

Cited by 0SourceScholar
2023

Towards Better Entity Linking with Multi-View Enhanced Distillation

ACL 2023long

Dense retrieval is widely used for entity linking to retrieve entities from large-scale knowledge bases. Mainstream techniques are based on a dual-encoder framework, which encodes mentions and entities independently and calculates their relevances via rough interaction metrics, resulting in difficul…

2022

Meta-Learning Based Knowledge Extrapolation for Knowledge Graphs in the Federated Setting

IJCAI 2022poster

We study the knowledge extrapolation problem to embed new components (i.e., entities and relations) that come with emerging knowledge graphs (KGs) in the federated setting. In this problem, a model trained on an existing KG needs to embed an emerging KG with unseen entities and relations. To solve t…

2022

Molecular Contrastive Learning with Chemical Element Knowledge Graph

AAAI 2022technical

Molecular representation learning contributes to multiple downstream tasks such as molecular property prediction and drug design. To properly represent molecules, graph contrastive learning is a promising paradigm as it utilizes self-supervision signals and has no requirements for human annotations.…

2022

Neural-Symbolic Entangled Framework for Complex Query Answering

NeurIPS 2022accept

Answering complex queries over knowledge graphs (KG) is an important yet challenging task because of the KG incompleteness issue and cascading errors during reasoning. Recent query embedding (QE) approaches embed the entities and relations in a KG and the first-order logic (FOL) queries into a low d…

Cited by 23SourcePDFScholar
2022

Robust Pressure Matching with ATF Perturbation Constraints for Sound Field Control

ICASSP 2022accepted

Sound field control systems deployed in room acoustic environments require knowing the acoustic channel impulse responses between the loudspeakers and matching microphones, which are challenging to estimate accurately due to perturbations caused by such factors as temperature changes and sensors’ po…

Cited by 0SourceScholar
2022

Ruleformer: Context-aware Rule Mining over Knowledge Graph

COLING 2022main

Rule mining is an effective approach for reasoning over knowledge graph (KG). Existing works mainly concentrate on mining rules. However, there might be several rules that could be applied for reasoning for one relation, and how to select appropriate rules for completion of different triples has not…

2022

Towards Robust Neural Machine Translation with Iterative Scheduled Data-Switch Training

COLING 2022main

Most existing methods on robust neural machine translation (NMT) construct adversarial examples by injecting noise into authentic examples and indiscriminately exploit two types of examples. They require the model to translate both the authentic source sentence and its adversarial counterpart into t…

2021

CSGNN: Contrastive Self-Supervised Graph Neural Network for Molecular Interaction Prediction

IJCAI 2021poster

Molecular interactions are significant resources for analyzing sophisticated biological systems. Identification of multifarious molecular interactions attracts increasing attention in biomedicine, bioinformatics, and human healthcare communities. Recently, a plethora of methods have been proposed to…

2020

Bridging the Gap between Training and Inference for Neural Machine Translation (Extended Abstract)

IJCAI 2020poster

Neural Machine Translation (NMT) generates target words sequentially in the way of predicting the next word conditioned on the context words. At training time, it predicts with the ground truth words as context while at inference it has to generate the entire sequence from scratch. This discrepancy…

2020

Logic-guided Semantic Representation Learning for Zero-Shot Relation Classification

COLING 2020main

Relation classification aims to extract semantic relations between entity pairs from the sentences. However, most existing methods can only identify seen relation classes that occurred during training. To recognize unseen relations at test time, we explore the problem of zero-shot relation classific…

Cited by 44SourcePDFScholar
2020

Neural Entity Summarization with Joint Encoding and Weak Supervision

IJCAI 2020poster

In a large-scale knowledge graph (KG), an entity is often described by a large number of triple-structured facts. Many applications require abridged versions of entity descriptions, called entity summaries. Existing solutions to entity summarization are mainly unsupervised. In this paper, we present…

2020

Towards Playing Full MOBA Games with Deep Reinforcement Learning

NeurIPS 2020poster

MOBA games, e.g., Honor of Kings, League of Legends, and Dota 2, pose grand challenges to AI systems such as multi-agent, enormous state-action space, complex action control, etc. Developing AI for playing MOBA games has raised much attention accordingly. However, existing work falls short in handli…

2019

2.5D Multizone Reproduction with Active Control of Scattered Sound Fields

ICASSP 2019accepted

Multizone reproduction has been focused on reproducing sounds in an empty listening space. However, there are always scatterers such as human heads in sound zones, generating scattered sound fields and causing degraded system performance. In this work, we develop a modal-domain method for 2.5D multi…

Cited by 0SourceScholar
2019

Robust Sparse Multichannel Active Noise Control

ICASSP 2019accepted

Multichannel active noise control (MC-ANC) aims to cancel low-frequency noise in an enclosure. If noise sources are distributed sparsely in space, adding an ℓ <inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</inf> -norm constraint to the standard MC-AN…

Cited by 0SourceScholar
2018

2.5D Multizone Reproduction Using Weighted Mode Matching

ICASSP 2018accepted

The mode matching based multizone reproduction has mainly been focused on a purely 2D theory which is inadequate to fit the 3D reality. Its extension to the 3D theory however requires many secondary sources and a high computational complexity. In this paper, a weighted mode matching approach is deve…

Cited by 0SourceScholar
2018

Reference Signal Generation for Broadband ANC Systems in Reverberant Rooms

ICASSP 2018accepted

One major issue of implementing broadband active noise control systems in reverberant rooms is the lack of reference signals. In this work, by exploiting the spatial sound field characteristics, a time-domain sound field separation method is developed to generate the reference signal for broadband a…

Cited by 0SourceScholar
2017

An Optimal Transportation Based Univariate Neuroimaging Index

ICCV 2017poster

The alterations of brain structures and functions have been considered closely correlated to the change of cognitive performance due to neurodegenerative diseases such as Alzheimer's disease. In this paper, we introduce a variational framework to compute the optimal transformation (OT) in 3D space a…

Cited by 7PDFScholar
2017

Online secondary path modelling in wave-domain active noise control

ICASSP 2017accepted

The performance of an ANC system largely depends on the availability of an accurate secondary path model. This is however a major challenge in multichannel ANC where the computational complexity increases significantly with the number of secondary sources and error sensors. This paper proposes wave-…

Cited by 0SourceScholar
2016

Sparse complex FxLMS for active noise cancellation over spatial regions

ICASSP 2016accepted

In this paper, we investigate active noise control over large 2D spatial regions when the noise source is sparsely distributed. The l <inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</inf> relaxation technique originated from compressive sensing is ado…

Cited by 0SourceScholar
2016

Spatial feature learning for robust binaural sound source localization using a composite feature vector

ICASSP 2016accepted

The performance of binaural speech source localization systems can be significantly impacted by an imperfect selection of spatial localization cues, due to the limited bandwidth of speech, and the effects of noise. In order to mitigate these impacts, this paper presents a novel method that combines…

Cited by 0SourceScholar
2015

Binaural localization of speech sources in 3-D using a composite feature vector of the HRTF

ICASSP 2015accepted

Binaural localization of speech sources in 3-D, using head-related transfer functions (HRTFs), always suffers elevation ambiguity due to the limited high frequency spectral information available at the receivers. This paper presents a method that overcomes this limitation by exploiting the interaura…

Cited by 0SourceScholar