← Search

Zhao Lv

26 accepted papers

2026

BrainHGT: A Hierarchical Graph Transformer for Interpretable Brain Network Analysis

AAAI 2026technical

Graph Transformer shows remarkable potential in brain network analysis due to its ability to model graph structures and complex node relationships. Most existing methods typically model the brain as a flat network, ignoring its modular structure, and their attention mechanisms treat all brain region

Cited by 1SourcePDFScholar
2026

Dual-stream Relation-modeling Disentanglement for Cloth-Changing Person Re-Identification

AAAI 2026technical

Cloth-changing person re-identification (CC-ReID) aims to identify individuals across non-overlapping cameras despite clothing variations. Existing methods are often constrained by two primary limitations: approaches using auxiliary modalities typically rely on a single specific cue, limiting their

Cited by 0SourcePDFScholar
2026

PCRNet: Phase-aware Complex Refinement Network for EEG-based Auditory Attention Decoding

ICML 2026poster

Auditory attention decoding (AAD) based on Electroencephalography (EEG) aims to identify the attended speaker in multi-speaker environments. However, existing methods typically overlook the crucial phase information of EEG signals, which limits their ability to distinguish structured neural patterns…

Cited by 0SourceScholar
2026

Trifuse: Enhancing Attention-Based GUI Grounding via Multimodal Fusion

ICML 2026poster

GUI grounding maps natural language instructions to the correct interface elements, serving as the perception foundation for GUI agents. Existing approaches predominantly rely on fine-tuning multimodal large language models (MLLMs) using large-scale GUI datasets to predict target element coordinates…

Cited by 0SourceScholar
2025

BSDB-Net: Band-Split Dual-Branch Network with Selective State Spaces Mechanism for Monaural Speech Enhancement

AAAI 2025technical

Although the complex spectrum-based speech enhancement (SE) methods have achieved significant performance, coupling amplitude and phase can lead to a compensation effect, where amplitude information is sacrificed to compensate for the phase that is harmful to SE. In addition, to further improve the…

Cited by 0SourcePDFScholar
2025

COLA: Collaborative Multi-Agent Framework with Dynamic Task Scheduling for GUI Automation

EMNLP 2025

With the rapid advancements in Large Language Models (LLMs), an increasing number of studies have leveraged LLMs as the cognitive core of agents to address complex task decision-making challenges. Specially, recent research has demonstrated the potential of LLM-based agents on automating GUI operati

2025

Community-Aware Graph Transformer for Brain Disorder Identification

IJCAI 2025

Abnormal brain functional network is an effective biomarker for brain disease diagnosis. Most existing methods focus on mining discriminative information from whole-brain connectivity patterns. However, multi-level collaboration is the foundation of efficient brain function, in addition to the whole

2025

EEG Correlation Analysis-guided Graph Local Enhanced Feature Learning For Emotion Recognition

ICASSP 2025accepted

EEG-based emotion recognition is a key technology in brain-computer interfaces. Many previous studies have applied deep learning methods to mine emotion-related features in EEG to decode emotions. However, they overlooked the importance of electrode correlations and varying brain region activation d…

Cited by 0SourceScholar
2025

ID-RemovalNet: Identity Removal Network for EEG Privacy Protection with Enhancing Decoding Tasks

IJCAI 2025

Electroencephalogram (EEG) contains not only decoding task information but also personal identity privacy information. If it is stolen or attacked, the user's brain-computer interaction behavior may be maliciously manipulated. Existing EEG identity privacy protection generally adopts generative or a

Cited by 0SourcePDFScholar
2025

Improved Feature Extraction Network for Neuro-Oriented Target Speaker Extraction

ICASSP 2025accepted

The recent rapid development of auditory attention decoding (AAD) offers the possibility of using electroencephalography (EEG) as auxiliary information for target speaker extraction. However, effectively modeling long sequences of speech and resolving the identity of the target speaker from EEG sign…

Cited by 0SourceScholar
2025

ListenNet: A Lightweight Spatio-Temporal Enhancement Nested Network for Auditory Attention Detection

IJCAI 2025

Auditory attention detection (AAD) aims to identify the direction of the attended speaker in multi-speaker environments from brain signals, such as Electroencephalography (EEG) signals. However, existing EEG-based AAD methods overlook the spatio-temporal dependencies of EEG signals, limiting their d

2025

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction

IJCAI 2025

The brain-assisted target speaker extraction (TSE) aims to extract the attended speech from mixed speech by utilizing the brain neural activities, for example Electroencephalography (EEG). However, existing models overlook the issue of temporal misalignment between speech and EEG modalities, which h

2025

MHANet: Multi-scale Hybrid Attention Network for Auditory Attention Detection

IJCAI 2025

Auditory attention detection (AAD) aims to detect the target speaker in a multi-talker environment from brain signals, such as electroencephalography (EEG), which has made great progress. However, most AAD methods solely utilize attention mechanisms sequentially and overlook valuable multi-scale con

2025

Region-Based Optimization in Continual Learning for Audio Deepfake Detection

AAAI 2025technical

Rapid advancements in speech synthesis and voice conversion bring convenience but also new security risks, creating an urgent need for effective audio deepfake detection. Although current models perform well, their effectiveness diminishes when confronted with the diverse and evolving nature of real…

2025

SSM2Mel: State Space Model to Reconstruct Mel Spectrogram from the EEG

ICASSP 2025accepted

Decoding speech from brain signals is a challenging research problem that holds significant importance for studying speech processing in the brain. Although breakthroughs have been made in reconstructing the mel spectrograms of audio stimuli perceived by subjects at the word or letter level using no…

Cited by 0SourceScholar
2025

Transformer Based Multi-view Learning for Integrating Static and Dynamic Complementarity of Brain Function

ICASSP 2025accepted

Dynamic temporal information and static connectivity information derived from functional magnetic resonance imaging (fMRI) can assist in the diagnosis of neurological disorders. However, existing disease diagnosis methods primarily rely on information from a single view, neglecting the advantages of…

Cited by 0SourceScholar
2024

A Non-parametric Graph Clustering Framework for Multi-View Data

AAAI 2024technical

Multi-view graph clustering (MVGC) derives encouraging grouping results by seamlessly integrating abundant information inside heterogeneous data, and has captured surging focus recently. Nevertheless, the majority of current MVGC works involve at least one hyper-parameter, which not only requires…

Cited by 19SourcePDFScholar
2024

Bilateral Masking with prompt for Knowledge Graph Completion

NAACL 2024findings

The pre-trained language model (PLM) has achieved significant success in the field of knowledge graph completion (KGC) by effectively modeling entity and relation descriptions. In recent studies, the research in this field has been categorized into methods based on word matching and sentence matchin…

Cited by 1SourcePDFScholar
2024

DARNet: Dual Attention Refinement Network with Spatiotemporal Construction for Auditory Attention Detection

NeurIPS 2024poster

At a cocktail party, humans exhibit an impressive ability to direct their attention. The auditory attention detection (AAD) approach seeks to identify the attended speaker by analyzing brain signals, such as EEG signals. However, current AAD algorithms overlook the spatial distribution information…

2024

DBPNet: Dual-Branch Parallel Network with Temporal-Frequency Fusion for Auditory Attention Detection

IJCAI 2024poster

Auditory attention decoding (AAD) aims to recognize the attended speaker based on electroencephalography (EEG) signals in multi-talker environments. Most AAD methods only focus on the temporal or frequency domain, but neglect the relationships between these two domains, which results in the inabilit…

Cited by 15SourcePDFScholar
2024

Progressive Distillation Based on Masked Generation Feature Method for Knowledge Graph Completion

AAAI 2024technical

In recent years, knowledge graph completion (KGC) models based on pre-trained language model (PLM) have shown promising results. However, the large number of parameters and high computational cost of PLM models pose challenges for their application in downstream tasks. This paper proposes a progress…

2024

Pruning via Merging: Compressing LLMs via Manifold Alignment Based Layer Merging

EMNLP 2024main

While large language models (LLMs) excel in many domains, their complexity and scale challenge deployment in resource-limited environments. Current compression techniques, such as parameter pruning, often fail to effectively utilize the knowledge from pruned parameters. To address these challenges,…

2024

UNO Arena for Evaluating Sequential Decision-Making Capability of Large Language Models

EMNLP 2024main

Sequential decision-making refers to algorithms that take into account the dynamics of the environment, where early decisions affect subsequent decisions. With large language models (LLMs) demonstrating powerful capabilities between tasks, we can’t help but ask: Can Current LLMs Effectively Make Seq…

Cited by 2SourcePDFScholar
2023

Learning From Yourself: A Self-Distillation Method For Fake Speech Detection

ICASSP 2023accepted

In this paper, we propose a novel self-distillation method for fake speech detection (FSD), which can significantly improve the performance of FSD without increasing the model complexity. For FSD, some fine-grained information is very important, such as spectrogram defects, mute segments, and so on,…

Cited by 0SourceScholar
2022

Csenet: Complex Squeeze-and-Excitation Network for Speech Depression Level Prediction

ICASSP 2022accepted

Automatic speech depression level prediction (SDLP) is a very challenging problem in affective computing. There are many studies that have acquired quite good performances for SDLP. However, most of the input speech features of these studies are based on the amplitude spectrogram, which loses the ph…

Cited by 0SourceScholar