← Search

Fei Li

56 accepted papers

2026

Beyond Implicit Constraint: Explicit Low-Rank Structured Subspace Learning for Fast Attributed Graph Clustering

IJCAI 2026

Attributed graph clustering has achieved remarkable success by synergistically integrating topological structures and node attributes. While subspace learning has emerged as a dominant paradigm for node partitioning, most existing methods rely on implicit low-rank constraints, which often fail to ca

Cited by 0Scholar
2026

Enhance-then-Balance Modality Collaboration for Robust Multimodal Sentiment Analysis

CVPR 2026

Multimodal sentiment analysis (MSA) seeks to infer human emotions by integrating heterogeneous signals from text, audio, and visual modalities.Although recent approaches attempt to leverage cross-modal complementarity, they often struggle to fully utilize weaker modalities.In practice, the expressiv

Cited by 0SourcecodeScholar
2026

Generating Attribute-Aware Human Motions from Textual Prompt

AAAI 2026technical

Text-driven human motion generation has recently attracted considerable attention, allowing models to generate human motions based on textual descriptions. However, current methods neglect the influence of human attributes—such as age, gender, weight, and height—which are key factors shaping human m

Cited by 0SourcePDFScholar
2026

Interest-Shift-Aware Logical Reasoning for Efficient Long-Sequence Recommendation

AAAI 2026technical

Logical reasoning-based recommendation methods formulate logical expressions to characterize user-item interaction patterns, incorporating regularization constraints to ensure consistency with logical rules. However, these methods face two critical challenges: (1) As sequence length increases, they

Cited by 0SourcePDFScholar
2026

KVmix: Gradient-Based Layer Importance-Aware Mixed-Precision Quantization for KV Cache

AAAI 2026technical

The high memory demands of the Key-Value (KV) Cache during the inference of Large Language Models (LLMs) severely restrict their deployment in resource-constrained platforms. Quantization can effectively alleviate the memory pressure caused by KV Cache. However, existing methods either rely on stati

Cited by 0SourcePDFScholar
2026

PaSE: Prototype-aligned Calibration and Shapley-based Equilibrium for Multimodal Sentiment Analysis

AAAI 2026technical

Multimodal Sentiment Analysis (MSA) seeks to understand human emotions by integrating textual, acoustic, and visual signals. Although multimodal fusion is designed to leverage cross-modal complementarity, real-world scenarios often exhibit modality competition: dominant modalities tend to overshadow

Cited by 0SourcePDFScholar
2025

Cross-Document Cross-Lingual NLI via RST-Enhanced Graph Fusion and Interpretability Prediction

EMNLP 2025

Natural Language Inference (NLI) is a fundamental task in natural language processing. While NLI has developed many subdirections such as sentence-level NLI, document-level NLI and cross-lingual NLI, Cross-Document Cross-Lingual NLI (CDCL-NLI) remains largely unexplored. In this paper, we propose a

Cited by 0SourcePDFScholar
2025

DALR: Dual-level Alignment Learning for Multimodal Sentence Representation Learning

ACL 2025finding

Previous multimodal sentence representation learning methods have achieved impressive performance. However, most approaches focus on aligning images and text at a coarse level, facing two critical challenges: cross-modal misalignment bias and intra-modal semantic divergence, which significantly degr…

2025

DiscoSG: Towards Discourse-Level Text Scene Graph Parsing through Iterative Graph Refinement

EMNLP 2025

Vision-Language Models (VLMs) generate discourse-level, multi-sentence visual descriptions, challenging text scene graph parsers built for single-sentence caption-to-graph mapping. Current approaches typically merge sentence-level parsing outputs for discourse input, often missing phenomena like cro

2025

Dual Multi-Scale GCN with Deformable Temporal Kernel for Skeleton-based Action Recognition

ICASSP 2025accepted

Skeleton sequences for action recognition are with complex temporal dynamics due to various factors such as speed variation and different activities. It is crucial and essential to model variation changes in the temporal dimension. In recent years, skeleton sequence is always modeled as a graph stru…

Cited by 0SourceScholar
2025

EVOREFUSE: Evolutionary Prompt Optimization for Evaluation and Mitigation of LLM Over-Refusal to Pseudo-Malicious Instructions

NeurIPS 2025poster

Large language models (LLMs) frequently refuse to respond to pseudo-malicious instructions: semantically harmless input queries triggering unnecessary LLM refusals due to conservative safety alignment, significantly impairing user experience. Collecting such instructions is crucial for evaluating an…

Cited by 0SourcecodeScholar
2025

Enhancing Hyperbole and Metaphor Detection with Their Bidirectional Dynamic Interaction and Emotion Knowledge

ACL 2025long

Text-based hyperbole and metaphor detection are of great significance for natural language processing (NLP) tasks. However, due to their semantic obscurity and expressive diversity, it is rather challenging to identify them. Existing methods mostly focus on superficial text features, ignoring the as…

2025

Harnessing Dimensional Contrast and Information Compensation for Sentence Embedding Enhancement

ICASSP 2025accepted

Unsupervised sentence embedding learning excels through positive sample construction and instance-level contrastive learning (ICL). However, this approach can lead to over-compression and dimensional contamination from noisy data augmentation and unconstrained ICL processes. To mitigate these issues…

Cited by 0SourceScholar
2025

Multi-Granular Multimodal Clue Fusion for Meme Understanding

AAAI 2025technical

With the continuous emergence of various social media platforms frequently used in daily life, the multimodal meme understanding (MMU) task has been garnering increasing attention. MMU aims to explore and comprehend the meanings of memes from various perspectives by performing tasks such as metaphor…

Cited by 0SourcePDFScholar
2025

PEMV: Improving Spatial Distribution for Emotion Recognition in Conversations Using Proximal Emotion Mean Vectors

NAACL 2025findings

Emotion Recognition in Conversation (ERC) aims to identify the emotions expressed in each utterance within a dialogue. Existing research primarily focuses on the analysis of contextual structure in dialogue and the interactions between different emotions. Nonetheless, ERC datasets often contain diff…

Cited by 0SourcePDFScholar
2025

PointSR: Self-Regularized Point Supervision for Drone-View Object Detection

CVPR 2025poster

Point-Supervised Object Detection (PSOD) in a discriminative style has recently gained significant attention for its impressive detection performance and cost-effectiveness. However, accurately predicting high-quality pseudo-box labels for drone-view images, which often feature densely packed small…

Cited by 0SourcePDFScholar
2025

TRIDENT: Enhancing Large Language Model Safety with Tri-Dimensional Diversified Red-Teaming Data Synthesis

ACL 2025long

Large Language Models (LLMs) excel in various natural language processing tasks but remain vulnerable to generating harmful content or being exploited for malicious purposes. Although safety alignment datasets have been introduced to mitigate such risks through supervised fine-tuning (SFT), these da…

2025

Zero-Shot Conversational Stance Detection: Dataset and Approaches

ACL 2025finding

Stance detection, which aims to identify public opinion towards specific targets using social media data, is an important yet challenging task. With the increasing number of online debates among social media users, conversational stance detection has become a crucial research area. However, existing…

2024

Benchmarks and Challenges in Pose Estimation for Egocentric Hand Interactions with Objects

ECCV 2024poster

"We interact with the world with our hands and see it through our own (egocentric) perspective. A holistic understanding of such interactions from egocentric views is important for tasks in robotics, AR/VR, action recognition and motion generation. Accurately reconstructing such interactions in is c…

2024

Compositional Generalization for Multi-Label Text Classification: A Data-Augmentation Approach

AAAI 2024technical

Despite significant advancements in multi-label text classification, the ability of existing models to generalize to novel and seldom-encountered complex concepts, which are compositions of elementary ones, remains underexplored. This research addresses this gap. By creating unique data splits acros…

2024

Development of Negative-Pressure Artificial Muscles With Fiber Constraints and Pre-Stretched Soft Skin

RA-L 2024

Negative-pressure artificial muscles based on internal support and flexible skin provide a way to develop high-performance artificial muscles. However, the random and disordered skin wrinkles formed during the contraction may lead to uncertainty in the actuation behavior, and the flexible but inexte

Cited by 3SourceScholar
2024

Enhancing Cross-Document Event Coreference Resolution by Discourse Structure and Semantic Information

COLING 2024main

Existing cross-document event coreference resolution models, which either compute mention similarity directly or enhance mention representation by extracting event arguments (such as location, time, agent, and patient), lackingmthe ability to utilize document-level information. As a result, they str…

2024

Harnessing Holistic Discourse Features and Triadic Interaction for Sentiment Quadruple Extraction in Dialogues

AAAI 2024technical

Dialogue Aspect-based Sentiment Quadruple (DiaASQ) is a newly-emergent task aiming to extract the sentiment quadruple (i.e., targets, aspects, opinions, and sentiments) from conversations. While showing promising performance, the prior DiaASQ approach unfortunately falls prey to the key crux of DiaA…

Cited by 7SourcePDFScholar
2024

Harvesting Events from Multiple Sources: Towards a Cross-Document Event Extraction Paradigm

ACL 2024findings

Document-level event extraction aims to extract structured event information from unstructured text. However, a single document often contains limited event information and the roles of different event arguments may be biased due to the influence of the information source.This paper addresses the li…

2024

MindMap: Constructing Evidence Chains for Multi-Step Reasoning in Large Language Models

AAAI 2024technical

Large language models (LLMs) have demonstrated remarkable performance in various natural language processing tasks. However, they still face significant challenges in automated reasoning, particularly in scenarios involving multi-step reasoning. In this paper, we focus on the logical reasoning probl…

Cited by 1SourcePDFScholar
2024

Recognizing Everything from All Modalities at Once: Grounded Multimodal Universal Information Extraction

ACL 2024findings

In the field of information extraction (IE), tasks across a wide range of modalities and their combinations have been traditionally studied in isolation, leaving a gap in deeply recognizing and analyzing cross-modal information. To address this, this work for the first time introduces the concept of…

2024

Refining and Synthesis: A Simple yet Effective Data Augmentation Framework for Cross-Domain Aspect-based Sentiment Analysis

ACL 2024findings

Aspect-based Sentiment Analysis (ABSA) is extensively researched in the NLP community, yet related models face challenges due to data sparsity when shifting to a new domain. Hence, data augmentation for cross-domain ABSA has attracted increasing attention in recent years. However, two key points hav…

Cited by 2SourcePDFScholar
2024

Reverse Multi-Choice Dialogue Commonsense Inference with Graph-of-Thought

AAAI 2024technical

With the proliferation of dialogic data across the Internet, the Dialogue Commonsense Multi-choice Question Answering (DC-MCQ) task has emerged as a response to the challenge of comprehending user queries and intentions. Although prevailing methodologies exhibit effectiveness in addressing single-ch…

2024

Revisiting Structured Sentiment Analysis as Latent Dependency Graph Parsing

ACL 2024long

Structured Sentiment Analysis (SSA) was cast as a problem of bi-lexical dependency graph parsing by prior studies.Multiple formulations have been proposed to construct the graph, which share several intrinsic drawbacks:(1) The internal structures of spans are neglected, thus only the boundary tokens…

Cited by 0SourcePDFScholar
2024

What Factors Influence LLMs’ Judgments? A Case Study on Question Answering

COLING 2024main

Large Language Models (LLMs) are now being considered as judges of high efficiency to evaluate the quality of answers generated by candidate models. However, their judgments may be influenced by complex scenarios and inherent biases, raising concerns about their reliability. This study aims to bridg…

Cited by 3SourcePDFScholar
2023

DiaASQ: A Benchmark of Conversational Aspect-based Sentiment Quadruple Analysis

ACL 2023findings

The rapid development of aspect-based sentiment analysis (ABSA) within recent decades shows great potential for real-world society. The current ABSA works, however, are mostly limited to the scenario of a single text piece, leaving the study in dialogue contexts unexplored. To bridge the gap between…

2023

Dialogue State Distillation Network with Inter-slot Contrastive Learning for Dialogue State Tracking

AAAI 2023technical

In task-oriented dialogue systems, Dialogue State Tracking (DST) aims to extract users' intentions from the dialogue history. Currently, most existing approaches suffer from error propagation and are unable to dynamically select relevant information when utilizing previous dialogue states. Moreover,…

Cited by 7SourcePDFScholar
2023

FACTUAL: A Benchmark for Faithful and Consistent Textual Scene Graph Parsing

ACL 2023findings

Textual scene graph parsing has become increasingly important in various vision-language applications, including image caption evaluation and image retrieval. However, existing scene graph parsers that convert image captions into scene graphs often suffer from two types of errors. First, the generat…

2023

Multi-Frequency Representation Enhancement with Privilege Information for Video Super-Resolution

ICCV 2023poster

CNN's limited receptive field restricts its ability to capture long-range spatial-temporal dependencies, leading to unsatisfactory performance in video super-resolution. To tackle this challenge, this paper presents a novel multi-frequency representation enhancement module (MFE) that performs spatia…

Cited by 20PDFScholar
2023

Reasoning Implicit Sentiment with Chain-of-Thought Prompting

ACL 2023short

While sentiment analysis systems try to determine the sentiment polarities of given targets based on the key opinion expressions in input texts, in implicit sentiment analysis (ISA) the opinion cues come in an implicit and obscure manner. Thus detecting implicit sentiment requires the common-sense a…

2022

Effective Token Graph Modeling using a Novel Labeling Strategy for Structured Sentiment Analysis

ACL 2022long

The state-of-the-art model for structured sentiment analysis casts the task as a dependency parsing problem, which has some limitations: (1) The label proportions for span prediction and span relation prediction are imbalanced. (2) The span lengths of sentiment tuple components may be very large in…

2022

Entity-centered Cross-document Relation Extraction

EMNLP 2022main

Relation Extraction (RE) is a fundamental task of information extraction, which has attracted a large amount of research attention. Previous studies focus on extracting the relations within a sentence or document, while currently researchers begin to explore cross-document RE. However, current cross…

2022

Global Inference with Explicit Syntactic and Discourse Structures for Dialogue-Level Relation Extraction

IJCAI 2022poster

Recent research attention for relation extraction has been paid to the dialogue scenario, i.e., dialogue-level relation extraction (DiaRE). Existing DiaRE methods either simply concatenate the utterances in a dialogue into a long piece of text, or employ naive words, sentences or entities to build d…

2022

Inheriting the Wisdom of Predecessors: A Multiplex Cascade Framework for Unified Aspect-based Sentiment Analysis

IJCAI 2022poster

So far, aspect-based sentiment analysis (ABSA) has involved with total seven subtasks, in which, however the interactions among them have been left unexplored sufficiently. This work presents a novel multiplex cascade framework for unified ABSA and maintaining such interactions. First, we model tota…

2022

Joint Alignment of Multi-Task Feature and Label Spaces for Emotion Cause Pair Extraction

COLING 2022main

Emotion cause pair extraction (ECPE), as one of the derived subtasks of emotion cause analysis (ECA), shares rich inter-related features with emotion extraction (EE) and cause extraction (CE). Therefore EE and CE are frequently utilized as auxiliary tasks for better feature learning, modeled via mul…

2022

LasUIE: Unifying Information Extraction with Latent Adaptive Structure-aware Generative Language Model

NeurIPS 2022accept

Universally modeling all typical information extraction tasks (UIE) with one generative language model (GLM) has revealed great potential by the latest study, where various IE predictions are unified into a linearized hierarchical expression under a GLM. Syntactic structure information, a type of ef…

2022

Mastering the Explicit Opinion-Role Interaction: Syntax-Aided Neural Transition System for Unified Opinion Role Labeling

AAAI 2022technical

Unified opinion role labeling (ORL) aims to detect all possible opinion structures of 'opinion-holder-target' in one shot, given a text. The existing transition-based unified method, unfortunately, is subject to longer opinion terms and fails to solve the term overlap issue. Current top performance…

2022

OneEE: A One-Stage Framework for Fast Overlapping and Nested Event Extraction

COLING 2022main

Event extraction (EE) is an essential task of information extraction, which aims to extract structured event information from unstructured text. Most prior work focuses on extracting flat events while neglecting overlapped or nested ones. A few models for overlapped and nested EE includes several su…

2022

Unified Named Entity Recognition as Word-Word Relation Classification

AAAI 2022technical

So far, named entity recognition (NER) has been involved with three major types, including flat, overlapped (aka. nested), and discontinuous NER, which have mostly been studied individually. Recently, a growing interest has been built for unified NER, tackling the above three jobs concurrently with…

2021

A Span-Based Model for Joint Overlapped and Discontinuous Named Entity Recognition

ACL 2021long

Research on overlapped and discontinuous named entity recognition (NER) has received increasing attention. The majority of previous work focuses on either overlapped or discontinuous entities. In this paper, we propose a novel span-based model that can recognize both overlapped and discontinuous ent…

2021

Encoder-Decoder Based Unified Semantic Role Labeling with Label-Aware Syntax

AAAI 2021technical

Currently the unified semantic role labeling (SRL) that achieves predicate identification and argument role labeling in an end-to-end manner has received growing interests. Recent works show that leveraging the syntax knowledge significantly enhances the SRL performances. In this paper, we investiga…

2021

Modularized Interaction Network for Named Entity Recognition

ACL 2021long

Although the existing Named Entity Recognition (NER) models have achieved promising performance, they suffer from certain drawbacks. The sequence labeling-based NER models do not perform well in recognizing long entities as they focus only on word-level information, while the segment-based NER model…

Cited by 40SourcePDFScholar
2021

Rethinking Boundaries: End-To-End Recognition of Discontinuous Mentions with Pointer Networks

AAAI 2021technical

A majority of research interests in irregular (e.g., nested or discontinuous) named entity recognition (NER) have been paid on nested entities, while discontinuous entities received limited attention. Existing work for discontinuous NER, however, either suffers from decoding ambiguity or predicting…

2020

HiTrans: A Transformer-Based Context- and Speaker-Sensitive Model for Emotion Detection in Conversations

COLING 2020main

Emotion detection in conversations (EDC) is to detect the emotion for each utterance in conversations that have multiple speakers. Different from the traditional non-conversational emotion detection, the model for EDC should be context-sensitive (e.g., understanding the whole conversation rather tha…

2019

Lung Nodule Detection with a 3D ConvNet via IoU Self-normalization and Maxout Unit

ICASSP 2019accepted

The automatic pulmonary nodule detection in thoracic computed tomography (CT) scans plays a crucial role in the early diagnosis of lung cancer. In this paper, we propose a novel framework with a 3D convolutional network (ConvNet) for pulmonary nodule detection. To improve the efficiency and flexibil…

Cited by 0SourceScholar
2019

Two-stream Multi-focus Image Fusion Based on the Latent Decision Map

ICASSP 2019accepted

The multi-focus image fusion with deep learning methods is mostly regarded as a two or three-category problem. Current systems utilize sliding windows to classify each pixel into focused or defocused, which is time consuming and requires post-processing such as denoising. In this paper, we propose a…

Cited by 0SourceScholar
2015

Hyperspectral Compressive Sensing Using Manifold-Structured Sparsity Prior

ICCV 2015poster

To reconstruct hyperspectral image (HSI) accurately from a few noisy compressive measurements, we present a novel manifold-structured sparsity prior based hyperspectral compressive sensing (HCS) method in this study. A matrix based hierarchical prior is first proposed to represent the spectral struc…

Cited by 18PDFScholar
2015

Reweighted Laplace Prior Based Hyperspectral Compressive Sensing for Unknown Sparsity

CVPR 2015poster

Compressive sensing(CS) has been exploited for hypespectral image(HSI) compression in recent years. Though it can greatly reduce the costs of computation and storage, the reconstruction of HSI from a few linear measurements is challenging. The underlying sparsity of HSI is crucial to improve the rec…

Cited by 39SourcePDFScholar