← Search

Qian Li

61 accepted papers

2026

Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models

AAAI 2026technical

Large Multimodal Models (LMMs) face notable challenges when encountering multimodal knowledge conflicts, particularly under retrieval-augmented generation (RAG) frameworks, where the contextual information from external sources may contradict the model’s internal parametric knowledge, leading to unr

Cited by 0SourcePDFScholar
2026

Cross-Scale Collaboration between LLMs and Lightweight Sequential Recommenders with Domain-Specific Latent Reasoning

AAAI 2026technical

Sequential recommendation aims to predict the next item based on historical interactions. To further enhance the reasoning capability in sequential recommendation, LLMs are employed to predict the next item or generate semantic IDs for item representation, given LLMs

Cited by 0SourcePDFScholar
2026

Directing Uncertainty-Aware Information Flow for Robust Diffusion Prediction

AAAI 2026technical

Information diffusion prediction is crucial for understanding social network dynamics, yet existing methods often neglect user participation uncertainty. This oversight typically stems from an implicit participation homogeneity assumption, which treats all observed interactions as equally reliable p

Cited by 0SourcePDFScholar
2026

Distilling Unsigned Distance Function for Surface Reconstruction from 3D Gaussian Splatting

CVPR 2026

Unsigned distance fields (UDFs) are well suited for representing open surfaces, but learning them from multi-view images is challenging because ground-truth surfaces are unavailable for supervision in most cases and the gradient of a UDF is undefined on the underlying surface. Prior methods optimize

Cited by 0SourceScholar
2026

EchoBat: Echo-Vision Enhancement and Echo-Layered Sampling for Video LLMs Hallucination Mitigation

AAAI 2026technical

Recent advancements in multimodal large language models (MLLMs) have shown remarkable progress in video understanding. However, video MLLMs (VideoMLLMs) still suffer from hallucinations, generating nonsensical or irrelevant content. This issue partly stems from over-reliance on pre-trained knowledge

Cited by 0SourcePDFScholar
2026

FlyPrompt: Brain-Inspired Random-Expanded Routing with Temporal-Ensemble Experts for General Continual Learning

ICLR 2026poster

General continual learning (GCL) challenges intelligent systems to learn from single-pass, non-stationary data streams without clear task boundaries. While recent advances in continual parameter-efficient tuning (PET) of pretrained models show promise, they typically rely on multiple training epochs…

Cited by 0SourcecodeScholar
2026

MoTRa: Motion-Aware Target Representation Learning for End-to-End Multi-Object Tracking

IJCAI 2026

Multi-object tracking (MOT) has long faced challenges with identity switches, especially for targets with low appearance discriminability and complex motion. Existing end-to-end trackers typically enhance robustness by modeling long-term temporal information across target-level representations, yet

Cited by 0Scholar
2026

Multi-Modal Fact Knowledge Generation for Imbalanced Cross-Source Entity Alignment

AAAI 2026technical

Multi-modal imbalanced cross-source entity alignment aims to identify equivalent entity pairs across multi-modal knowledge graphs (MMKGs) that encompass diverse data sources with imbalanced modality, which poses significant challenges due to the non-uniform distribution of information across differe

Cited by 0SourcePDFScholar
2026

PDAR-RSITR: A Progressive Decoupling-Aggregation-Refinement Framework for Remote Sensing Image-Text Retrieval

IJCAI 2026

Remote Sensing Image-Text Retrieval (RSITR) aims to achieve precise retrieval between remote sensing images and textual descriptions. However, existing methods neglect the multi-dimensional cognitive attributes inherent in remote sensing data and struggle to handle them simultaneously, leading to su

Cited by 0Scholar
2026

SCo-Cloud: Satellite Constellation Collaboration for Cloud-Aware Onboard-Computed Imaging and Transmission

AAAI 2026technical

Satellite-acquired optical remote sensing imagery is extensively applied in time-critical applications like traffic surveillance and evaluation of natural disasters. However, clouds, as a common atmospheric phenomenon, frequently obscure observation. Current approaches aim to restore visibility in c

Cited by 0SourcePDFScholar
2026

UniDrag: Unified Multi-Field Prediction and Robust Shape Optimization for Vehicle Aerodynamics

ICML 2026poster

High-fidelity vehicle aerodynamics analysis is bottlenecked by costly CFD simulations. Neural surrogates accelerate prediction but lack inverse design capabilities, while existing generative optimization methods suffer from unstable convergence and frequent engineering constraint violations. We pres…

Cited by 0SourceScholar
2026

VKG-QA: Visual Knowledge Graph-based Question Answer for Large Multimodal Models

CVPR 2026

Understanding and reasoning over structured knowledge is a fundamental capability for intelligent systems. While Large Language Models (LLMs) have leveraged textual knowledge graphs for relational reasoning, linearizing graph structures into text often leads to token inefficiency and loss of higher-

Cited by 0SourcecodeScholar
2025

Context-Alignment: Activating and Enhancing LLMs Capabilities in Time Series

ICLR 2025poster

Recently, leveraging pre-trained Large Language Models (LLMs) for time series (TS) tasks has gained increasing attention, which involves activating and enhancing LLMs' capabilities. Many methods aim to activate LLMs' capabilities based on token-level alignment, but overlook LLMs' inherent strength i…

2025

Learning SQL Like a Human: Structure-Aware Curriculum Learning for Text-to-SQL Generation

EMNLP 2025

The Text-to-SQL capabilities of large language allow users to interact with databases using natural language. While current models struggle with handling complex queries, especially involving multi-table joins and reasoning. To address this gap, we propose to construct a model, namely SAC-SQL, with

Cited by 0SourcePDFScholar
2025

Lightweight Contenders: Navigating Semi-Supervised Text Mining through Peer Collaboration and Self Transcendence

NAACL 2025findings

The semi-supervised learning (SSL) strategy in lightweight models requires reducing annotated samples and facilitating cost-effective inference. However, the constraint on model parameters, imposed by the scarcity of training labels, limits the SSL performance. In this paper, we introduce PS-NET, a…

2025

Multimodal Knowledge Retrieval-Augmented Iterative Alignment for Satellite Commonsense Conversation

IJCAI 2025

Satellite technology has significantly influenced our daily lives, manifested in applications such as navigation and communication. With its development, a vast amount of multimodal satellite commonsense data has been generated, thus leading to an urgent demand for conversation about satellite data.

Cited by 0SourcePDFScholar
2025

N2GON: Neural Networks for Graph-of-Net with Position Awareness

ICML 2025poster

Graphs, fundamental in modeling various research subjects such as computing networks, consist of nodes linked by edges. However, they typically function as components within larger structures in real-world scenarios, such as in protein-protein interactions where each protein is a graph in a larger n…

Cited by 0SourcePDFScholar
2025

OS-GCL: A One-Shot Learner in Graph Contrastive Learning

IJCAI 2025

Graph contrastive learning (GCL) enhances the self-supervised learning capacity for graph representation learning. Nevertheless, the previous research has neglected to consider one fundamental nature of GCL -- graph contrastive learning operates as a one-shot learner, guided by the widely utilized n

Cited by 0SourcePDFScholar
2025

Right Time to Learn: Promoting Generalization via Bio-inspired Spacing Effect in Knowledge Distillation

ICML 2025poster

Knowledge distillation (KD) is a powerful strategy for training deep neural networks (DNNs). While it was originally proposed to train a more compact “student” model from a large “teacher” model, many recent efforts have focused on adapting it as an effective way to promote generalization of the mod…

2025

T-T: Table Transformer for Tagging-based Aspect Sentiment Triplet Extraction

IJCAI 2025

Aspect sentiment triplet extraction (ASTE) aims to extract triplets composed of aspect terms, opinion terms, and sentiment polarities from given sentences. The table tagging method is a popular approach to addressing this task, which encodes a sentence into a 2-dimensional table, allowing for the ta

2025

Towards Explaining the Power of Constant-depth Graph Neural Networks for Structured Linear Programming

ICLR 2025poster

Graph neural networks (GNNs) have recently emerged as powerful tools for solving complex optimization problems, often being employed to approximate solution mappings. Empirical evidence shows that even shallow GNNs (with fewer than ten layers) can achieve strong performance in predicting optimal sol…

Cited by 0SourcePDFScholar
2025

Variational Multi-Modal Hypergraph Attention Network for Multi-Modal Relation Extraction

IJCAI 2025

Multi-modal relation extraction (MMRE) is a challenging task that seeks to identify relationships between entities with textual and visual attributes. However, existing methods struggle to handle the complexities posed by multiple entity pairs within a single sentence that share similar contextual i

2025

When GNNs meet symmetry in ILPs: an orbit-based feature augmentation approach

ICLR 2025poster

A common characteristic in integer linear programs (ILPs) is symmetry, allowing variables to be permuted without altering the underlying problem structure. Recently, GNNs have emerged as a promising approach for solving ILPs. However, a significant challenge arises when applying GNNs to ILPs with s…

2024

Collapse-Aware Triplet Decoupling for Adversarially Robust Image Retrieval

ICML 2024poster

Adversarial training has achieved substantial performance in defending image retrieval against adversarial examples. However, existing studies in deep metric learning (DML) still suffer from two major limitations: *weak adversary* and *model collapse*. In this paper, we address these two limitations…

2024

Few-Shot Multimodal Named Entity Recognition Based on Mutlimodal Causal Intervention Graph

COLING 2024main

Multimodal Named Entity Recognition (MNER) models typically require a significant volume of labeled data for effective training to extract relations between entities. In real-world scenarios, we frequently encounter unseen relation types. Nevertheless, existing methods are predominantly tailored for…

Cited by 1SourcePDFScholar
2024

Focus on Hiders: Exploring Hidden Threats for Enhancing Adversarial Training

CVPR 2024poster

Adversarial training is often formulated as a min-max problem however concentrating only on the worst adversarial examples causes alternating repetitive confusion of the model i.e. previously defended or correctly classified samples are not defensible or accurately classifiable in subsequent adversa…

Cited by 6SourcePDFScholar
2024

HiFT: A Hierarchical Full Parameter Fine-Tuning Strategy

EMNLP 2024main

Full-parameter fine-tuning (FPFT) has become the go-to choice for adapting language models (LMs) to downstream tasks due to its excellent performance. As LMs grow in size, fine-tuning the full parameters of LMs requires a prohibitively large amount of GPU memory. Existing approaches utilize zeroth-o…

2024

LLM-based Multi-Level Knowledge Generation for Few-shot Knowledge Graph Completion

IJCAI 2024poster

Knowledge Graphs (KGs) are pivotal in various NLP applications but often grapple with incompleteness, especially due to the long-tail problem where infrequent, unpopular relationships drastically reduce the KG completion performance. In this paper, we focus on Few-shot Knowledge Graph Completion (FK…

Cited by 6SourcePDFScholar
2024

MKGL: Mastery of a Three-Word Language

NeurIPS 2024spotlight

Large language models (LLMs) have significantly advanced performance across a spectrum of natural language processing (NLP) tasks. Yet, their application to knowledge graphs (KGs), which describe facts in the form of triplets and allow minimal hallucinations, remains an underexplored frontier. In th…

Cited by 1SourcePDFScholar
2024

On the Power of Small-size Graph Neural Networks for Linear Programming

NeurIPS 2024poster

Graph neural networks (GNNs) have recently emerged as powerful tools for addressing complex optimization problems. It has been theoretically demonstrated that GNNs can universally approximate the solution mapping functions of linear programming (LP) problems. However, these theoretical results typic…

Cited by 0SourcePDFScholar
2024

Physical 3D Adversarial Attacks against Monocular Depth Estimation in Autonomous Driving

CVPR 2024poster

Deep learning-based monocular depth estimation (MDE) extensively applied in autonomous driving is known to be vulnerable to adversarial attacks. Previous physical attacks against MDE models rely on 2D adversarial patches so they only affect a small localized region in the MDE map but fail under vari…

2024

ReGCL: Rethinking Message Passing in Graph Contrastive Learning

AAAI 2024technical

Graph contrastive learning (GCL) has demonstrated remarkable efficacy in graph representation learning. However, previous studies have overlooked the inherent conflict that arises when employing graph neural networks (GNNs) as encoders for node-level contrastive learning. This conflict pertains to t…

2024

Reference Line Network: On Simultaneous Gaussian Line Detection and Connection Graph Inference

ICASSP 2024accepted

Reference line detection is a challenging problem due to localization uncertainty and severe occlusion. To deal with the two issues, we propose a general framework for reference line detection with two modules: Gaussian line detection and connection graph inference. The first module outputs a set of…

Cited by 0SourceScholar
2024

Synonym Replacement and Generation Enhancement for Document Augmentation

ICASSP 2024accepted

Document AI, or Document Intelligence pertains to the technology used for document comprehension and analysis. Given the multimodality of documents, the importance of multimodal learning cannot be overstated in the field of document intelligence research. Multimodal data augmentation, as a crucial a…

Cited by 0SourceScholar
2023

BPNet: Bézier Primitive Segmentation on 3D Point Clouds

IJCAI 2023poster

This paper proposes BPNet, a novel end-to-end deep learning framework to learn Bézier primitive segmentation on 3D point clouds. The existing works treat different primitive types separately, thus limiting them to finite shape categories. To address this issue, we seek a generalized primitive segmen…

2023

Contrastive Learning with Generated Representations for Inductive Knowledge Graph Embedding

ACL 2023findings

With the evolution of Knowledge Graphs (KGs), new entities emerge which are not seen before. Representation learning of KGs in such an inductive setting aims to capture and transfer the structural patterns from existing entities to new entities. However, the performance of existing methods in induct…

2023

Discrete Point-Wise Attack Is Not Enough: Generalized Manifold Adversarial Attack for Face Recognition

CVPR 2023poster

Classical adversarial attacks for Face Recognition (FR) models typically generate discrete examples for target identity with a single state image. However, such paradigm of point-wise attack exhibits poor generalization against numerous unknown states of identity and can be easily defended. In this…

2023

Dual-Gated Fusion with Prefix-Tuning for Multi-Modal Relation Extraction

ACL 2023findings

Multi-Modal Relation Extraction (MMRE) aims at identifying the relation between two entities in texts that contain visual clues. Rich visual content is valuable for the MMRE task, but existing works cannot well model finer associations among different modalities, failing to capture the truly helpful…

2023

Hebbian and Gradient-based Plasticity Enables Robust Memory and Rapid Learning in RNNs

ICLR 2023poster

Rapidly learning from ongoing experiences and remembering past events with a flexible memory system are two core capacities of biological intelligence. While the underlying neural mechanisms are not fully understood, various evidence supports that synaptic plasticity plays a critical role in memory…

2023

Learning to Initialize: Can Meta Learning Improve Cross-task Generalization in Prompt Tuning?

ACL 2023long

Prompt tuning (PT) which only tunes the embeddings of an additional sequence of tokens per task, keeping the pre-trained language model (PLM) frozen, has shown remarkable performance in few-shot learning. Despite this, PT has been shown to rely heavily on good initialization of the prompt embeddings…

Cited by 15SourcePDFScholar
2023

Multi-Modal Knowledge Graph Transformer Framework for Multi-Modal Entity Alignment

EMNLP 2023long findings

Multi-Modal Entity Alignment (MMEA) is a critical task that aims to identify equivalent entity pairs across multi-modal knowledge graphs (MMKGs). However, this task faces challenges due to the presence of different types of information, including neighboring entities, multi-modal attributes, and ent…

Cited by 0SourcecodeScholar
2022

Alleviating Sparsity of Open Knowledge Graphs with Ternary Contrastive Learning

EMNLP 2022finding

Sparsity of formal knowledge and roughness of non-ontological construction make sparsity problem particularly prominent in Open Knowledge Graphs (OpenKGs). Due to sparse links, learning effective representation for few-shot entities becomes difficult. We hypothesize that by introducing negative samp…

2022

CoSCL: Cooperation of Small Continual Learners Is Stronger than a Big One

ECCV 2022poster

"Continual learning requires incremental compatibility with a sequence of tasks. However, the design of model architecture remains an open question: In general, learning all tasks with a shared set of parameters suffers from severe interference between tasks; while learning each task with a dedicate…

2022

How Does Knowledge Graph Embedding Extrapolate to Unseen Data: A Semantic Evidence View

AAAI 2022technical

Knowledge Graph Embedding (KGE) aims to learn representations for entities and relations. Most KGE models have gained great success, especially on extrapolation scenarios. Specifically, given an unseen triple (h, r, t), a trained model can still correctly predict t from (h, r, ?), or h from (?, r, t…

2021

AFEC: Active Forgetting of Negative Transfer in Continual Learning

NeurIPS 2021poster

Continual learning aims to learn a sequence of tasks from dynamic data distributions. Without accessing to the old training samples, knowledge transfer from the old tasks to each new task is difficult to determine, which might be either positive or negative. If the old knowledge interferes with the…

2021

AdaAttN: Revisit Attention Mechanism in Arbitrary Neural Style Transfer

ICCV 2021poster

Fast arbitrary neural style transfer has attracted widespread attention from academic, industrial and art communities due to its flexibility in enabling various applications. Existing solutions either attentively fuse deep style feature into deep content feature without considering feature distribut…

Cited by 444PDFcodeScholar
2021

Generation and Extraction Combined Dialogue State Tracking with Hierarchical Ontology Integration

EMNLP 2021main

Recently, the focus of dialogue state tracking has expanded from single domain to multiple domains. The task is characterized by the shared slots between domains. As the scenario gets more complex, the out-of-vocabulary problem also becomes severer. Current models are not satisfactory for solving th…

Cited by 11SourcePDFScholar
2021

Reasoning Operational Decisions for Robots via Time Series Causal Inference

ICRA 2021poster

Justifying operational decisions for robots is a challenging task as the operator or the robot itself has to understand the underlying physical interaction between the robot and the environment to predict the potential outcome. It is desirable to understand how the decision influences the operationa…

Cited by 9SourceScholar
2020

HPRNN: A Hierarchical Sequence Prediction Model for Long-Term Weather Radar Echo Extrapolation

ICASSP 2020accepted

Weather radar echo extrapolation has been one of the most important means for weather forecasting and precipitation nowcasting. However, the effective forecasting time of the most current extrapolation methods is usually short. In this paper, to meet the demand for long-term extrapolation in actual…

Cited by 0SourceScholar
2020

Stochastic Batch Augmentation with An Effective Distilled Dynamic Soft Label Regularizer

IJCAI 2020poster

Data augmentation have been intensively used in training deep neural network to improve the generalization, whether in original space (e.g., image space) or representation space. Although being successful, the connection between the synthesized data and the original data is largely ignored in traini…

Cited by 0SourcePDFScholar
2019

Group-Wise Deep Object Co-Segmentation With Co-Attention Recurrent Neural Network

ICCV 2019poster

Effective feature representations which should not only express the images individual properties, but also reflect the interaction among group images are essentially crucial for real-world co-segmentation. This paper proposes a novel end-to-end deep learning approach for group-wise object co-segment…

Cited by 75PDFScholar