← Search

kaiwen wei

23 accepted papers

2026

GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding

CVPR 2026

Recent advances in multimodal large language models (MLLMs) have led to remarkable progress in visual grounding, enabling fine-grained cross-modal alignment between textual queries and image regions. However, transferring such capabilities to remote sensing imagery remains challenging, as targets ar

Cited by 0SourcecodeScholar
2026

How Far Can LLM Agents Reason with Tables? Benchmarking Multi-Turn Agentic Table Question Answering in the Wild

ICML 2026poster

Recent advances in large language models (LLMs) have substantially expanded the scope of Table Question Answering (TableQA). However, existing benchmarks primarily treat TableQA as a passive, single-turn natural language understanding task, lacking the capacity to evaluate autonomous reasoning and t…

Cited by 0SourceScholar
2026

MIRAGE: Scaling Test-Time Inference with Parallel Graph-Retrieval-Augmented Reasoning Chains

AAAI 2026technical

Large reasoning models (LRMs) have shown significant progress in test-time scaling through chain-of-thought prompting. Current approaches like search-o1 integrate retrieval augmented generation (RAG) into multi-step reasoning processes but rely on a single, linear reasoning path while incorporating

Cited by 0SourcePDFScholar
2026

SMLDR: Spectral Memory Learner with Dual-Retrieval for Time Series Forecasting

IJCAI 2026

Time series forecasting aims to predict future values using historical observations, which is crucial for many practical applications with complex temporal dynamics. Recent frequency-domain forecasting methods have utilized spectral representations to model periodicity, but they usually rely on an i

Cited by 0Scholar
2025

Chain-of-Specificity: Enhancing Task-Specific Constraint Adherence in Large Language Models

COLING 2025main

Large Language Models (LLMs) exhibit remarkable generative capabilities, enabling the generation of valuable information. Despite these advancements, previous research found that LLMs sometimes struggle with adhering to specific constraints, such as being in a specific place or at a specific time, a…

Cited by 1SourcePDFScholar
2025

EEG Decoding and Visual Reconstruction via 3D Geometric with Nonstationarity Modelling

ICASSP 2025accepted

Electroencephalogram (EEG) signal processing has advanced in revealing the mechanisms of human visual perception, but existing methods often overlook two key EEG properties: (1) 3D geometric relationships between EEG electrodes, which reflects the ability to model the brain in stereoscopic terms; an…

Cited by 0SourceScholar
2025

FedLEKE: Federated Locate-then-Edit Knowledge Editing for Multi-Client Collaboration

ACL 2025finding

Locate-then-Edit Knowledge Editing (LEKE) is a key technique for updating large language models (LLMs) without full retraining. However, existing methods assume a single-user setting and become inefficient in real-world multi-client scenarios, where decentralized organizations (e.g., hospitals, fina…

2025

Latent Distribution Decouple for Uncertain-Aware Multimodal Multi-label Emotion Recognition

ACL 2025finding

Multimodal multi-label emotion recognition (MMER) aims to identify the concurrent presence of multiple emotions in multimodal data. Existing studies primarily focus on improving fusion strategies and modeling modality-to-label dependencies. However, they often overlook the impact of aleatoric uncert…

2025

P²Net: Parallel Pointer-based Network for Key Information Extraction with Complex Layouts

ACL 2025finding

Key Information Extraction (KIE) is a challenging multimodal task aimed at extracting structured value entities from visually rich documents. Despite recent advancements, two major challenges remain. First, existing datasets typically feature fixed layouts and a limited set of entity categories, whi…

Cited by 0SourcePDFScholar
2025

SARA: Salience-Aware Reinforced Adaptive Decoding for Large Language Models in Abstractive Summarization

ACL 2025long

LLMs have improved the fluency and informativeness of abstractive summarization but remain prone to hallucinations, where generated content deviates from the source document. Recent PMI decoding strategies mitigate over-reliance on prior knowledge by comparing output probabilities with and without s…

Cited by 0SourcePDFScholar
2025

T2R-BENCH: A Benchmark for Real World Table-to-Report Task

EMNLP 2025

Extensive research has been conducted to explore the capabilities of large language models (LLMs) in table reasoning. However, the essential task of transforming tables information into reports remains a significant challenge for industrial applications. This task is plagued by two critical issues:

2024

CAMEL: Capturing Metaphorical Alignment with Context Disentangling for Multimodal Emotion Recognition

AAAI 2024technical

Understanding the emotional polarity of multimodal content with metaphorical characteristics, such as memes, poses a significant challenge in Multimodal Emotion Recognition (MER). Previous MER researches have overlooked the phenomenon of metaphorical alignment in multimedia content, which involves n…

Cited by 9SourcePDFScholar
2024

GOME: Grounding-based Metaphor Binding With Conceptual Elaboration For Figurative Language Illustration

EMNLP 2024main

The illustration or visualization of figurative language, such as linguistic metaphors, is an emerging challenge for existing Large Language Models (LLMs) and multimodal models. Due to their comparison of seemingly unrelated concepts in metaphors, existing LLMs have a tendency of over-literalization…

Cited by 0SourcePDFScholar
2024

Multimodal Event Causality Reasoning with Scene Graph Enhanced Interaction Network

AAAI 2024technical

Multimodal event causality reasoning aims to recognize the causal relations based on the given events and accompanying image pairs, requiring the model to have a comprehensive grasp of visual and textual information. However, existing studies fail to effectively model the relations of the objects wi…

Cited by 3SourcePDFScholar
2024

TAeKD: Teacher Assistant Enhanced Knowledge Distillation for Closed-Source Multilingual Neural Machine Translation

COLING 2024main

Knowledge Distillation (KD) serves as an efficient method for transferring language knowledge from open-source large language models (LLMs) to more computationally efficient models. However, challenges arise when attempting to apply vanilla KD methods to transfer knowledge from closed-source Multili…

2024

Video Event Extraction with Multi-View Interaction Knowledge Distillation

AAAI 2024technical

Video event extraction (VEE) aims to extract key events and generate the event arguments for their semantic roles from the video. Despite promising results have been achieved by existing methods, they still lack an elaborate learning strategy to adequately consider: (1) inter-object interaction, whi…

Cited by 2SourcePDFScholar
2023

1% VS 100%: Parameter-Efficient Low Rank Adapter for Dense Predictions

CVPR 2023poster

Fine-tuning large-scale pre-trained vision models to downstream tasks is a standard technique for achieving state-of-the-art performance on computer vision benchmarks. However, fine-tuning the whole model with millions of parameters is inefficient as it requires storing a same-sized new model copy f…

Cited by 53SourcePDFScholar
2023

Event Causality Extraction via Implicit Cause-Effect Interactions

EMNLP 2023long main

Event Causality Extraction (ECE) aims to extract the cause-effect event pairs from the given text, which requires the model to possess a strong reasoning ability to capture event causalities. However, existing works have not adequately exploited the interactions between the cause and effect event th…

Cited by 0SourceScholar
2023

Guide the Many-to-One Assignment: Open Information Extraction via IoU-aware Optimal Transport

ACL 2023long

Open Information Extraction (OIE) seeks to extract structured information from raw text without the limitations of close ontology. Recently, the detection-based OIE methods have received great attention from the community due to their parallelism. However, as the essential step of those models, how…

Cited by 14SourcePDFScholar
2023

Let Me Check the Examples: Enhancing Demonstration Learning via Explicit Imitation

ACL 2023short

Demonstration learning aims to guide the prompt prediction by providing answered demonstrations in the few shot settings. Despite achieving promising results, existing work only concatenates the answered examples as demonstrations to the prompt template (including the raw context) without any additi…

2023

Narrative Order Aware Story Generation via Bidirectional Pretraining Model with Optimal Transport Reward

EMNLP 2023long findings

To create a captivating story, a writer often plans a sequence of logically coherent events and ingeniously manipulates the narrative order to generate flashback in place. However, existing storytelling systems suffer from both insufficient understanding of event correlations and inadequate awarenes…

Cited by 0SourceScholar
2022

Assist Non-native Viewers: Multimodal Cross-Lingual Summarization for How2 Videos

EMNLP 2022main

Multimodal summarization for videos aims to generate summaries from multi-source information (videos, audio transcripts), which has achieved promising progress. However, existing works are restricted to monolingual video scenarios, ignoring the demands of non-native video viewers to understand the c…

2021

Trigger is Not Sufficient: Exploiting Frame-aware Knowledge for Implicit Event Argument Extraction

ACL 2021long

Implicit Event Argument Extraction seeks to identify arguments that play direct or implicit roles in a given event. However, most prior works focus on capturing direct relations between arguments and the event trigger. The lack of reasoning ability brings many challenges to the extraction of implici…

Cited by 76SourcePDFScholar