← Search

Lei Fang

13 accepted papers

2026

Analyzing and Mitigating Object Hallucination: A Training Bias Perspective

AAAI 2026technical

As scaling up training data has significantly improved the general multimodal capabilities of Large Vision-Language Models (LVLMs), they still suffer from the hallucination issue, generating text that is inconsistent with the visual input. This phenomenon motivates us to systematically investigate t

Cited by 0SourcePDFScholar
2026

BFMPF-Net: Bidirectional Frequency-Domain Modulation Progressive Fusion Network for Road Crack Segmentation

ICRA 2026poster

Recently, deep learning–based methods for road crack segmentation have achieved promising performance, particularly in robotic vision applications such as automated inspection and maintenance. However, most frequency-domain methods employ a decoupled processing strategy, overlooking the dynamic modu…

Cited by 0Scholar
2026

Towards Effective Code-Integrated Reasoning

AAAI 2026technical

In this paper, we investigate code-integrated reasoning (CIR), where models generate code when necessary and integrate feedback by executing it through a code interpreter. To acquire this capability, models must learn when and how to use external code tools effectively, which is supported by tool-au

Cited by 0SourcePDFScholar
2025

CAFE: Retrieval Head-based Coarse-to-Fine Information Seeking to Enhance Multi-Document QA Capability

EMNLP 2025

Advancements in Large Language Models (LLMs) have extended their input context length, yet they still struggle with retrieval and reasoning in long-context inputs. Existing methods propose to utilize the prompt strategy and Retrieval-Augmented Generation (RAG) to alleviate this limitation. However,

2025

Integrating Spectro-Temporal Cross Aggregation and Multi-Scale Dynamic Learning for Audio Deepfake Detection

ICASSP 2025accepted

Audio deepfake refers to the technology of synthesizing speech using deep learning or large model algorithms. Compared to human voice, synthetic deepfake speech exhibits artifacts at global and local levels, which can be leveraged by audio deepfake detection (ADD) to distinguish real and fake speech…

Cited by 0SourceScholar
2025

Rethinking Invariance in In-context Learning

ICLR 2025poster

In-Context Learning (ICL) has emerged as a pivotal capability of auto-regressive large language models, yet it is hindered by a notable sensitivity to the ordering of context examples regardless of their mutual independence. To address this issue, recent studies have introduced several variant algor…

2025

SimpleDeepSearcher: Deep Information Seeking via Web-Powered Reasoning Trajectory Synthesis

EMNLP 2025

Retrieval-augmented generation (RAG) systems have advanced large language models (LLMs) in complex deep search scenarios requiring multi-step reasoning and iterative information retrieval. However, existing approaches face critical limitations that lack high-quality training trajectories or suffer f

2025

Smart-Searcher: Incentivizing the Dynamic Knowledge Acquisition of LLMs via Reinforcement Learning

EMNLP 2025

Large Language Models (LLMs) are powerful but prone to hallucinations due to static knowledge. Retrieval-Augmented Generation (RAG) helps by injecting external information, but current methods often are costly, generalize poorly, or ignore the model’s internal knowledge.In this paper, we introduce S

2024

A White-Box False Positive Adversarial Attack Method on Contrastive Loss Based Offline Handwritten Signature Verification Models

AISTATS 2024poster

In this paper, we tackle the challenge of white-box false positive adversarial attacks on contrastive loss based offline handwritten signature verification models. We propose a novel attack method that treats the attack as a style transfer between closely related but distinct writing styles. To guid…

2024

Accurate Forgetting for Heterogeneous Federated Continual Learning

ICLR 2024poster

Recent years have witnessed a burgeoning interest in federated learning (FL). However, the contexts in which clients engage in sequential learning remain under- explored. Bridging FL and continual learning (CL) gives rise to a challenging practical problem: federated continual learning (FCL). Existi…

2024

CLAP: Collaborative Adaptation for Patchwork Learning

ICLR 2024spotlight

In this paper, we investigate a new practical learning scenario, where the data distributed in different sources/clients are typically generated with various modalities. Existing research on learning from multi-source data mostly assume that each client owns the data of all modalities, which may lar…

Cited by 1SourcePDFScholar
2022

Leveraging Sparse Coding for EEG Based Emotion Recognition in Shooting

ICASSP 2022accepted

Emotion recognition in shooting is of great importance for improving athletes’ training methods. However, there is no open and high confident electroencephalography (EEG) dataset about shooting due to the difficulty of data acquisition, which made it a challenge for related studies. In this paper, w…

Cited by 0SourceScholar
2021

TWT: Table with Written Text for Controlled Data-to-Text Generation

EMNLP 2021finding

Large pre-trained neural models have recently shown remarkable progress in text generation. In this paper, we propose to generate text conditioned on the structured data (table) and a prefix (the written text) by leveraging the pre-trained models. We present a new data-to-text dataset, Table with Wr…

Cited by 13SourcePDFScholar