SCIR: A Self-Correcting Iterative Refinement Framework for Enhanced Information Extraction Based on Schema
Yushen Fang, Jianjun Li, Mingqian Ding, Chang Liu, Xinchi Zou, Wenqi Yang
Abstract
Although Large language Model (LLM)-powered information extraction (IE) systems have shown impressive capabilities, current fine-tuning paradigms face two major limitations: high training costs and difficulties in aligning with LLM preferences. To address these issues, we propose a novel universal IE paradigm—the Self-Correcting Iterative Refinement (SCIR) framework—along with a Multi-task Bilingual (Chinese-English) Self-Correcting (MBSC) dataset containing over 100,000 entries. The SCIR framework achieves plug-and-play compatibility with existing LLMs and IE systems through its Dual-Path Self-Correcting module and feedback-driven optimization, thereby significantly reducing training costs. Concurrently, the MBSC dataset tackles the challenge of preference alignment by indirectly distilling GPT-4
BibTeX
@inproceedings{aaai2026_sciraselfcorrect,
title = {SCIR: A Self-Correcting Iterative Refinement Framework for Enhanced Information Extraction Based on Schema},
author = {Yushen Fang and Jianjun Li and Mingqian Ding and Chang Liu and Xinchi Zou and Wenqi Yang},
booktitle = {AAAI 2026},
year = {2026}
}