Enhancing Information Extraction with METORIE: A Metaphor and Trap-Based Dataset for Cross-Domain Fine-Tuning
Zhengyuan Pan, Yilian Peng, Zhongquan Jian, Yanhao Chen, Wentao Qiu, Haonan Ma, Junfeng Yao, Meihong Wang
Abstract
This research proposes the METORIE dataset <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>, a novel resource designed to improve the reasoning capabilities of large language models (LLMs), such as LLaMA3 and GLM4, in information extraction (IE) tasks. The METORIE dataset is derived from brain teasers that incorporate complex logical and metaphorical elements and is designed to train LLMs to navigate intricate reasoning paths and interpret layered expressions. Our findings demonstrate that the METORIE dataset markedly enhances LLMs’ performance across both general and specialized IE tasks. The results of fine-tuning with the METORIE dataset, mixed with a small number of IE datasets, are close to, if not exceeding, those of LLMs of the same parametric size on IE tasks using much larger datasets. Through controlled experiments, we establish that metaphors of medium complexity optimize IE performance, while higher complexities tend to overstretch LLMs’ inference limits. METORIE-fine-tuned LLMs also demonstrate exceptional performance in legal and medical domains, suggesting that enhanced metaphor understanding and logical deduction are key to improving LLMs’ adaptability and efficiency in vertical domains.
BibTeX
@inproceedings{icassp2025_enhancinginforma,
title = {Enhancing Information Extraction with METORIE: A Metaphor and Trap-Based Dataset for Cross-Domain Fine-Tuning},
author = {Zhengyuan Pan and Yilian Peng and Zhongquan Jian and Yanhao Chen and Wentao Qiu and Haonan Ma and Junfeng Yao and Meihong Wang and Qingqiang Wu},
booktitle = {ICASSP 2025},
year = {2025}
}