Evidence Reasoning and Curriculum Learning for Document-level Relation Extraction

关系抽取计算机科学关系（数据库）判决人工智能任务（项目管理）自然语言处理信息抽取强化学习情报检索数据挖掘经济管理

作者

Tianyu Xu,Jianfeng Qu,Wen Hua,Zhixu Li,Jiajie Xu,An Liu,Lei Zhao,Xiaofang Zhou

出处

期刊：IEEE Transactions on Knowledge and Data Engineering [Institute of Electrical and Electronics Engineers]
日期：2023-01-01 卷期号：: 1-14 被引量：1

标识

DOI：10.1109/tkde.2023.3292974

摘要

Document-level Relation Extraction (RE) is a promising task aiming at identifying relations of multiple entity pairs in a document. Compared with the sentence-level counterpart, it has raised two significant challenges: a) In most cases, a relational fact can be adequately expressed via a small subset of sentences from the document, namely evidence. But the traditional method cannot model such strong semantic correlations between evidence sentences that collaborate to describe a specific relation; b) The data of this task is extremely long-tail in terms of too many NA instances and imbalanced relational types. Such data can mislead the tail prediction bias to the head categories in the RE model. In this paper, we present a novel E vidence reasoning and C urriculum learning method for D oc RE (DRE-EC) to address these challenges. Particularly, we first formulate evidence extraction as a sequential decision problem through a crafted reinforcement learning mechanism with an efficient path searching strategy to reduce the action space. Providing the evidence for each entity pair as a customized-filtered document in advance helps infer the relations better. To address the long-tail issue, we further develop a hybrid curriculum learning method at the NA-level (NC) and relation-level (RC) with our customized difficulty measure score. In NC, the NA samples are scheduled in an easy-to-hard scheme and gradually added, resulting in the data distribution from ideal and balanced to real and unbalanced. In RC, the scheme is switched into hard-to-easy to enhance the hard and tail samples. In addition, we propose a new Equalization adaptive Focal Loss(EFLoss) that can adjust to the changing data distribution and focus more on the tail categories. We conduct various experiments on two document-level RE benchmarks and achieve a remarkable improvement over previous competitive baselines. Furthermore, we provide detailed analyses of the advantages and effectiveness of our method.

求助该文献

最长约 10秒，即可获得该文献文件

Evidence Reasoning and Curriculum Learning for Document-level Relation Extraction

今日热心研友