计算机科学
正确性
任务(项目管理)
人工智能
自然语言处理
基石
推论
算法
艺术
视觉艺术
经济
管理
作者
Mingyue Jiang,Houzhen Bao,Kaiyi Tu,Xiao-Yi Zhang,Zuohua Ding
标识
DOI:10.1109/issre52982.2021.00033
摘要
Natural language inference (NLI) is a fundamental NLP task that forms the cornerstone of deep natural language understanding. Unfortunately, evaluation of NLI models is challenging. On one hand, due to the lack of test oracles, it is difficult to automatically judge the correctness of NLI's prediction results. On the other hand, apart from knowing how well a model performs, there is a further need for understanding the capabilities and characteristics of different NLI models. To mitigate these issues, we propose to apply the technique of metamorphic testing (MT) to NLI. We identify six categories of metamorphic relations, covering a wide range of properties that are expected to be possessed by NLI task. Based on this, MT can be conducted on NLI models without using test oracles, and MT results are able to interpret NLI models' capabilities from varying aspects. We further demonstrate the validity and effectiveness of our approach by conducting experiments on five NLI models. Our experiments expose a large number of prediction failures from subject NLI models, and also yield interpretations for common characteristics of NLI models.
科研通智能强力驱动
Strongly Powered by AbleSci AI