计算机科学
人工智能
数据预处理
大数据
数据质量
数据科学
管道(软件)
技术债务
机器学习
数据收集
质量(理念)
软件
数据挖掘
软件开发
工程类
公制(单位)
哲学
运营管理
统计
数学
认识论
程序设计语言
标识
DOI:10.1016/j.dsm.2023.06.001
摘要
Artificial intelligence (AI) relies on data and algorithms. State-of-the-art (SOTA) AI smart algorithms have been developed to improve the performance of AI-oriented structures. However, model-centric approaches are limited by the absence of high-quality data. Data-centric AI is an emerging approach for solving machine learning (ML) problems. It is a collection of various data manipulation techniques that allow ML practitioners to systematically improve the quality of the data used in an ML pipeline. However, data-centric AI approaches are not well documented. Researchers have conducted various experiments without a clear set of guidelines. This survey highlights six major data-centric AI aspects that researchers are already using to intentionally or unintentionally improve the quality of AI systems. These include big data quality assessment, data preprocessing, transfer learning, semi-supervised learning, MLOps, and the effect of adding more data. In addition, it highlights recent data-centric techniques adopted by ML practitioners. We addressed how adding data might harm datasets and how HoloClean can be used to restore and clean them. Finally, we discuss the causes of technical debt in AI. Technical debt builds up when software design and implementation decisions run into "or outright collide with "business goals and timelines. This survey lays the groundwork for future data-centric AI discussions by summarizing various data-centric approaches.
科研通智能强力驱动
Strongly Powered by AbleSci AI