Bug characterization in machine learning-based systems

被引：0

作者：

Mohammad Mehdi Morovati

Amin Nikanjam

Florian Tambon

Foutse Khomh

Zhen Ming (Jack) Jiang

机构：

[1] Polytechnique Montréal,SWAT Lab.

[2] York University,undefined

来源：

Empirical Software Engineering | 2024年 / 29卷

关键词：

Software bug; Software testing; ML-based systems; ML bug; Deep learning; Software maintenance; Empirical study;

D O I：

暂无

中图分类号：

学科分类号：

摘要：

The rapid growth of applying Machine Learning (ML) in different domains, especially in safety-critical areas, increases the need for reliable ML components, i.e., a software component operating based on ML. Since corrective maintenance, i.e. identifying and resolving systems bugs, is a key task in the software development process to deliver reliable software components, it is necessary to investigate the usage of ML components, from the software maintenance perspective. Understanding the bugs’ characteristics and maintenance challenges in ML-based systems can help developers of these systems to identify where to focus maintenance and testing efforts, by giving insights into the most error-prone components, most common bugs, etc. In this paper, we investigate the characteristics of bugs in ML-based software systems and the difference between ML and non-ML bugs from the maintenance viewpoint. We extracted 447,948 GitHub repositories that used one of the three most popular ML frameworks, i.e., TensorFlow, Keras, and PyTorch. After multiple filtering steps, we select the top 300 repositories with the highest number of closed issues. We manually investigate the extracted repositories to exclude non-ML-based systems. Our investigation involved a manual inspection of 386 sampled reported issues in the identified ML-based systems to indicate whether they affect ML components or not. Our analysis shows that nearly half of the real issues reported in ML-based systems are ML bugs, indicating that ML components are more error-prone than non-ML components. Next, we thoroughly examined 109 identified ML bugs to identify their root causes, and symptoms, and calculate their required fixing time. The results also revealed that ML bugs have significantly different characteristics compared to non-ML bugs, in terms of the complexity of bug-fixing (number of commits, changed files, and changed lines of code). Based on our results, fixing ML bugs is more costly and ML components are more error-prone, compared to non-ML bugs and non-ML components respectively. Hence, paying significant attention to the reliability of the ML components is crucial in ML-based systems. These results deepen the understanding of ML bugs and we hope that our findings help shed light on opportunities for designing effective tools for testing and debugging ML-based systems.

引用

共 50 条

[1] Bug characterization in machine learning-based systems
Morovati, Mohammad Mehdi
Nikanjam, Amin
Tambon, Florian
Khomh, Foutse
Jiang, Zhen Ming
EMPIRICAL SOFTWARE ENGINEERING, 2024, 29 (01)
[2] Machine Learning-based Anomaly Detection for Post-silicon Bug Diagnosis
DeOrio, Andrew
Li, Qingkun
Burgess, Matthew
Bertacco, Valeria
DESIGN, AUTOMATION & TEST IN EUROPE, 2013, : 491 - 496
[3] Study of Information Retrieval and Machine Learning-Based Software Bug Localization Models
Tamanna
Sangwan, Om Prakash
ADVANCES IN COMPUTING AND INTELLIGENT SYSTEMS, ICACM 2019, 2020, : 503 - 510
[4] A Machine Learning-Based Framework for Water Quality Index Estimation in the Southern Bug River
Masood, Adil
Niazkar, Majid
Zakwan, Mohammad
Piraei, Reza
WATER, 2023, 15 (20)
[5] Zonal Machine Learning-Based Protection for Distribution Systems
Poudel, Binod P.
Bidram, Ali
Reno, Matthew J.
Summers, Adam
IEEE ACCESS, 2022, 10 : 66634 - 66645
[6] Bugs in machine learning-based systems: a faultload benchmark
Morovati, Mohammad Mehdi
Nikanjam, Amin
Khomh, Foutse
Jiang, Zhen Ming
EMPIRICAL SOFTWARE ENGINEERING, 2023, 28 (03)
[7] Machine Learning-Based Systems for Intrusion Detection in VANETs
Idris, Hala Eldaw
Hosni, Ines
INTELLIGENT SYSTEMS AND APPLICATIONS, VOL 3, INTELLISYS 2024, 2024, 1067 : 603 - 614
[8] Machine Learning-Based Analytical Systems: Food Forensics
Ranbir
Kumar, Manish
Singh, Gagandeep
Singh, Jasvir
Kaur, Navneet
Singh, Narinder
ACS OMEGA, 2022, 7 (51): : 47518 - 47535
[9] Bugs in machine learning-based systems: a faultload benchmark
Mohammad Mehdi Morovati
Amin Nikanjam
Foutse Khomh
Zhen Ming (Jack) Jiang
Empirical Software Engineering, 2023, 28
[10] On Distribution Shift in Learning-based Bug Detectors
He, Jingxuan
Beurer-Kellner, Luca
Vechev, Martin
INTERNATIONAL CONFERENCE ON MACHINE LEARNING, VOL 162, 2022,

← 1 2 3 4 5 →