Please use this identifier to cite or link to this item:
http://localhost:8081/jspui/handle/123456789/21590| Title: | Prediction of Refactoring in Software Systems |
| Authors: | Baghel, Dheeraj Kumar |
| Issue Date: | Jun-2023 |
| Publisher: | IIT Roorkee |
| Abstract: | Refactoring is defined as “the practice of altering the inner workings of a software system while keeping its external functionality intact.” by Martin Fowler [35]. it helps to improve code quality i.e., easier code maintenance, code readability, etc. Earlier it was based on developers’ intuition and expertise, but later researchers explored the static analysis tool, like PMD, ESLint, etc. but they were producing a large number of FP due to which developers were losing faith in themselves. Later Researchers shows that instead of a static approach we should use data-driven ap proach. So, we are extending previous work of Aniche et al. [4] who described “Software Refactoring” as a binary classification problem and used RMv1.0 to detect refactor ing instances. we modified data collection tool and integrated RMv2.0 to fetch the refactoring instances. we collected the large dataset from open-source Java projects hosted on GitHub. For non-refactoring instances, we chose stable commit threshold ’K’ = 50. For feature extraction, we used the various tools such as code metrics calculator (CK) tool for the extraction of source code metrices. Feature are extracted at various levels such as class, method, etc. Finally, the gathered information from RM, CK, and other metric are merged into either refactorings or stable-instances, which are then stored in the database. After the processing of refactoring, the process-metric and ownership-metric undergo updates. Additionally, if the class undergoes renaming or movement, the commit-tracker is also updated. After that we removed erroneous entries from dataset and provide data as input to various supervised machine learning algorithm like Logistic Regression, Random Forest. Lastly we observed Logistic Regression showed an average accuracy of 84% and considered as baseline model whereas Random Forest (RF) consistently outper forms all other binary classifiers with an improved accuracy of 93.6% compared to the previous accuracy of 93.3%. |
| URI: | http://localhost:8081/jspui/handle/123456789/21590 |
| Research Supervisor/ Guide: | Kumar, Sandeep |
| metadata.dc.type: | Dissertations |
| Appears in Collections: | MASTERS' THESES (CSE) |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| 21535009_Dheeraj Kumar Baghel.pdf | 1.9 MB | Adobe PDF | View/Open |
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.
