Clustering Dycom: An Online Cross-Company Software Effort Estimation Study

L. MINKU; S. HOU

doi:10.1145/3127005.3127007

Clustering Dycom: An Online Cross-Company Software Effort Estimation Study

L. MINKU, S. HOU

Computer Science

Research output: Chapter in Book/Report/Conference proceeding › Conference contribution

8 Citations (Scopus)

111 Downloads (Pure)

Abstract

Background: Software Effort Estimation (SEE) can be formulated as an online learning problem, where new projects are completed over time and may become available for training. In this scenario, a Cross-Company (CC) SEE approach called Dycom can drastically reduce the number of Within-Company (WC) projects needed for training, saving the high cost of collecting such training projects. However, Dycom relies on splitting CC projects into different subsets in order to create its CC models. Such splitting can have a significant impact on Dycom's predictive performance. Aims: This paper investigates whether clustering methods can be used to help finding good CC splits for Dycom. Method: Dycom is extended to use clustering methods for creating the CC subsets. Three different clustering methods are investigated, namely Hierarchical Clustering, K-Means, and Expectation-Maximisation. Clustering Dycom is compared against the original Dycom with CC subsets of different sizes, based on four SEE databases. A baseline WC model is also included in the analysis. Results: Clustering Dycom with K-Means can potentially help to split the CC projects, managing to achieve similar or better predictive performance than Dycom. However, K-Means still requires the number of CC subsets to be pre-defined, and a poor choice can negatively affect predictive performance. EM enables Dycom to automatically set the number of CC subsets while still maintaining or improving predictive performance with respect to the baseline WC model. Clustering Dycom with Hierarchical Clustering did not offer significant advantage in terms of predictive performance. Conclusion: Clustering methods can be an effective way to automatically generate Dycom's CC subsets.

Original language	English
Title of host publication	Proceedings of the 13th International Conference on Predictive Models and Data Analytics in Software Engineering (PROMISE)
Place of Publication	Toronto, Canada
Publisher	ACM/IEEE
Pages	12-21
Number of pages	10
ISBN (Print)	9781450353052
DOIs	https://doi.org/10.1145/3127005.3127007
Publication status	E-pub ahead of print - 8 Nov 2017

Access to Document

10.1145/3127005.3127007Licence: None: All rights reserved

Minku_&_Hou_Clustering_Dycom_PROMISE_2017
Checked for eligibility 09/11/2018 doi>10.1145/3127005.3127007
Accepted author manuscript, 643 KBLicence: None: All rights reserved

https://dl.acm.org/citation.cfm?id=3127007Licence: None: All rights reserved

Cite this

@inproceedings{68ddeec93cee4489a3672a15dc6f55e1,

title = "Clustering Dycom: An Online Cross-Company Software Effort Estimation Study",

abstract = "Background: Software Effort Estimation (SEE) can be formulated as an online learning problem, where new projects are completed over time and may become available for training. In this scenario, a Cross-Company (CC) SEE approach called Dycom can drastically reduce the number of Within-Company (WC) projects needed for training, saving the high cost of collecting such training projects. However, Dycom relies on splitting CC projects into different subsets in order to create its CC models. Such splitting can have a significant impact on Dycom's predictive performance. Aims: This paper investigates whether clustering methods can be used to help finding good CC splits for Dycom. Method: Dycom is extended to use clustering methods for creating the CC subsets. Three different clustering methods are investigated, namely Hierarchical Clustering, K-Means, and Expectation-Maximisation. Clustering Dycom is compared against the original Dycom with CC subsets of different sizes, based on four SEE databases. A baseline WC model is also included in the analysis. Results: Clustering Dycom with K-Means can potentially help to split the CC projects, managing to achieve similar or better predictive performance than Dycom. However, K-Means still requires the number of CC subsets to be pre-defined, and a poor choice can negatively affect predictive performance. EM enables Dycom to automatically set the number of CC subsets while still maintaining or improving predictive performance with respect to the baseline WC model. Clustering Dycom with Hierarchical Clustering did not offer significant advantage in terms of predictive performance. Conclusion: Clustering methods can be an effective way to automatically generate Dycom's CC subsets.",

author = "L. MINKU and S. HOU",

year = "2017",

month = nov,

day = "8",

doi = "10.1145/3127005.3127007",

language = "English",

isbn = "9781450353052",

pages = "12--21",

booktitle = "Proceedings of the 13th International Conference on Predictive Models and Data Analytics in Software Engineering (PROMISE)",

publisher = "ACM/IEEE",

}

TY - GEN

T1 - Clustering Dycom: An Online Cross-Company Software Effort Estimation Study

AU - MINKU, L.

AU - HOU, S.

PY - 2017/11/8

Y1 - 2017/11/8

N2 - Background: Software Effort Estimation (SEE) can be formulated as an online learning problem, where new projects are completed over time and may become available for training. In this scenario, a Cross-Company (CC) SEE approach called Dycom can drastically reduce the number of Within-Company (WC) projects needed for training, saving the high cost of collecting such training projects. However, Dycom relies on splitting CC projects into different subsets in order to create its CC models. Such splitting can have a significant impact on Dycom's predictive performance. Aims: This paper investigates whether clustering methods can be used to help finding good CC splits for Dycom. Method: Dycom is extended to use clustering methods for creating the CC subsets. Three different clustering methods are investigated, namely Hierarchical Clustering, K-Means, and Expectation-Maximisation. Clustering Dycom is compared against the original Dycom with CC subsets of different sizes, based on four SEE databases. A baseline WC model is also included in the analysis. Results: Clustering Dycom with K-Means can potentially help to split the CC projects, managing to achieve similar or better predictive performance than Dycom. However, K-Means still requires the number of CC subsets to be pre-defined, and a poor choice can negatively affect predictive performance. EM enables Dycom to automatically set the number of CC subsets while still maintaining or improving predictive performance with respect to the baseline WC model. Clustering Dycom with Hierarchical Clustering did not offer significant advantage in terms of predictive performance. Conclusion: Clustering methods can be an effective way to automatically generate Dycom's CC subsets.

AB - Background: Software Effort Estimation (SEE) can be formulated as an online learning problem, where new projects are completed over time and may become available for training. In this scenario, a Cross-Company (CC) SEE approach called Dycom can drastically reduce the number of Within-Company (WC) projects needed for training, saving the high cost of collecting such training projects. However, Dycom relies on splitting CC projects into different subsets in order to create its CC models. Such splitting can have a significant impact on Dycom's predictive performance. Aims: This paper investigates whether clustering methods can be used to help finding good CC splits for Dycom. Method: Dycom is extended to use clustering methods for creating the CC subsets. Three different clustering methods are investigated, namely Hierarchical Clustering, K-Means, and Expectation-Maximisation. Clustering Dycom is compared against the original Dycom with CC subsets of different sizes, based on four SEE databases. A baseline WC model is also included in the analysis. Results: Clustering Dycom with K-Means can potentially help to split the CC projects, managing to achieve similar or better predictive performance than Dycom. However, K-Means still requires the number of CC subsets to be pre-defined, and a poor choice can negatively affect predictive performance. EM enables Dycom to automatically set the number of CC subsets while still maintaining or improving predictive performance with respect to the baseline WC model. Clustering Dycom with Hierarchical Clustering did not offer significant advantage in terms of predictive performance. Conclusion: Clustering methods can be an effective way to automatically generate Dycom's CC subsets.

U2 - 10.1145/3127005.3127007

DO - 10.1145/3127005.3127007

M3 - Conference contribution

SN - 9781450353052

SP - 12

EP - 21

BT - Proceedings of the 13th International Conference on Predictive Models and Data Analytics in Software Engineering (PROMISE)

PB - ACM/IEEE

CY - Toronto, Canada

ER -

Clustering Dycom: An Online Cross-Company Software Effort Estimation Study

Abstract

Access to Document

Fingerprint

Cite this