dc.creatorAtlántida Irene Sánchez Vivar
dc.creatorEduardo Francisco Morales Manzanares
dc.creatorJesús Antonio González Bernal
dc.date2013
dc.date.accessioned2023-07-25T16:25:31Z
dc.date.available2023-07-25T16:25:31Z
dc.identifierhttp://inaoe.repositorioinstitucional.mx/jspui/handle/1009/2395
dc.identifier.urihttps://repositorioslatinoamericanos.uchile.cl/handle/2250/7807571
dc.descriptionImbalanced data sets, in the class distribution, is common to many real world applications. As many classifiers tend to degrade their performance over the minority class, several approaches have been proposed to deal with this problem. In this paper, we propose two new cluster-based oversampling methods, SOI-C and SOI-CJ. The proposed methods create clusters from the minority class instances and generate synthetic instances inside those clusters. In contrast with other oversampling methods, the proposed approaches avoid creating new instances in majority class regions. They are more robust to noisy examples (the number of new instances generated per cluster is proportional to the cluster's size). The clusters are automatically generated. Our new methods do not need tuning parameters, and they can deal both with numerical and nominal attributes. The two methods were tested with twenty artificial datasets and twenty three datasets from the UCI Machine Learning repository. For our experiments, we used six classifiers and results were evaluated with TPR, precision, F-measure, and AUC measures, which are more suitable for class imbalanced datasets. We performed ANOVA and paired t-tests to show that the proposed methods are competitive and in many cases significantly better than the rest of the oversampling methods used during the comparison.
dc.formatapplication/pdf
dc.languageeng
dc.publisherWorld Scientific Publishing Company
dc.relationcitation:Sánchez, A., et al., (2013). Synthetic Oversampling of Instances Using Clustering, International Journal on Artificial Intelligence Tools, Vol. 22 (2): 1-22
dc.rightsinfo:eu-repo/semantics/openAccess
dc.rightshttp://creativecommons.org/licenses/by-nc-nd/4.0
dc.subjectinfo:eu-repo/classification/Imbalanced datasets/Imbalanced datasets
dc.subjectinfo:eu-repo/classification/Oversampling/Oversampling
dc.subjectinfo:eu-repo/classification/Cluster-based oversampling/Cluster-based oversampling
dc.subjectinfo:eu-repo/classification/Jittering/Jittering
dc.subjectinfo:eu-repo/classification/cti/1
dc.subjectinfo:eu-repo/classification/cti/12
dc.subjectinfo:eu-repo/classification/cti/1203
dc.subjectinfo:eu-repo/classification/cti/1203
dc.titleSynthetic Oversampling of Instances Using Clustering
dc.typeinfo:eu-repo/semantics/article
dc.typeinfo:eu-repo/semantics/acceptedVersion
dc.audiencestudents
dc.audienceresearchers
dc.audiencegeneralPublic


Este ítem pertenece a la siguiente institución