Optimizing Iris Plant Classification with Ensemble Models and XAI

A Comprehensive Analysis of Model Performance

Authors

  • Ahmad Subadri Sunan Kalijaga State Islamic University Yogyakarta image/svg+xml
  • Ishmah Afiyah Sunan Kalijaga State Islamic University Yogyakarta image/svg+xml
  • Fiki Sanora Sunan Kalijaga State Islamic University Yogyakarta image/svg+xml
  • Arya Indrawan Sunan Kalijaga State Islamic University Yogyakarta image/svg+xml
  • Maria Ulfah Siregar Sunan Kalijaga State Islamic University Yogyakarta image/svg+xml

DOI:

https://doi.org/10.14421/jiska.5967

Keywords:

Iris Plant, Ensemble Models, XAI, Model Performance, Machine Learning

Abstract

This study aims to improve the performance of Iris plant classification by integrating ensemble learning techniques with Explainable Artificial Intelligence (XAI) to achieve high accuracy while enhancing model interpretability. Random Forest, XGBoost, and AdaBoost algorithms are combined within a Voting Ensemble framework and evaluated using the Iris Plants dataset, which comprises 150 data samples distributed equally across three Iris species (50 samples per class: Iris setosa, Iris versicolor, and Iris virginica). The dataset exhibits a perfectly balanced class distribution, ensuring that no class imbalance correction was required. The Voting Ensemble model was evaluated using a hold-out test set (80:20 split) and further validated through 5-Fold Stratified Cross-Validation, yielding a mean cross-validation accuracy of 95.83% (±2.64%) and a test set accuracy of 93.33%. To enhance model transparency, the SHAP (SHapley Additive Explanations) method is applied to explain the contribution of each feature to the prediction outcomes. The Voting Ensemble model achieved an ROC AUC score of 0.9900 (macro-average), with Precision, Recall, and F1-Score each reaching 0.9333 (macro-average). Feature importance analysis reveals that petal length and petal width are the primary factors in the Iris species classification process. The strong correlation (r = 0.9991) between feature importance scores in the Random Forest model and SHAP values confirms the consistency and reliability of the model’s interpretability. These findings demonstrate that integrating ensemble learning with XAI not only improves predictive performance but also strengthens transparency and trust in machine learning models, particularly for plant classification tasks.

Author Biography

  • Ahmad Subadri, Sunan Kalijaga State Islamic University Yogyakarta

    Ahmad Subadri is affiliated with Universitas Islam Negeri Sunan Kalijaga Yogyakarta. His research interests include information systems, web-based application development, and data analysis.

References

Arjuna Priandika1, A. R. I. (2025). Application of Ensemble Learning Technique for Classification of Anemia Types Penerapan Teknik Ensemble Learning untuk. 5(July), 972–980.

Bobek, S., Kuk, M., Szelazek, M., & Nalepa, G. J. (2022). Enhancing Cluster Analysis With Explainable AI and Multidimensional Cluster Prototypes. IEEE Access, 10(September), 101556–101574. https://doi.org/10.1109/ACCESS.2022.3208957

Choi, Y. R., & Lim, D. J. (2021). DDES: A Distribution-Based Dynamic Ensemble Selection Framework. IEEE Access, 9, 40743–40754. https://doi.org/10.1109/ACCESS.2021.3063254

Doyen, S., Taylor, H., Nicholas, P., Crawford, L., Young, I., & Sughrue, M. E. (2021). Hollow-tree super: A directional and scalable approach for feature importance in boosted tree models. PLoS ONE, 16(10 October), 1–16. https://doi.org/10.1371/journal.pone.0258658

Fisher, R. (2025). Iris plants dataset. Scikit Learn. https://scikit-learn.org/stable/datasets/toy_dataset.html

Ge, H., Ma, F., Li, Z., Tan, Z., & Du, C. (2021). Improved accuracy of phenological detection in rice breeding by using ensemble models of machine learning based on uav‐rgb imagery. Remote Sensing, 13(14). https://doi.org/10.3390/rs13142678

He, Z., Yang, Y., Fang, R., Zhou, S., Zhao, W., Bai, Y., Li, J., & Wang, B. (2023). Integration of shapley additive explanations with a random forest model for quantitative precipitation estimation of mesoscale convective systems. Frontiers in Environmental Science, 10(January), 1–15. https://doi.org/10.3389/fenvs.2022.1057081

Hernandez, M., Ramon-Julvez, U., & Ferraz, F. (2022). Explainable AI toward understanding the performance of the top three TADPOLE Challenge methods in the forecast of Alzheimer’s disease diagnosis. In PLoS ONE (Vol. 17, Issue 5, May). https://doi.org/10.1371/journal.pone.0264695

Jafarzadeh, H., Mahdianpari, M., Gill, E., Mohammadimanesh, F., & Homayouni, S. (2021). Bagging and boosting ensemble classifiers for classification of multispectral, hyperspectral, and polSAR data: A comparative evaluation. Remote Sensing, 13(21). https://doi.org/10.3390/rs13214405

Jiang, P., Suzuki, H., & Obi, T. (2023). XAI-based cross-ensemble feature ranking methodology for machine learning models. International Journal of Information Technology (Singapore), 15(4), 1759–1768. https://doi.org/10.1007/s41870-023-01270-2

Kaneko, H. (2023). Interpretation of Machine Learning Models for Data Sets with Many Features Using Feature Importance. ACS Omega, 8(25), 23218–23225. https://doi.org/10.1021/acsomega.3c03722

Lundberg, S. M., & Lee, S. I. (2017). A unified approach to interpreting model predictions. Advances in Neural Information Processing Systems, 2017-Decem(Section 2), 4766–4775.

Matloob, F., Ghazal, T. M., Taleb, N., Abbas, S., Soomro, T. R., & Member, S. (2021). Software Defect Prediction Using Ensemble Learning : A Systematic Literature Review. IEEE Access, 9, 98754–98771. https://doi.org/10.1109/ACCESS.2021.3095559

Mostofi, F., To, V., Ayözen, Y. E., & Tokdemir, O. B. (2022). Predicting the Impact of Construction Rework Cost Using an Ensemble Classifier. Sustainability MDPI, 14(22). https://doi.org/https://doi.org/10.3390/su142214800

Orlenko, A., & Moore, J. H. (2021). A comparison of methods for interpreting random forest models of genetic association in the presence of non-additive interactions. BioData Mining, 14(1), 1–22. https://doi.org/10.1186/s13040-021-00243-0

Rezk, N. G., Alshathri, S., Sayed, A., El-Din Hemdan, E., & El-Behery, H. (2024). XAI-Augmented Voting Ensemble Models for Heart Disease Prediction: A SHAP and LIME-Based Approach. Bioengineering, 11(10). https://doi.org/10.3390/bioengineering11101016

Rosyid, I. F., & Pramaditya, H. (2025). Visual Interpretation of Machine Learning Models (Random Forest) for Lung Cancer Risk Classification Using Explainable Artificial Intelligence (SHAP & LIME). In Jurnal Teknik Informatika (Jutif) (Vol. 6, Issue 4, pp. 2187–2206). https://doi.org/10.52436/1.jutif.2025.6.4.4925

Ryo, M. (2022). Explainable artificial intelligence and interpretable machine learning for agricultural data analysis. Artificial Intelligence in Agriculture, 6, 257–265. https://doi.org/10.1016/j.aiia.2022.11.003

Tasci, E., Zhuge, Y., Kaur, H., Camphausen, K., & Krauze, A. V. (2022). Hierarchical Voting-Based Feature Selection and Ensemble Learning Model Scheme for Glioma Grading with Clinical and Molecular Characteristics. International Journal of Molecular Sciences, 23(22). https://doi.org/https://doi.org/10.3390/ijms232214155

Theissler, A., Thomas, M., Burch, M., & Gerschner, F. (2022). Knowledge-Based Systems ConfusionVis : Comparative evaluation and selection of multi-class classifiers based on confusion matrices. Knowledge-Based Systems, 247, 108651. https://doi.org/10.1016/j.knosys.2022.108651

Wang, H., Liang, Q., Hancock, J. T., & Khoshgoftaar, T. M. (2024). Feature selection strategies: a comparative analysis of SHAP-value and importance-based methods. Journal of Big Data, 11(1). https://doi.org/10.1186/s40537-024-00905-w

Wu, Y., He, J., Ji, Y., Huang, G., Yao, H., Zhang, P., Xu, W., Guo, M., & Li, Y. (2019). Enhanced Classification Models for Iris Dataset. Procedia Computer Science, 162, 946–954. https://doi.org/10.1016/j.procs.2019.12.072

Wu, Y., Liu, L., Xie, Z., Chow, K. H., & Wei, W. (2021). Boosting ensemble accuracy by revisiting ensemble diversity metrics. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 16464–16472. https://doi.org/10.1109/CVPR46437.2021.01620

Yates, L. A., Aandahl, Z., Richards, S. A., & Brook, B. W. (2023). Cross-validation for model selection: A review with examples from ecology. Ecological Monographs, 93(1), 1–24. https://doi.org/10.1002/ecm.1557

Zhou, Z., & Hooker, G. (2021). Unbiased measurement of feature importance in tree-based methods. ACM Transactions on Knowledge Discovery from Data, 15(2). https://doi.org/10.1145/3429445

Downloads

Published

2026-05-25

Issue

Section

Articles

How to Cite

Optimizing Iris Plant Classification with Ensemble Models and XAI: A Comprehensive Analysis of Model Performance. (2026). JISKA (Jurnal Informatika Sunan Kalijaga), 11(2), 195-212. https://doi.org/10.14421/jiska.5967

Similar Articles

1-10 of 111

You may also start an advanced similarity search for this article.

Most read articles by the same author(s)