Klasifikasi Suara Paru – Paru yang ditingkatkan menggunakan SpecAugmentasi dan Mel Spectrogram

Authors

  • Jaenal Arifin Telkom University
  • Dony Hutabarat Universitas Islam Negeri Sultan Maulana Hasanuddin
  • Moh Nur Shodiq

DOI:

https://doi.org/10.61124/sinta.v3i4.398

Keywords:

Klasifikasi Suara Paru-Paru, Mel Spectrogram, SpecAugmentasi, Validasi Silang 10 Fold

Abstract

Penyakit paru-paru merupakan kontributor utama morbiditas dan mortalitas di seluruh dunia. Suara napas pada individu dapat bervariasi dari keadaan normal hingga patologis. Studi ini bertujuan untuk meningkatkan klasifikasi menggunakan SpecAugment dan Mel Spectrogram dengan validasi silang 10 fold. Dataset dikumpulkan dari database Kaggle, dataset International Conference on Biomedical and Health Informatics (ICBHI), dan dataset Rumah Sakit Fortis India. Dataset suara paru-paru terdiri dari 1363 pasien. Pra-pemrosesan meliputi konversi data audio ke sensor float, frekuensi sampling 16 kHz, dan transformasi ke Mel Spectrogram. Augmentasi data dilakukan menggunakan google brain specAugment. Dalam studi ini, kami menggunakan validasi silang 10 fold. Hasil eksperimen memperoleh nilai akurasi terbaik sebesar 95,72%, presisi 96,06%, recall 95,71%, F1-Score 95,74%, dan AUC 99,51%. Studi ini berkontribusi pada penelitian AI di bidang elektromedis dan pemrosesan sinyal. Peningkatan dataset suara paru-paru dengan menambahkan spesifikasi meningkatkan akurasi dan keandalan diagnosis gangguan suara paru-paru.

References

World Health Organization, “The top 10 causes of death,” https://www.who.int/news-room/fact-sheets/detail/the-top-10-causes-of-death.

X. Liao et al., “Automated detection of abnormal respiratory sound from electronic stethoscope and mobile phone using MobileNetV2,” Biocybern. Biomed. Eng., vol. 43, no. 4, pp. 763–775, Oct. 2023, doi: 10.1016/j.bbe.2023.11.001.

Y. Y. Ang, L. R. Aw, V. Koh, and R. X. Tan, “Characterization and cross-comparison of digital stethoscopes for telehealth remote patient auscultation,” Med. Nov. Technol. Devices, vol. 19, Sep. 2023, doi: 10.1016/j.medntd.2023.100256.

N. Tharapecharat et al., “Digital stethoscope with processing and recording based on cloud,” in BMEiCON 2023 - 15th Biomedical Engineering International Conference, Institute of Electrical and Electronics Engineers Inc., 2023. doi: 10.1109/BMEiCON60347.2023.10322053.

X. Li, Y. Shang, J. Wei, and Y. Zhou, “Research on electronic stethoscope system and signal processing algorithm,” in Journal of Physics: Conference Series, Institute of Physics, 2023. doi: 10.1088/1742-6596/2634/1/012037.

Y. S. Rao, B. Mehta, A. Kathiriya, A. Dharvia, D. Pancholi, and H. Patel, “Smart Stethoscope for Remote Healthcare Monitoring Using IoT: A Cost-Effective Solution,” in 2023 3rd Asian Conference on Innovation in Technology, ASIANCON 2023, Institute of Electrical and Electronics Engineers Inc., 2023. doi: 10.1109/ASIANCON58793.2023.10270283.

J. J. Seah, J. Zhao, D. Y. Wang, and H. P. Lee, “Review on the Advancements of Stethoscope Types in Chest Auscultation,” Diagnostics, vol. 13, no. 9, pp. 1–19, May 2023, doi: 10.3390/diagnostics13091545.

J. S. Park, K. Kim, J. H. Kim, Y. J. Choi, K. Kim, and D. I. Suh, “A machine learning approach to the development and prospective evaluation of a pediatric lung sound classification model,” Sci. Rep., vol. 13, no. 1, Dec. 2023, doi: 10.1038/s41598-023-27399-5.

A. Roy and U. Satija, “ILDNet: A Novel Deep Learning Framework for Interstitial Lung Disease Identification Using Respiratory Sounds,” in 2024 International Conference on Signal Processing and Communications, SPCOM 2024, Institute of Electrical and Electronics Engineers Inc., 2024. doi: 10.1109/SPCOM60851.2024.10631581.

T. Wanasinghe, S. Bandara, S. Madusanka, D. Meedeniya, M. Bandara, and I. D. L. T. Diez, “Lung Sound Classification with Multi-Feature Integration Utilizing Lightweight CNN Model,” IEEE Access, vol. 12, pp. 21262–21276, Feb. 2024, doi: 10.1109/ACCESS.2024.3361943.

C. Wu, D. Lei, and Z. Xu, “Respiratory Disease Classification Model Based on Feature Fusion,” in 2023 4th International Conference on Intelligent Computing and Human-Computer Interaction, ICHCI 2023, Institute of Electrical and Electronics Engineers Inc., 2023, pp. 148–155. doi: 10.1109/ICHCI58871.2023.10277774.

M. Chaiani, S. A. Selouani, and M. Boudraa, “Voice Disorder Detection Using Enhanced Auditory Perception-Scaled Spectrograms,” in 2022 45th International Conference on Telecommunications and Signal Processing, TSP 2022, Institute of Electrical and Electronics Engineers Inc., 2022, pp. 49–54. doi: 10.1109/TSP55681.2022.9851379.

R. Islam and M. Tarique, “Spectrogram and Mel-Spectrogram Based Dysphonic Voice Detection Using Convolutional Neural Network,” in International Conference on Electrical, Computer, and Energy Technologies, ICECET 2024, Institute of Electrical and Electronics Engineers Inc., 2024. doi: 10.1109/ICECET61485.2024.10698112.

ICBHI Challenge, “ICBHI 2017 Challenge - Respiratory Sound Database.” Accessed: Aug. 22, 2024. [Online]. Available: https://bhichallenge.med.auth.gr/ICBHI_2017_Challenge

N. Baghel, V. Nangia, and M. K. Dutta, “ALSD-Net: Automatic lung sounds diagnosis network from pulmonary signals,” Neural Comput. Appl., vol. 33, no. 24, pp. 17103–17118, Dec. 2021, doi: 10.1007/s00521-021-06302-1.

https://www.kaggle.com/datasets/vbookshelf/respiratory-sound-database, “Respiratory Sound Database.”

L. D. Yang, R. B. Yue, J. Wang, and M. Liu, “Neural Network Model Based on the Tensor Network for Audio Tagging of Domestic Activities,” Front. Phys., vol. 10, Apr. 2022, doi: 10.3389/fphy.2022.863291.

R. Mahum, A. Irtaza, A. Javed, H. A. Mahmoud, and H. Hassan, “DeepDet: YAMNet with BottleNeck Attention Module (BAM) TTS synthesis detection,” EURASIP J. Audio Speech Music Process., vol. 2024, no. 1, pp. 1–16, Dec. 2024, doi: 10.1186/s13636-024-00335-9.

N. H. Valliappan, S. D. Pande, and S. Reddy Vinta, “Enhancing Gun Detection With Transfer Learning and YAMNet Audio Classification,” IEEE Access, vol. 12, pp. 58940–58949, 2024, doi: 10.1109/ACCESS.2024.3392649.

A. Fava et al., “Pre-processing techniques to enhance the classification of lung sounds based on deep learning,” Biomed. Signal Process. Control, vol. 92, Jun. 2024, doi: 10.1016/j.bspc.2024.106009.

R. Gupta, R. Singh, C. M. Travieso-González, R. Burget, and M. Kishore Dutta, “DeepRespNet: A deep neural network for classification of respiratory sounds,” Biomed. Signal Process. Control, vol. 93, Jul. 2024, doi: 10.1016/j.bspc.2024.106191.

K. N. Lal, “A lung sound recognition model to diagnoses the respiratory diseases by using transfer learning,” Multimed. Tools Appl., vol. 82, no. 23, pp. 36615–36631, Sep. 2023, doi: 10.1007/s11042-023-14727-0.

H. Wang, Y. Zou, and W. Wang, “Specaugment++: A hidden space data augmentation method for acoustic scene classification,” in Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, International Speech Communication Association, 2021, pp. 21–25. doi: 10.21437/Interspeech.2021-140.

G. Shanthakumari and E. Priya, “Interpretation of Lung Sounds Using Spectrogram-Based Statistical Features,” in Futuristic Communication and Network Technologies, A. Sivasubramanian, P. N. Shastry, and P. C. Hong, Eds., Singapore: Springer Nature Singapore, 2022, pp. 815–823.

J. Ramonaite and G. Korvel, “Noisy Phoneme Recognition Using 2D Convolution Neural Network,” in 2023 IEEE 10th Jubilee Workshop on Advances in Information, Electronic and Electrical Engineering, AIEEE 2023 - Proceedings, Institute of Electrical and Electronics Engineers Inc., 2023. doi: 10.1109/AIEEE58915.2023.10134866.

Y. Zhang, A. Herygers, T. Patel, Z. Yue, and O. Scharenborg, “Exploring Data Augmentation in Bias Mitigation Against Non-Native-Accented Speech,” in 2023 IEEE Automatic Speech Recognition and Understanding Workshop, ASRU 2023, Institute of Electrical and Electronics Engineers Inc., 2023. doi: 10.1109/ASRU57964.2023.10389756.

J. Han, M. Matuszewski, O. Sikorski, H. Sung, and H. Cho, “Randmasking Augment: A Simple and Randomized Data Augmentation For Acoustic Scene Classification,” in ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings, Institute of Electrical and Electronics Engineers Inc., 2023. doi: 10.1109/ICASSP49357.2023.10095001.

D. S. Park et al., “Specaugment: A simple data augmentation method for automatic speech recognition,” in Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, Graz, Austria: INTERSPEECH, Sep. 2019, pp. 2613–2617. doi: 10.21437/Interspeech.2019-2680.

O. O. Abayomi-Alli, R. Damaševičius, A. Qazi, M. Adedoyin-Olowe, and S. Misra, “Data Augmentation and Deep Learning Methods in Sound Classification: A Systematic Review,” Electronics (Basel)., vol. 11, no. 22, Nov. 2022, doi: 10.3390/electronics11223795.

A. Y. Chang et al., “GAP-AUG: GAMMA PATCH-WISE CORRECTION AUGMENTATION METHOD FOR RESPIRATORY SOUND CLASSIFICATION,” in ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings, Seoul: Institute of Electrical and Electronics Engineers Inc., Apr. 2024, pp. 551–555. doi: 10.1109/ICASSP48485.2024.10447967.

T. T. Wong and P. Y. Yeh, “Reliable Accuracy Estimates from k-Fold Cross Validation,” IEEE Trans. Knowl. Data Eng., vol. 32, no. 8, pp. 1586–1594, Aug. 2020, doi: 10.1109/TKDE.2019.2912815.

S. M. Malakouti, M. B. Menhaj, and A. A. Suratgar, “The usage of 10-fold cross-validation and grid search to enhance ML methods performance in solar farm power generation prediction,” Clean. Eng. Technol., vol. 15, Aug. 2023, doi: 10.1016/j.clet.2023.100664.

E. Bergman, L. Purucker, and F. Hutter, “Don’t Waste Your Time: Early Stopping Cross-Validation,” in AutoML 2024 (International Conference on Automated Machine Learning), Paris, France: Proceedings of Machine Learning Research (PMLR), Volume 256, Sep. 2024, pp. 1–32. [Online]. Available: https://github.com/automl/DontWasteYourTime-early-stopping

K. S. Gill and R. Gupta, “Chronic Kidney Disease Detection Using GridSearchCV Cross Validation Method,” in 2023 International Conference on Recent Advances in Electrical, Electronics and Digital Healthcare Technologies, REEDCON 2023, Institute of Electrical and Electronics Engineers Inc., 2023, pp. 318–322. doi: 10.1109/REEDCON57544.2023.10151392.

Published

08/06/2026

How to Cite

Arifin, J., Dony Hutabarat, & Moh Nur Shodiq. (2026). Klasifikasi Suara Paru – Paru yang ditingkatkan menggunakan SpecAugmentasi dan Mel Spectrogram. Jurnal SINTA: Sistem Informasi Dan Teknologi Komputasi, 3(4). https://doi.org/10.61124/sinta.v3i4.398