INVESTIGATION OF SPEECH DISFLUENCIES CLASSIFICATION ON DIFFERENT THRESHOLD SELECTION TECHNIQUES USING ENERGY FEATURE EXTRACTION

Authors

  • R. Hamzah Faculty of Computer and Mathematical Sciences, UiTM Shah Alam, Selangor, Malaysia
  • N.Jamil Faculty of Computer and Mathematical Sciences, UiTM Shah Alam, Selangor, Malaysia

DOI:

https://doi.org/10.24191/mjoc.v4i1.4979

Keywords:

Filled Pause and Elongation, Naïve Bayes, Energy Feature Extraction, Automatic Speech Recognition

Abstract

Filled pause and Elongation are the two types of speech disfluencies that need more suitable acoustical features to be classified correctly since they are always being misclassified. This work concentrates on developing an accurate and robust energy feature extraction for modelling filled pause and elongation by investigating different energy features using local maxima points of the speech energy. Method: In this paper, we extracted peak values from each frame of a voiced signal by implementing different thresholding techniques to classify filled pause and elongation. These energy features are evaluated by using statistical naïve Bayes classifier to see the contribution on the classification processes. Various samples of sustained syllables and filled pauses of spontaneous speech were extracted from Malaysian Parliamentary Debate Database of the year 2008. A naïve Bayes was used as a classifier. We performed F-measure evaluation to investigate the significant differences in mean of filled pause and elongation samples. Results: Results revealed that our proposed LM-E has increase the classification with up to 71% and 75% F-measure for elongation and filled pause. Conclusion: The best achieved accuracies in both filled pause and elongation classification were varied depending on the types of thresholding techniques applied during the local maxima of speech energy extraction. The most contributed thresholding technique is our proposed technique which is by using the adaptive height as the threshold that extracts the local maxima of the speech energy (LM-E).

References

Abbas, E. I., & Refeis, A. A. (2013). Influence of Noisy Environment on the Speech Recognition Rate Based on the Altera FPGA. Engineering and Technology Journal, 31(13 Part (A) Engineering), 2513-2530.

Audhkhasi, K., Kandhway, K., Deshmukh, O. D., & Verma, A. (2009, April). Formant-based technique for automatic filled-pause detection in spontaneous spoken English. In 2009 IEEE International Conference on Acoustics, Speech and Signal Processing (pp. 4857-4860). IEEE.

Bertot, E. M., Beaujean, P. P., & Vendittis, D. (2014, July). Refining envelope analysis methods using wavelet de-noising to identify bearing faults. In European Conference of the Prognostics and Health Management Society.

Bouckaert, R. R. (2004, December). Naive bayes classifiers that perform well with continuous variables. In Australasian joint conference on artificial intelligence (pp. 1089-1094). Springer, Berlin, Heidelberg.

Deng, L., & O'Shaughnessy, D. (2018). Speech processing: a dynamic and optimization-oriented approach. CRC Press.

Doellinger, M., Burger, M., Hoppe, U., Bosco, E., & Eysholdt, U. (2011). Effects of consonant-vowel transitions in speech stimuli on cortical auditory evoked potentials in adults. The open neurology journal, 5, 37.

Dougherty, J., Kohavi, R., & Sahami, M. (1995). Supervised and unsupervised discretization of continuous features. In Machine Learning Proceedings 1995 (pp. 194-202). Morgan Kaufmann.

Espy-Wilson, C. (1986, April). A phonetically based semivowel recognition system. In ICASSP'86. IEEE International Conference on Acoustics, Speech, and Signal Processing(Vol. 11, pp. 2775-2778). IEEE.

Elkan, C. (2012). Evaluating classifiers. San Diego: University of California.

Gabrea, M., & O'Shaughnessy, D. (2000). Detection of filled pauses in spontaneous conversational speech. In Sixth International Conference on Spoken Language Processing.

Ganapathy, S. (2012). Signal analysis using autoregressive models of amplitude modulation (Doctoral dissertation, Johns Hopkins University).

Garg, G., & Ward, N. (2006). Detecting filled pauses in tutorial dialogs. Report of University of Texas at El Paso, El Paso.

Goto, M., Itou, K., & Hayamizu, S. (1999). A real-time filled pause detection system for spontaneous speech recognition. In Sixth European Conference on Speech Communication and Technology.

Izzad, M., Jamil, N., & Bakar, Z. A. (2013, January). Speech/non-speech detection in Malay language spontaneous speech. In 2013 International Conference on Computing, Management and Telecommunications (ComManTel) (pp. 219-224). IEEE.

Jalil, M., Butt, F. A., & Malik, A. (2013, May). Short-time energy, magnitude, zero crossing rate and autocorrelation measurement for discriminating voiced and unvoiced segments of speech signals. In 2013 The International Conference on Technological Advances in Electrical, Electronics and Computer Engineering (TAEECE) (pp. 208-212). IEEE.

Karpiński, M. (2013). Acoustic Features of Filled Pauses in Polish Task-Oriented Dialogues. Archives of Acoustics, 38(1), 63-73.

Kaushik, M., Trinkle, M., & Hashemi-Sakhtsari, A. (2010). Automatic detection and removal of disfluencies from spontaneous speech. In Proceedings of the Australasian International Conference on Speech Science and Technology (SST).

Kitayama, K., Goto, M., Itou, K., & Kobayashi, T. (2003). Speech starter: Noise-robust endpoint detection by using filled pauses. In Eighth European Conference on Speech Communication and Technology.

Li, Y. X., He, Q. H., & Li, T. (2008, May). A novel detection method of filled pause in mandarin spontaneous speech. In Seventh IEEE/ACIS International Conference on Computer and Information Science (icis 2008) (pp. 217-222). IEEE.

Li, Y. X., He, Q. H., Li, W., & Wang, Z. F. (2010, November). Two-level approach for detecting non-lexical audio events in spontaneous speech. In 2010 International Conference on Audio, Language and Image Processing (pp. 771-777). IEEE.

Meseguer, N. A. (2009). Speech analysis for automatic speech recognition. Norwegian University of Science and Technology, Department of Electronics and Telecommunications, 14-19.

Mohd Yusof, S. A., & Yaacob, S. (2008). Classification of Malaysian vowels using formant based features. Journal of ICT, 7, 27-40.

Murakami, Y., & Mizuguchi, K. (2010). Applying the Naïve Bayes classifier with kernel density estimation to the prediction of protein–protein interaction sites. Bioinformatics, 26(15), 1841-1848.

Ogata, J., Goto, M., & Itou, K. (2009, April). The use of acoustically detected filled and silent pauses in spontaneous speech recognition. In 2009 IEEE International Conference on Acoustics, Speech and Signal Processing (pp. 4305-4308). IEEE.

Qin, H., Ma X., Herawan, T., and Zain, J. M. (2012). “DFIS: A Novel Data Filling Approach for Incomplete Soft Set”, Journal of Applied Mathematics and Computer Science, 22 (4), 817-828.

Rosenberg, A., & Hirschberg, J. (2006). On the correlation between energy and pitch accent in read english speech. In Ninth International Conference on Spoken Language Processing.

Schwartzman, A., Gavrilov, Y., & Adler, R. J. (2011). Multiple testing of local maxima for detection of peaks in 1D. Annals of statistics, 39(6), 3290.

Singh, B., Rani, V., & Mahajan, N. (2012). Preprocessing in ASR for computer machine interaction with humans: A review. International Journal of Advanced Research in Computer Science and Software Engineering, 2(3), 396-399.

Stouten, F. (2008). Feature extraction and event detection for automatic speech recognition (Doctoral dissertation, Ghent University).

Stouten, F., & Martens, J. P. (2003). A feature-based filled pause detection system for Dutch. In Automatic Speech Recognition and Understanding, 2003. ASRU'03. 2003 IEEE Workshop on (pp. 309-314). IEEE.

Stouten, F., Duchateau, J., Martens, J. P., & Wambacq, P. (2006). Coping with disfluencies in spontaneous speech recognition: Acoustic detection and linguistic context manipulation. Speech Communication, 48(11), 1590-1606.

Stouten, F., Duchateau, J., Martens, J. P., & Wambacq, P. (2006). Coping with disfluencies in spontaneous speech recognition: Acoustic detection and linguistic context manipulation. Speech Communication, 48(11), 1590-1606.

Veiga, A., Candeias, S., Lopes, C., & Perdigão, F. (2011, August). Characterization of Hesitations Using Acoustic Models. In ICPhS (pp. 2054-2057).

Verkhodanova, V., & Shapranov, V. (2014, October). Filled Pauses and Lengthenings Detection Based on the Acoustic Features for the Spontaneous Russian Speech. In International Conference on Speech and Computer (pp. 227-234). Springer, Cham.

Zapata, J., & Kirkedal, A. S. (2015). Assessing the Performance of Automatic Speech Recognition Systems When Used by Native and Non-Native Speakers of Three Major Languages in Dictation Workflows. In Proceedings of the 20th Nordic Conference of Computational Linguistics (NODALIDA 2015) (pp. 201-210).

Žgank, A., Rotovnik, T., & Sepesy Maučec, M. (2008). Slovenian spontaneous speech recognition and acoustic modeling of filled pauses and onomatopoeas. WSEAS Transaction on Signal Processing, 4(7), 388-397.

Published

2019-06-01

How to Cite

R. Hamzah, & N.Jamil. (2019). INVESTIGATION OF SPEECH DISFLUENCIES CLASSIFICATION ON DIFFERENT THRESHOLD SELECTION TECHNIQUES USING ENERGY FEATURE EXTRACTION. Malaysian Journal of Computing, 4(1). https://doi.org/10.24191/mjoc.v4i1.4979