Predicting the Intelligibility of Vocoded Speech

被引:46
|
作者
Chen, Fei [1 ]
Loizou, Philipos C. [1 ]
机构
[1] Univ Texas Dallas, Dept Elect Engn, Richardson, TX 75080 USA
来源
EAR AND HEARING | 2011年 / 32卷 / 03期
关键词
NORMAL-HEARING LISTENERS; ELECTRIC HEARING; TEMPORAL CUES; PHONEME RECOGNITION; SIGNAL PROCESSORS; ACOUSTIC HEARING; NOISE; ENVELOPE; CHANNELS; INDEX;
D O I
10.1097/AUD.0b013e3181ff3515
中图分类号
R36 [病理学]; R76 [耳鼻咽喉科学];
学科分类号
100104 ; 100213 ;
摘要
Objectives: The purpose of this study is to evaluate the performance of a number of speech intelligibility indices in terms of predicting the intelligibility of vocoded speech. Design: Noise-corrupted sentences were vocoded in a total of 80 conditions, involving three different signal-to-noise ratio levels (-5, 0, and 5 dB) and two types of maskers (steady state noise and two-talker). Tone-vocoder simulations and combined electric-acoustic stimulation (EAS) simulations were used. The vocoded sentences were presented to normal-hearing listeners for identification, and the resulting intelligibility scores were used to assess the correlation of various speech intelligibility measures. These included measures designed to assess speech intelligibility, including the speech transmission index (STI) and articulation index based measures, as well as distortions in hearing aids (e. g., coherence-based measures). These measures employed primarily either the temporal-envelope or the spectral-envelope information in the prediction model. The underlying hypothesis in the present study is that measures that assess temporal-envelope distortions, such as those based on the STI, should correlate highly with the intelligibility of vocoded speech. This is based on the fact that vocoder simulations preserve primarily envelope information, similar to the processing implemented in current cochlear implant speech processors. Similarly, it is hypothesized that measures such as the coherence-based index that assess the distortions present in the spectral envelope could also be used to model the intelligibility of vocoded speech. Results: Of all the intelligibility measures considered, the coherence-based and the STI-based measures performed the best. High correlations (r = 0.9 to 0.96) were maintained with the coherence-based measures in all noisy conditions. The highest correlation obtained with the STI-based measure was 0.92, and that was obtained when high modulation rates (100 Hz) were used. The performance of these measures remained high in both steady-noise and fluctuating masker conditions. The correlations with conditions involving tone-vocoded speech were found to be a bit higher than the correlations with conditions involving EAS-vocoded speech. Conclusions: The present study demonstrated that some of the speech intelligibility indices that have been found previously to correlate highly with wideband speech can also be used to predict the intelligibility of vocoded speech. Both the coherence-based and STI-based measures have been found to be good measures for modeling the intelligibility of vocoded speech. The highest correlation (r = 0.96) was obtained with a derived coherence measure that placed more emphasis on information contained in vowel/consonant spectral transitions and less emphasis on information contained in steady sonorant segments. High (100 Hz) modulation rates were found to be necessary in the implementation of the STI-based measures for better modeling of the intelligibility of vocoded speech. We believe that the difference in modulation rates needed for modeling the intelligibility of wideband versus vocoded speech can be attributed to the increased importance of higher modulation rates in situations where the amount of spectral information available to the listeners is limited (eight channels in our study). Unlike the traditional STI method that has been found to perform poorly in terms of predicting the intelligibility of processed speech wherein nonlinear operations are involved, the STI-based measure used in the present study has been found to perform quite well. In summary, the present study took the first step in modeling the intelligibility of vocoded speech. Access to such intelligibility measures is of high significance as they can be used to guide the development of new speech coding algorithms for cochlear implants.
引用
收藏
页码:331 / 338
页数:8
相关论文
共 50 条
  • [1] Speech intelligibility and talker gender classification with noise-vocoded and tone-vocoded speech
    Villard, Sarah
    Kidd, Gerald, Jr.
    JASA EXPRESS LETTERS, 2021, 1 (09):
  • [2] Predicting the Intelligibility of Cochlear-implant Vocoded Speech from Objective Quality Measure
    Chen, Fei
    JOURNAL OF MEDICAL AND BIOLOGICAL ENGINEERING, 2012, 32 (03) : 189 - 193
  • [3] Predicting the intelligibility of vocoded and wideband Mandarin Chinese
    Chen, Fei
    Loizou, Philipos C.
    JOURNAL OF THE ACOUSTICAL SOCIETY OF AMERICA, 2011, 129 (05): : 3281 - 3290
  • [4] Effects of factor elimination on intelligibility of noise-vocoded Japanese speech
    Kishida, Takuya
    Nakajima, Yoshitaka
    Ueda, Kazuo
    Remijin, Gerard B.
    INTERNATIONAL JOURNAL OF PSYCHOLOGY, 2016, 51 : 819 - 819
  • [5] Role of binaural hearing in speech intelligibility and spatial release from masking using vocoded speech
    Garadat, Soha N.
    Litovsky, Ruth Y.
    Yu, Gongqiang
    Zeng, Fan-Gang
    Journal of the Acoustical Society of America, 2009, 126 (05): : 2522 - 2535
  • [6] Role of binaural hearing in speech intelligibility and spatial release from masking using vocoded speech
    Garadat, Soha N.
    Litovsky, Ruth Y.
    Yu, Gongqiang
    Zeng, Fan-Gang
    JOURNAL OF THE ACOUSTICAL SOCIETY OF AMERICA, 2009, 126 (05): : 2522 - 2535
  • [7] Effects of envelope bandwidth on the intelligibility of sine- and noise-vocoded speech
    Souza, Pamela
    Rosen, Stuart
    JOURNAL OF THE ACOUSTICAL SOCIETY OF AMERICA, 2009, 126 (02): : 792 - 805
  • [8] Neural correlates of intelligibility in speech investigated with noise vocoded speech- A positron emission tomography study
    Scott, Sophie K.
    Rosen, Stuart
    Lang, Harriet
    Wise, Richard J. S.
    JOURNAL OF THE ACOUSTICAL SOCIETY OF AMERICA, 2006, 120 (02): : 1075 - 1083
  • [9] A Binaural Model Predicting Speech Intelligibility in the Presence of Stationary Noise and Noise-Vocoded Speech Interferers for Normal-Hearing and Hearing-Impaired Listeners
    Lavandier, Mathieu
    Buchholz, Joerg M.
    Rana, Baljeet
    ACTA ACUSTICA UNITED WITH ACUSTICA, 2018, 104 (05) : 909 - 913
  • [10] A Deep Denoising Autoencoder Approach to Improving the Intelligibility of Vocoded Speech in Cochlear Implant Simulation
    Lai, Ying-Hui
    Chen, Fei
    Wang, Syu-Siang
    Lu, Xugang
    Tsao, Yu
    Lee, Chin-Hui
    IEEE TRANSACTIONS ON BIOMEDICAL ENGINEERING, 2017, 64 (07) : 1568 - 1578