Tytuł artykułu
Autorzy
Identyfikatory
Warianty tytułu
The systems for automatic speech signal processing, analysis and recognition
Języki publikacji
Abstrakty
Automatyczne rozpoznawanie i rozumienie języka mówionego jest pierwszym i prawdopodobnie najważniejszym krokiem na drodze do konstrukcji systemów automatycznych, które pozwolą na swobodną konwersację i interakcję człowieka z maszyną (komputerem). Prowadzone w ostatnich dekadach ubiegłego stulecia intensywne prace na tym fascynującym polu badawczym doprowadziły już do wielu wymiernych i ciekawych rezultatów. W artykule dokonano skrótowego przeglądu stosowanych obecnie technik, służących do automatycznego przetwarzania, rozpoznawania i interpretowania sygnału mowy. Obecnie uzyskane rezultaty badawcze doprowadziły już do powstania wielu praktycznie wykorzystywanych aplikacji, wśród których największą popularnością cieszą się: oprogramowanie pozwalające na automatyczną konwersję języka mówionego do tekstu (możliwość głosowego wprowadzania danych do komputera) oraz wykorzystywane w telekomunikacji systemy pozwalające na głosowe wybranie numeru wewnętrznego. W artykule przedstawiono podstawy technik wykorzystywanych w procesie konstrukcji komputerowych systemów dialogowych. Szczególną uwagę skupiono na zagadnieniach percepcyjnego przetwarzania sygnału mowy, które stanowią podstawę rozpoznania poszczególnych fonemów zawartych w wypowiedzi, będących podstawowymi, nieredukowalnymi składnikami (atomami) sylab i wyrazów. Bowiem, prawidłowe rozpoznanie fonemów stanowi klucz do konstrukcji każdego systemu komputerowego rozpoznawania i rozumienia sygnału mowy.
Automatic recognition and understanding of spoken language is the first and probably the most important step towards natural humanmachine interaction. Research in the fascinating field in the past few decades has produced remarkable results. In this paper the development of the spoken language technology is summarised. Today, research results in spoken language processing have led to a number of successfull applications, ranging from dictation software for personal computers and telephone-call processing systems for automatic call routing In the paper the basic of natural language processing and recognition are given. The special attention is paid to the poblems of perceptual speech processing, which is an important step towards phoneme recognition. Phonermes are the most basic elements of speech so their proper recognition is the condition of utterance understanding by computer.
Wydawca
Rocznik
Tom
Strony
12--16
Opis fizyczny
Bibliogr. 47 poz.
Twórcy
Bibliografia
- 1. A. Shomali: Rozpoznawanie mówcy na podstawie długookreso- wego histogramu amplitud sygnału mowy, Rozprawa doktorska, AGH, Kraków, 1999.
- 2. L. Bu: Perceptual speech processing and phonetic feature map- ping for robust vowel recognition, IEEE Transactions on Speech and Audio Processing, vol. 8, no. 2, March 2000, p. 105-114. 3. J. B. Allen: How do humans process and recognize speech, IEEE Transactions on Speech and Audio Processing, vol. 2, no. 4, 1994, p. 567-577.
- 4. R. F. Lyon, C. Mead: An analog electronic cochlea, IEEE Tran- sactions on Acoustics, Speech, Signal Processing, vol. 36, July 1988, p. 1119-1134.
- 5. O. Ghitza: Auditory models and human performance in tasks re- lated to speech coding and speech recognition, IEEE Tran- sactions on Speech and Audio Processing, vol. 2, January 1994, p. 115-132.
- 6. K. Wang, S. Shamma: Self-normalization and noise-robustness in early auditory representation, IEEE Transactions on Speech and Audio Processing, vol. 2, July 1994, p. 421-435.
- 7. D. P. Morgan: Neural networks and speech processing, Boston, MA: Kluwer, 1991.
- 8. T. Kohonen: The self-organizing map, Proceedings of IEEE, vol. 78, September 1990, p. 1464-1480.
- 9. R. Lippmann: An introduction to computing with neural nets, IEEE ASSP Magazine, vol. 4, April 1987, p. 4-22.
- 10. E. Levin: Hidden control neural architecture modeling of nonline- ar time varying systems and its applications, IEEE Transactions on Neural Networks, vol. 4, January 1993, p. 109-116. 11. K. R. Farrell, J. Mammone, K. T. Assaleh: Speaker recognition using neural networks and conventional classifiers, IEEE Trans- actions on Speech and Audio Processing, vol. 2, january 1994, p. 194-205.
- 12. A. J. Robinson: An application of recurrent nets to phone proba- bility estimation, IEEE Transactions on Neural Networks, vol. 5, March 1994, p. 298-305.
- 13. H. Bourlard, N. Morgan: Continuous speech recognition by con- nectionist statistical methods, IEEE Transactions on Neural Ne- tworks, vol. 4, November 1993, p. 893-909.
- 14. C. Dugast, L. Devillers, X. Aubert. Combining TDNN and HMM in hybrid system for improved continous-speech recognition, IEEE Transactions on Speech and Audio Processing, vol. 2, January 1994, p. 217-223.
- 15. G. Zavaliagkos, Y. Zhao, R. Schwartz, J. Makhoul: A hybrid seg- mental neural net/hidden Markov model system for continuous speech recognition, IEEE Transactions on Speech and Audio Processing, vol. 2, January 1994, p. 151-160.
- 16. G. Rigoll: Maximum mutual information neural networks for hybri- de connectionist-HMM speech recognition system, IEEE Trans- actions on Speech and Audio Processing, vol. 2, January 1994, p. 175-184.
- 17. J. M. Zurada: Introduction to artificial neural systems, Singapore: West, 1992.
- 18. L. Rabiner, B. H. Juang: Fundamentals of speech recognition, Englewood Cliffs, NJ: Prentice-Hall, 1993.
- 19. H. Hermansky, N. Morgan: RASTA processing of speech, IEEE Transactions on Speech and Audio Processing, vol. 2, October 1994, p. 578-589.
- 20. J. B. Allen: Cochlea modeling, Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing, 1981, p. 768-769.
- 21. M. Kapusta, M. Gajer, A. Shomali: Zastosowanie sieci Kohonena do rozpoznawania mowy patologicznej, Pomiary Automatyka Kontrola nr 7/2000, p. 10-15.
- 22. M. Kapusta, M. Gajer, A. Shomali: Wykorzystanie techniki obli- czeń neuronowych do przetwarzania i rozpoznawania sygnałów mowy, Pomiary Automatyka Kontrola nr 7/2000, p. 16-18. 23. P. Buser, M. Imbert. Audition, Cambridge, MA: MIT Press, 1992. 24. W. A. Yost. Fundamentals of hearing, an introduction, San Diego, CA: Academic, 1994.
- 25. N. Jayant, J. Johnston, R. Safranek: Signal compression based on models of human perception, Proceedings of IEEE, vol. 81, October 1993, p. 1385-1422.
- 26. W. Chang, C. Wang: A masking-threshold-adapted weighting fil- ter for excitation search, IEEE Transactions on Speech and Au- dio Processing, vol. 4, March 1996, p. 124-132. 27. N. Virag: Speech enhancement based on masking properties of the auditory system, Proceedings of IEEE Conference on Acou- stics, Speech and Signal Processing, 1995, p. 796-799.
- 28. A. A. Azirani, R. L. B. Jeannes, G. Faucon: Optimizing speech enhancement by exploiting masking properties of the human ear, Proceedings of IEEE Conference on Acoustics, Speech and Si- gnal Processing, 1995, p. 800-803.
- 29. J. B. Allen: Harvey Fletcher's role in the creation of communica- tion acoustics, J. Acoust. Soc. Amer., vol. 99, no. 4, 1996, p. 1825-1839.
- 30. J. Allen, S. Neely. Modeling the relation between the intensity JND and loudness for pure tones and wide-band noise, J. Aco- ust. Soc. Amer., vol. 102, no. 6, 1997, p. 3628-3646.
- 31. J. D. Markel, A. H. Gray. Linear prediction of speech, Santa Barbara, CA: Speech Community Research Laboratory, 1976.
- 32. J. W. Picone: Signal modeling techniques in speech recognition,
- Proceedings of IEEE, vol. 81, April 1993, p. 1215-1247. 33. D. Greenwood: The mel-scale's disqualifying bias and a consi- stency of pitch-difference equisections in 1965 with equal cochle- ar distances and equal frequency ratios, Hearing Research, vol. 102, 1997, p. 199–248.
- 34. T. D. Chiueh, H. K. Tsai: Multivalued associative memories based on recurrent networks, IEEE Transactions on Neural Networks, vol. 4, March 1993, p. 364-366.
- 35. P. R. Krishnaiah, L. N. Kanal: Classification, pattern recognition and reduction of dimensionality, New York: Elsevier, 1982. 36. K. Miyawaki, W. Strange, R. Verbrugge, A. M. Liberman, J. J. Jenkins, O. Fujimura: An effect of linguistic experience: The discrimination of /r/ and /l/ by native speakers of Japanese and English, Percept. Psychophys., 1975, p. 331-340.
- 37. W. Strange, T. Halwes: Confidence rating in speech perception
- discrimination testing, Percept. Psychophys., 1971, p. 182-186. 38. J. K. Chen, F. K. Soong: An N-best candidates-based discrimina- tive training for speech recognition applications, IEEE Transac- tions on Speech and Audio Processing, vol. 2, January 1994, p. 206-216.
- 39. L. S. Lee, C. Y. Gu, F. H. Li, C. H. Chang, Y. H. Lin, S. L. Tu, S. H. Hsieh, C. H. Chen: Golden Mandarin (1) A real-time Man- darin speech dictation machine for Chinese language with very large vocabullary, IEEE Transactions on Speech and Audio Pro- cessing, vol. 1, April 1993, p. 158-178.
- 40. V. Zue, S. Seneff, J. R. Glass, J. Poliforni, C. Pao, T. J. Hazen, L. Hetherington: Jupiter: atelephone-based conversational inter- face for weather information, IEEE Transactions on Speech and Audio Processing, vol. 8, no. 1, January 2000, p. 85-95. 41. D. Goddeau: Galaxy: A human-language interface to on-line travel information, Proc. ICSLP, Yokohama, Japan, p. 707-710. 42. S. Seneff. Galaxy-II: A reference architecture for conversational system development, Proc. ICSLP, Sydney, Australia, 1998, p. 931-934.
- 43. J. Glass, J. Chang, M. McCandless: A probabilistic framework for feature-based speech recognition, Proc. ICSLP, Philadelphia, PA, October 1996, p. 2277-2280.
- 44. J. Glass, T. Hazen, L. Hetherington: Real-time telephone-based speech recognition in the Jupiter domain, Proc. ICASSP, Phoe- nix, AZ, 1999, p. 61-64.
- 45. S. Seneff. TINA: A natural language system for spoken language applications, Comput. Linguist., vol. 18, 1992, p. 61-86.
- 46. J. Glass, J. Poliforni, S. Seneff. Multilingual language generation across multiple domains, Proc. ICSLP, Yokohama, Japan, 1994, p. 983-986.
- 47. V. Zue: Conversational interfaces: Advances and challenges, Proc. Eurospeech, Rhodes, Greece, 1997, p. KN9-KN18.
Typ dokumentu
Bibliografia
Identyfikator YADDA
bwmeta1.element.baztech-article-BWA9-0002-0057
JavaScript jest wyłączony w Twojej przeglądarce internetowej. Włącz go, a następnie odśwież stronę, aby móc w pełni z niej korzystać.