返回导师列表
PT
Professor Thomas Hain
Professor · School of Computer Science Regent Court
University of Sheffield · United Kingdomspeechaudio and multimedia technologymachine learningnon-linear methods in speech processinglow bit-rate speech codingmulti-modal systemsimage classificationmicrophone arrays
简介
Thomas Hain obtained the degree 'Dipl.-Ing' in Electrical/Communication Engineering in 1994 from the University of Technology, Vienna. He joined the Speech Technology Group at Philips Speech Processing which he left in a senior position.In 1997 he joined the Speech, Vision and Robotics Group at the Cambridge University Engineering Department as Research Associate and PhD Student. He took up a Lectureship at the SVR group in 2001.In 2004 he joined the Speech and Hearing Group to work as Lecturer in Computer Science. He was promoted to Senior Lecturer in 2008 and Reader in 2011.
代表成果
- Young SJ, Evermann G, Gales MJF, Hain T, Kershaw D, Moore GL, Odell JJ, Ollason D, Povey D, Valtchev V & Woodland PC (2004) The HTK Book. Cambridge, England: Cambridge University Engineering Department.
- Young S, Evermann G, Gales M, Hain T, Kershaw D, Xunying L, Moore G, Odell J, Ollason D, Povey D , Ragni A et al () The HTK Book (for HTK Version 3.5, documentation alpha version). Cambridge University Engineering Department: Cambridge University Engineering Department.
- Sudro PN, Ragni A & Hain T (2025) A comparative study of generative models for child voice conversion.. CoRR, abs/2512.12129.
- Farooq MU & Hain T (2025) Enhancing Low-Resource Speech Recognition With Non-Linear Cross-Lingual Mappings. IEEE Transactions on Audio, Speech and Language Processing, 33, 4653-4666.
- Song H, Zhang L, Gao M, Zhang H, Hain T & Shan L (2025) MS-EmoBoost: a novel strategy for enhancing self-supervised speech emotion representations. Scientific Reports, 15(1). View this article in WRRO
- Hasan M, Jefferson N, Hain T & Dawson J (2022) Automatic detection of behavioural codes in team interactions. Computer Speech & Language, 74, 101339-101339.
- Ravenscroft W, Goetze S & Hain T (2022) Att-TasNet: attending to encodings in time-domain audio speech separation of noisy, reverberant speech mixtures. Frontiers in Signal Processing, 2. View this article in WRRO
- Shi Y, Huang Q & Hain T (2021) H-VECTORS : improving the robustness in utterance-level speaker embeddings using a hierarchical attention model. Neural Networks, 142, 329-339. View this article in WRRO
- El Hannani A, Errattahi R, Salmam FZ, Hain T & Ouahmane H (2021) Evaluation of the effectiveness and efficiency of state-of-the-art features and models for automatic speech recognition error detection. Journal of Big Data, 8.
- Errattahia R, Hannani AEL, Hain T & Ouahmane H (2019) System-independent ASR error detection and classification using Recurrent Neural Network. Computer Speech and Language, 55, 187-199. View this article in WRRO
数据校验于 9/6/2026数据来源