Speech Processing Lab
Speech and Vision Research
Building robust speech systems for recognition, synthesis, and retrieval.
About the Lab
The objective of the Speech and Vision Lab (SVL) is to conduct goal oriented basic research, addressing fundamental issues involved in building robust speech-to-text systems, natural sounding text-to-speech systems, spoken/audio information retrieval and biometrics using speech and video.
Towards these goals the lab works on a speech translation system from one Indian language to another, secure access to information using speech mode, biometrics involving speech, image, text and audio-visual information, content-based information storage and retrieval, and the development of a phonetic engine for Indian languages.
Our ResearchRecent Publications
View all publicationsResearch Areas
Speech Recognition
Building robust speech-to-text systems that work across diverse Indian languages, accents, and acoustic environments.
Text-to-Speech Synthesis
Developing natural sounding text-to-speech systems for Indian languages with focus on prosody and naturalness.
Speaker Recognition
Research on biometric identification using speech signals, including speaker verification and identification systems.
Language Identification
Automatic identification of spoken language from audio signals, particularly for Indian languages.
Emotion Recognition
Speech processing in emotion conditions, including emotion detection and emotion-aware speech conversion.
Projects
Stutter Detection Platform
CompletedSpeech analysis platform created for All India Institute of Speech and Hearing, Mysore for detecting stuttering patterns.
Speech-to-Speech Translation
ActivePerformance measurement platform for broadcast speeches and talks, as part of the Bhashini project.