CONTACT

Speech Processing Lab

Speech and Vision Research

Building robust speech systems for recognition, synthesis, and retrieval.

About the Lab

The objective of the Speech and Vision Lab (SVL) is to conduct goal oriented basic research, addressing fundamental issues involved in building robust speech-to-text systems, natural sounding text-to-speech systems, spoken/audio information retrieval and biometrics using speech and video.

Towards these goals the lab works on a speech translation system from one Indian language to another, secure access to information using speech mode, biometrics involving speech, image, text and audio-visual information, content-based information storage and retrieval, and the development of a phonetic engine for Indian languages.

Our Research

Recent Publications

View all publications
Vowel Based Non-Uniform Prosody Modification for Emotion Conversion
Harikrishna, S. R. Kadiri and A. K. Vuppala
2016
Neutral to Anger Speech Conversion Using Non-Uniform Duration Modification
A. K. Vuppala and S. R. Kadiri
2014
Automatic Detection of Breathy Voiced Vowels in Gujarati Speech
A. K. Vuppala and P. Bhaskararao
2013

Research Areas

Speech Recognition

Building robust speech-to-text systems that work across diverse Indian languages, accents, and acoustic environments.

Text-to-Speech Synthesis

Developing natural sounding text-to-speech systems for Indian languages with focus on prosody and naturalness.

Speaker Recognition

Research on biometric identification using speech signals, including speaker verification and identification systems.

Language Identification

Automatic identification of spoken language from audio signals, particularly for Indian languages.

Emotion Recognition

Speech processing in emotion conditions, including emotion detection and emotion-aware speech conversion.

Projects

Stutter Detection Platform

Completed

Speech analysis platform created for All India Institute of Speech and Hearing, Mysore for detecting stuttering patterns.

Speech-to-Speech Translation

Active

Performance measurement platform for broadcast speeches and talks, as part of the Bhashini project.

Publications

Vowel Based Non-Uniform Prosody Modification for Emotion Conversion
Harikrishna, S. R. Kadiri and A. K. Vuppala
2016
Neutral to Anger Speech Conversion Using Non-Uniform Duration Modification
A. K. Vuppala and S. R. Kadiri
2014
Automatic Detection of Breathy Voiced Vowels in Gujarati Speech
A. K. Vuppala and P. Bhaskararao
2013

Faculty

Professor
Ph.D (IIT Kharagpur)
Speech Recognition, Speaker Recognition, Language Identification, Emotion & Pathological Speech Processing
Speech Processing Lab
Chiranjeevi Yarra
Assistant Professor
Ph.D (IISc, Bangalore)
Speech Processing
Speech Processing Lab
Kishore S Prahallad
Adjunct Faculty
Ph.D (Carnegie Mellon University)
Text-to-Speech, Voice Conversion, Speaker Recognition
Speech Processing Lab