MT-NLP Lab
Machine Translation and Natural Language Processing
We work on building intelligent systems that understand, interpret and generate human language.
About the Lab
The MT-NLP Lab at LTRC, IIIT-H, focuses on a wide range of NLP problems including syntax and parsing, semantics, discourse processing, machine translation, and computational models inspired by linguistics and machine learning.
We also develop resources and tools for Indian languages and contribute to open-source NLP ecosystem.
Our ResearchResearch Areas
Computational Grammatical Model
Development of the Computational Paninian Grammar framework, treebanks for Hindi and Urdu, and studies in language typology.
Parsing
Constraint-based and data-driven parsers for Indian languages, along with shallow parsers, part-of-speech taggers, and morphological analyzers.
Machine Translation
Transfer-based approaches, automatic learning of transfer rules, and statistical machine translation from English to Indian languages.
Semantics
Purpose-net and unsupervised/semi-supervised word category disambiguation.
Dialogue and Discourse Analysis
Anaphora resolution in text and generation of sentences from words.
Projects
Indian Language to Indian Language MT (ILMT) Phase-II
CompletedMachine translation among Indian languages, developed as a consortium project funded by the Department of Information Technology, Govt. of India (2010-2015).
English to Indian Language MT System Phase-II
CompletedDevelopment of English to Indian language machine translation as a consortium project funded by the Department of Information Technology, Govt. of India (2010-2013).
SSMT - Pilot Systems Development
CompletedSpeech-to-speech machine translation pilot systems, funded by the Office of the Principal Scientific Advisor (2020-2021).
Bahubhashak - IL-IL MT (Pilot Project)
CompletedIndian language to Indian language machine translation pilot, funded by the Ministry of Electronics and Information Technology, Govt. of India (2020-2021).
SWAYAM - National MOOCs
CompletedTranscription, translation and subtitling of MOOC content, funded by the Ministry of Human Resource Development, Govt. of India (2019-2021).
Hindi to English MT for Judicial Domain (HEMTS)
CompletedMachine translation for the judicial domain, funded by the Ministry of Electronics and Information Technology, Govt. of India (2017-2020).
Multilingual Document Summarization in Quasi-Stationary Environment
CompletedDocument summarization research funded by the Defence Research & Development Organisation, Govt. of India (2019-2022).
Multi-Representational and Multi-Layered Treebank for Hindi and Urdu
CompletedTreebank development for Hindi and Urdu, funded by the NSF, USA (2008-2011).