To mark its Silver Jubilee, LTRC released BhashaVerse, a set of free-to-download language resources and machine translation models for Indian languages.
What was released
At the heart of BhashaVerse is a multitask encoder-decoder model that can translate across 36 Indian languages. The set of supported languages deliberately reaches beyond the most widely spoken ones to include lower-resource languages such as Tulu, Bodo, Bhojpuri, Magahi, and Santhali.
Alongside the translation models, LTRC also announced a BhashaVerse LLM decoder model for the same 36 languages, aimed at tasks such as summarisation and question answering, and released the Bhashik datasets of Indian language pairs.
Why it matters
Making models and data openly available lowers the barrier for researchers, students, and developers working on Indian languages, many of which have historically lacked the resources available for high-resource languages. Open releases like this one are part of LTRC’s long-running focus on both the basic and applied aspects of language technology.