As artificial intelligence increasingly becomes embedded in digital services, the availability of language technologies can determine who benefits from these advances. While mainstream AI models have expanded support for several Indian languages, many of Northeast India’s more than 220 indigenous languages remain largely underserved by digital speech and language technologies.
Shillong-based research and AI deployment lab MWire Labs is working to address this gap by developing foundational language AI infrastructure for the region. Its first product, Lemka, is an end-to-end speech stack offering speech-to-text (STT) and text-to-speech (TTS) capabilities for Khasi, Garo, Mizo, Meitei, Nagamese and Kokborok.
Lemka combines open-source resources with more than 1,000 hours of proprietary field audio to improve performance across real-world speech conditions, dialectal variations, code-switching, and diverse acoustic environments. The company is also developing the broader NE-Stack, covering language identification, OCR, translation and speech technologies across languages in the region.
In this interview with AI Spectrum, Badal Nyalang, Director, MWire Labs, discusses the technical and linguistic challenges of building AI for low-resource languages, the importance of community-driven data collection, how field audio can improve deployment readiness, and the potential applications of speech AI across healthcare, education, public services, media and digital commerce. He also outlines MWire Labs’ approach to expanding language infrastructure across Northeast India while keeping open-source development and local communities at the centre of its strategy.
What are the biggest technical challenges in developing speech recognition and text-to-speech systems for low-resource languages such as those spoken across Northeast India?
Low-resource Northeast languages have very little digital text or audio to begin with, and what exists is often unstandardized; spelling conventions, scripts, and even dialect boundaries are still evolving for several of these languages. On top of that we're dealing with heavy code-switching (with English and Hindi), rich morphology, and very little existing benchmark data to measure progress against. So a large part of the work is building the data and evaluation infrastructure before the modelling even starts.
Also, collecting speech data for many of the regional tribal languages are difficult because of the stigma and misconception and beliefs about language misrepresentation.
How does the 1,000+ hours of proprietary field audio collected across the region contribute to Lemka’s performance, and what makes field-collected data particularly important for capturing linguistic diversity, accents, and real-world speech patterns?
Most available speech data for Indian languages is scripted, studio-recorded, or scraped from a narrow set of sources. Field audio gives us spontaneous, real-world speech from different devices, different dialects across sub-regions, natural code-switching, background noise, and the accent variation you actually encounter in a clinic, a classroom, or a government office. Models trained only on clean, read speech tend to break down exactly in these real conditions, so the field data is what makes Lemka usable outside a lab setting. That being said, the TTS data was then again collected in a clean studio environment only.
How does Lemka’s STT and TTS architecture compare with mainstream multilingual speech models, particularly in terms of accuracy, latency, and performance in low-resource language environments?
Most mainstream multilingual speech models either don't support these languages at all or support them so thinly that accuracy drops sharply compared to high-resource languages. Lemka is built specifically for this language set rather than treating it as an afterthought bolted onto a larger multilingual model, so it's tuned to the phonetics, dialectal variation, and real-world acoustic conditions of the region, which is where general-purpose models tend to struggle most. Recent launches like SraVaani by IISC Artpark are a big leap; a lot of non-scheduled languages are included. However, there is still gaps; a lot of these are basic coverage and not deployment readiness, which we are trying to solve by moving one language at a time.
What steps has MWire Labs taken to ensure that the development of Lemka reflects local linguistic and cultural contexts, rather than simply adapting models built primarily for high-resource languages?
Working with native-speaker linguists and community contributors throughout data collection and verification. Outputs are checked by native speakers as part of the pipeline, not just benchmarked against a generic metric, so the system reflects how the language is actually spoken rather than a textbook approximation of it. Speaking of numbers, even our open-source speech-to-text gets 9.3 WER (Word error rate) for Garo and 14 WER for Khasi. This is the open-source version created with a mix of Vaani and a small amount of our proprietary data. Lemka on the other is built with a combination of open source + 1,000 hours of proprietary data, focusing on deployment.
How do you see speech AI such as Lemka being applied in areas such as education, healthcare, public services, local media, and digital commerce across Northeast India?
We see the clearest near-term use cases in education (assistive tools for learners who aren't comfortable in English or Hindi), healthcare (patient communication where language is currently a barrier), public services (voice access to government information for non-literate or non-English-speaking citizens), local media (content localization and dubbing), and digital commerce (voice interfaces for local businesses and customer support). The common thread is giving people access in the language they actually think and speak in.
Looking ahead, how does MWire Labs plan to scale its language AI infrastructure across the region’s 220+ indigenous languages, and what role could open-source data, local communities, and partnerships play in achieving this goal?
Lemka is one product built on top of a broader foundation layer we've been developing, the NE-Stack - which already includes language identification, OCR, translation, and speech models across the region's languages, not just the six Lemka currently covers. Open-source data and community partnerships are central to this: much of our underlying research is released openly, and we rely on local linguists and community contributors to keep expanding coverage responsibly, rather than trying to do it top-down.

