DeepTech Grants in India 2026 (webinar)
67% off today
Platform

Project Vaani

A programme by ARTPARK

A pioneering platform initiative building one of the world’s largest open-source, multimodal datasets for Indic languages.

About Project Vaani

This program involves owning the end-to-end lifecycle of large-scale AI data programs, from strategy and vendor execution to dataset quality, platform delivery, and ecosystem adoption. It includes defining program roadmaps, driving execution across multiple workstreams (vendors, internal teams, ML, platform), and managing budget allocation and resource planning. The program involves designing data collection strategies, identifying and managing vendors across India, building scalable vendor frameworks, and solving real-world constraints of data collection and annotation. It also focuses on defining and driving data quality frameworks, ensuring dataset usability for ASR, TTS, and LLM applications, and driving benchmarking efforts with partners and startups. Key aspects include owning the data platform experience, translating user needs, driving improvements in dataset usability and adoption, and building partnerships. Stakeholder management includes strategic partners (e.g., Google), Government ecosystems (e.g., Bhashini), startups, and research labs, and leading cross-functional teams for data curation and program operations.

Where to apply

Other ARTPARK programmes

Everything about ARTPARK