Bridging the knowledge gap for the majority of India's population — making science, biodiversity, and innovation accessible in Hindi and Telugu.
India records 117 million daily Wikipedia views, yet the vast majority of scientific content remains locked in English — out of reach for most Indian-language readers. We are fixing this at scale.
Inspired by the success of native-language knowledge platforms elsewhere in the world, we are building a lasting Indian-language scientific corpus for generations to come.
From biological sciences and chemistry to biodiversity, species taxonomy, and landmark inventions — every domain covered in both Hindi and Telugu.
Automated bot pipelines and NLP translation tools allow us to process WikiData and WikiSpecies at a scale impossible with manual effort alone.
Planned Wiki Pages
WikiSpecies Data Source
Target Languages
(Hindi + Telugu)
Project Duration
(Months)
Systematic creation of pages covering landmark scientists, scientific inventions, and core biological and chemical science topics. Data sourced from WikiData with automated translation pipelines.
Expansion to species-level coverage including scientific naming conventions, taxonomic classification, and biodiversity documentation. Data sourced from WikiSpecies (800K+ entries).
| Phase | Focus Domain | Data Source | Target Articles |
|---|---|---|---|
| Year 1 | Scientists & Scientific Inventions | WikiData | 12,000 – 15,000 |
| Year 2 | WikiSpecies & Biodiversity | WikiSpecies | ~15,000 |
Harvesting structured data from WikiData and WikiSpecies — over 800,000 species entries.
Automated Wikipedia bots programmed to generate, format, and upload article skeletons.
Translation models fine-tuned on low-resource Indic language corpora.
Human-in-the-loop validation by LTRC linguists and domain experts before publication.
All tools, datasets, and corpora released open-source for the global research community.
Language Technologies Research Centre (LTRC), International Institute of Information Technology Hyderabad (IIITH). Kohli Building, Gachibowli, Hyderabad – 500032.
Language Technologies Research Centre (LTRC), International Institute of Information Technology Hyderabad (IIITH). Kohli Building, Gachibowli, Hyderabad – 500032.
| Focus Area | Publication Highlights |
|---|
All datasets, bot scripts, translation models, and automated pipelines created under this project are published as open-source resources under the IndicWiki initiative at LTRC. We believe in open science — building tools that empower researchers, developers, and educators worldwide.
Species and topic-level datasets — IndicWiki's own generation repositories plus the external reference sources they draw on.
Native Hindi or Telugu speakers with scientific background can help review and refine bot-generated drafts.
See open datasets →Institutions and researchers working on low-resource NLP or Indic Wikimedia projects are welcome to collaborate.
Visit LTRC, IIIT Hyderabad →Milestones, datasets, and publications are tracked openly as the two-year roadmap unfolds.
View the roadmap →