A CSIR-ASPIRE Funded Research Initiative

Enhancing Scientific Wiki
Articles in Indic Languages

Bridging the knowledge gap for the majority of India's population — making science, biodiversity, and innovation accessible in Hindi and Telugu.

RM
Dr. Radhika Mamidi
Principal Investigator · LTRC, IIIT Hyderabad
The Mission

Democratizing Scientific Knowledge

🌐

The Knowledge Gap

India records 117 million daily Wikipedia views, yet the vast majority of scientific content remains locked in English — out of reach for most Indian-language readers. We are fixing this at scale.

🇮🇳

Language Empowerment

Inspired by the success of native-language knowledge platforms elsewhere in the world, we are building a lasting Indian-language scientific corpus for generations to come.

🔬

Scientific Coverage

From biological sciences and chemistry to biodiversity, species taxonomy, and landmark inventions — every domain covered in both Hindi and Telugu.

🤖

AI-Powered Scale

Automated bot pipelines and NLP translation tools allow us to process WikiData and WikiSpecies at a scale impossible with manual effort alone.

0

Planned Wiki Pages

0

WikiSpecies Data Source

0

Target Languages
(Hindi + Telugu)

0

Project Duration
(Months)

Language Impact

Projected Knowledge Expansion

Articles by Language
Projected article distribution across Indic language ecosystems
Coverage vs. English Wikipedia
Scientific topic coverage comparison — current vs. projected
Progress Dashboard

Article Creation Trajectory

Cumulative Articles Created (24-Month Projection)
Month-by-month cumulative article output across both target languages
Implementation Plan

Two-Year Roadmap

Year 1 · 2025–2026

Scientists, Inventions & Biological/Chemical Sciences

Systematic creation of pages covering landmark scientists, scientific inventions, and core biological and chemical science topics. Data sourced from WikiData with automated translation pipelines.

📄 12,000 – 15,000 Articles
Year 2 · 2026–2027

WikiSpecies, Biodiversity & Taxonomy

Expansion to species-level coverage including scientific naming conventions, taxonomic classification, and biodiversity documentation. Data sourced from WikiSpecies (800K+ entries).

📄 ~15,000 Species Pages
Phase Focus Domain Data Source Target Articles
Year 1 Scientists & Scientific Inventions WikiData 12,000 – 15,000
Year 2 WikiSpecies & Biodiversity WikiSpecies ~15,000
Technology

Automated Technological Pipeline

01 🗄️

Data Sourcing

Harvesting structured data from WikiData and WikiSpecies — over 800,000 species entries.

02 🤖

Bot Creation

Automated Wikipedia bots programmed to generate, format, and upload article skeletons.

03 🔤

NLP Translation

Translation models fine-tuned on low-resource Indic language corpora.

04

Quality Review

Human-in-the-loop validation by LTRC linguists and domain experts before publication.

05 🌍

Open Release

All tools, datasets, and corpora released open-source for the global research community.

Research Team

Leadership & Publications

RM

Dr. Radhika Mamidi

Principal Investigator

Language Technologies Research Centre (LTRC), International Institute of Information Technology Hyderabad (IIITH). Kohli Building, Gachibowli, Hyderabad – 500032.

KK

Krupal Kasyap

Program Manager

Language Technologies Research Centre (LTRC), International Institute of Information Technology Hyderabad (IIITH). Kohli Building, Gachibowli, Hyderabad – 500032.

Focus Area Publication Highlights
Open Source

Resources & Access

Everything Open. Always.

All datasets, bot scripts, translation models, and automated pipelines created under this project are published as open-source resources under the IndicWiki initiative at LTRC. We believe in open science — building tools that empower researchers, developers, and educators worldwide.

Domain-Specific Data Sources

Species and topic-level datasets — IndicWiki's own generation repositories plus the external reference sources they draw on.

Collaborate

Get Involved

Contribute an Article

Native Hindi or Telugu speakers with scientific background can help review and refine bot-generated drafts.

See open datasets →

Partner With Us

Institutions and researchers working on low-resource NLP or Indic Wikimedia projects are welcome to collaborate.

Visit LTRC, IIIT Hyderabad →

Follow Progress

Milestones, datasets, and publications are tracked openly as the two-year roadmap unfolds.

View the roadmap →