Find public datasets about India by topic. Review the source and licence, preview available CSV rows, then download after sign-in or use Python with a free token.
pip install desidata
# Create a free DD token in Profile & settings and set DD_TOKEN once.
import desidata
df = desidata.load("annual-report-2015-16-by-agriwelfare-gov-in-question-and-answer-dataset")
df.head()Catalogue data
Records by category label
Recorded downloads
7611,075
Datasets
3,26,911
Rows of data
5
Categories
761
Downloads
1
Languages
Popular datasets
Datasets with the most downloads on DesiData. Open a page to review its description, source, licence and available preview.
1013 extractive question-and-answer pairs built from Annual Report 2015-16, published by agriwelfare.gov.in. Every answer is a verbatim span of text the source prints, and each row carries the passage it sits in, its offset in that passage, the source quote, the page and the location in the document, so any row can be checked against the original. 380 of the 1013 pairs (37.5%) are explanatory questions and 633 restate a figure. 96.15% of rows pass the corpus quality gate.
National Family Health Survey 5 dataset
5729 extractive question-and-answer pairs built from National Health Profile 2022, published by cbhidghs.mohfw.gov.in. The National Health Profile 2022, published by the Central Bureau of Health Intelligence, provides a comprehensive overview of health statistics, trends, and infrastructure in India. It includes demographic indicators, vital statistics, socio-economic factors, health programs, disease prevalence, healthcare infrastructure, and financial data related to health services. The report is structured into chapters, tables, and lists, presenting state/UT-wise, district-wise, and national-level data. It also discusses health-related initiatives, schemes, and digital health applications, alongside training programs and awareness campaigns. Every answer is a verbatim span of text the source prints, and each row carries the passage it sits in, its offset in that passage, the source quote, the page and the location in the document, so any row can be checked against the original. 0 of the 5729 pairs (0.0%) are explanatory questions and 5729 restate a figure. 96.28% of rows pass the corpus quality gate."
How it works
DesiData helps you discover India-focused public datasets, review their details, and use them in analysis or AI projects.
Step 1
Search or filter the catalogue by topic and language. Dataset pages show a CSV preview when one is available.
Step 2
Check its published description, row count, file format, source and licence before deciding whether it fits your work.
Step 3
Download through the website after signing in, or use the Python package and API with a free DD token.
Categories
Counts reflect the category assigned in the catalog. Open a dataset to check its actual subject, source and licence.
For developers
Install the package, create a free DD token, then load published datasets into pandas or use the public catalogue API.
# install once
pip install desidata
import desidata
# Create a free DD token in Profile & settings and set DD_TOKEN in your environment.
# Then every dataset is one line away.
df = desidata.load("gender-policy-of-nabard-question-and-answer-dataset")
df.head()
# and searchable from Python or your terminal
desidata.search("agriculture")Why DesiData
Browse a single catalogue, review the information available for each dataset, and decide whether it fits your research or project.
Browse public datasets about India across topics such as farming, health, education, transport and the economy.
Search dataset titles and descriptions, then narrow the catalogue by topic or language.
Using data for AI
India-focused data can support research, analysis and AI projects. Its suitability depends on the dataset, licence and intended task.
Confirm the subject, source and coverage match the question you want to answer.
Review the listed licence and the original source terms before using data in a product or model.
Use the CSV preview when available to understand its fields. Check labels, missing values and duplicates for your task.
Before you download
Check the available source, licence, file information and preview to decide whether a dataset suits your project.
Publisher information and a source link are shown when they have been provided.
The listed licence is visible on the dataset page; missing licences are identified as unspecified.
Review the published row count, format, file size and last updated date.
Free to browse and preview. Sign in for website downloads, or use a free DD token with Python and API clients to download datasets.
$ pip install desidataNo credit card or paywall. Every dataset lists its source and licence.
761
Downloads
Current catalog category label
5 datasets
Zero-dependency install. load(), search() and catalog() cover the whole catalogue, with a 24-hour local cache.
desidata on PyPIDownload a generated starter notebook with code to load and explore a dataset. Open it in Colab when a hosted notebook is available.
Browse the notebook indexdesidata search, info, catalog and download — the whole platform from your terminal, no Python needed.
Ships with the packageUse a free DD token for downloads from pandas, Hugging Face, PyTorch, or your warehouse.
Grab a URL from any dataset pageEach dataset page shows the source information provided for that file, so you can follow it back to the publisher when a source link is available.
Inspect sample rows and column names on dataset pages when a preview is available.
Review the listed format, row count, file size and last updated date before you download.
Use the DesiData Python package or API with a free account token to load published datasets into your workflow.
A downloadable CSV is not automatically a model-training dataset. Confirm that it has the labels, structure and permissions your task needs.
Browse the current question-and-answer dataset collection .
Website downloads require sign-in. Python and API downloads use a free DD token.