For developers
Browse the catalogue and previews publicly. Sign in for website downloads, or create a free DD token for Python, command-line, and API downloads. The core package has zero dependencies.
DesiData datasets are free. You can browse catalogue metadata and previews without a token; Python, CLI, and API dataset downloads require one. Sign in and create a DD token in Profile & settings. Copy it when it is shown because the full token is displayed only once.
Set it before starting Python or the DesiData CLI. Keep it private: don't paste it into shared notebooks or commit it to GitHub. Revoke it from Profile & settings if it is exposed.
# macOS or Linux (current terminal)
export DD_TOKEN="paste-your-token-here"
# Windows PowerShell (current terminal)
$env:DD_TOKEN = "paste-your-token-here"Add your token under the notebook's Secrets panel as DD_TOKEN, enable notebook access to it, then run this setup cell before downloading data.
from google.colab import userdata
import os
# First save your token in Colab Secrets with the name "DD_TOKEN".
os.environ["DD_TOKEN"] = userdata.get("DD_TOKEN")This runs the real desidata package from PyPI inside your browser via WebAssembly — pandas included, hitting the live API. Edit the code, press Run. Catalogue metadata is public; file downloads require a personal DD token.
Three steps from nothing to a DataFrame. load() accepts a bare slug or a full desidata.in dataset URL, and any extra keyword arguments go straight to pandas.read_csv.
# 1. Install (once)
pip install desidata pandas
# 2. Create a free DD token in Profile & settings and set DD_TOKEN in your environment.
# 3. Load any dataset — one line
import desidata
df = desidata.load("gender-policy-of-nabard-question-and-answer-dataset")
# 4. Build.
df.head()Catalogue search and metadata remain public. Dataset loads and downloads requireDD_TOKEN and are cached locally for 24 hours.
import desidata
# load(slug) -> pandas DataFrame (kwargs pass through to read_csv)
df = desidata.load("crop-prices-district-wise", dtype={"year": int})
# search(query) -> list of matching datasets
results = desidata.search("nabard")
# catalog() -> every published dataset, optionally filtered
economy = desidata.catalog(category="Economy")
# info(slug) -> metadata dict (name, size, downloads, tags, ...)
meta = desidata.info("crop-prices-district-wise")
# download(slug, path) -> requires DD_TOKEN; returns bytes or writes a file
data = desidata.download("crop-prices-district-wise", "prices.csv")
# Caching: loads are cached for 24h in ~/.desidata/cache
df = desidata.load("crop-prices-district-wise", refresh=True) # force re-download
desidata.clear_cache() # wipe the cacheThe package ships a desidata command — browse public metadata or download files after setting your free DD_TOKEN; no Python scripting needed.
$ desidata search nabard
$ desidata info gender-policy-of-nabard-question-and-answer-dataset
$ desidata catalog --category Agriculture
$ desidata download crop-prices-district-wise -o prices.csv
$ desidata clear-cacheThe datasets are free. Sign in once and create a free DD token in Profile & settings to download with the API or Python package. Keep it in theDD_TOKEN environment variable or a notebook secret such as Colab Secrets; don't put it in code or commit it. Requests are linked to your account for download history and counts. The API works from R, Julia, JavaScript, curl, or any client that can send an Authorization header.
| Endpoint | Returns |
|---|---|
| GET/api/datasets | Public catalogue metadata (filters: ?query=, ?category=, ?language=) |
| GET/api/datasets/:slug | Public metadata for one dataset, including a protected download_url |
| GET/api/datasets/:slug/download | Short-lived download_url JSON; requires Authorization: Bearer DD_TOKEN |
| GET/api/datasets/:slug/notebook | Generated starter notebook (.ipynb) |
| GET/api/datasets/:slug/preview | First rows + column names |
| GET/api/notebooks/manifest | Index of every starter notebook |
# Public catalogue discovery needs no token
curl -s "https://www.desidata.in/api/datasets?category=Agriculture" | python -m json.tool
# Dataset files require a free DD_TOKEN from Profile & settings.
# The authenticated request returns a short-lived file URL as JSON.
file_url=$(curl -fsS -H "Authorization: Bearer $DD_TOKEN" \
-H "Accept: application/json" \
"https://www.desidata.in/api/datasets/crop-prices-district-wise/download" \
| python -c 'import json,sys; print(json.load(sys.stdin)["download_url"])')
# Fetch the file separately, without forwarding DD_TOKEN.
curl -fL "$file_url" -o prices.csvEach dataset page offers a generated notebook with code to load and explore the file. Download it from the page, or open it directly in Colab when a hosted copy is available. Add your free DD token as a Colab secret before running cells that download data.
Every error the package raises subclasses DesiDataError, so one except clause catches everything. Loads are cached in ~/.desidata/cache for 24 hours — override the location with DESIDATA_CACHE_DIR, the API root with DESIDATA_BASE_URL.
from desidata import DesiDataError, DesiDataNotFound, DesiDataAuthenticationError
try:
df = desidata.load("a-slug-that-does-not-exist")
except DesiDataNotFound:
print("Check the slug with desidata.search(...)");
# Every error subclasses DesiDataError:
# DesiDataNotFound -> HTTP 404, dataset doesn't exist
# DesiDataAuthenticationError -> DD_TOKEN is missing, invalid, or revoked
# DesiDataConnectionError -> network/DNS/timeout failure
# DesiDataServerError -> desidata.in returned an unexpected errorDataset access remains free. Sign in once to create a token, keep it in an environment variable, and revoke it from Profile & settings if it is exposed. Each dataset page lists its source and licence.