About desidata.in
India publishes an enormous amount of data. Very little of it is usable. Numbers arrive as scanned PDFs, headers change between years, states report the same field three different ways, and portals go down the week you need them.
We start from the official release
Census tables, MoSPI bulletins, CPCB station feeds, ECI statistical reports, UDISE+ state files. Each dataset names its origin and links back to it.
We clean, and we write down what we did
Unpivoting wide sheets, reconciling district splits, mapping names to LGD codes, recomputing percentages from raw counts. Every dataset carries an extraction note describing those choices.
We document the schema before you download
Column names, types and plain-English descriptions are on the dataset page, so you know whether a file is worth your afternoon.
We keep the licences intact
Source licences travel with the data. Our cleaning notes and derived columns are released under CC BY 4.0.