127 extractive question-and-answer pairs built from Census of India 2011, published by censusindia.gov.in. Every answer is a verbatim span of text the source prints, and each row carries the passage it sits in, its offset in that passage, the source quote, the page and the location in the document, so any row can be checked against the original. 69 of the 127 pairs (54.3%) are explanatory questions and 58 restate a figure. 100.00% of rows pass the corpus quality gate.
Use the API URL with your free DD token in notebooks, scripts, and pipelines.
https://www.desidata.in/api/datasets/post-enumeration-survey-census-of-india-2011-question-and-answer-dataset/downloadDataset downloads are free. For Python or API downloads, sign in once and create a free DD token; set it as DD_TOKEN or save it in your notebook's secrets. Requests are linked to your account so your download history and counts stay accurate.
# One-time install: pip install desidata
# Set DD_TOKEN in your environment first (create a free token in Profile & settings).
import desidata
df = desidata.load("post-enumeration-survey-census-of-india-2011-question-and-answer-dataset")
df.head()Sign in with Google to download.
Usable for analysis, but expect some cleaning before you rely on it.
Language Atlas 2011 - Question and Answer Dataset
Economy · 15 rows
West Bengal Census Document - Question and Answer Dataset
Economy · 5 rows
Census of India 2011: Bihar, Volume 2 - Question and Answer Dataset
Economy · 30 rows
Prakasam District Census 2011 Documentation - Question and Answer Dataset
Economy · 168 rows
Annual Report 2015-16 by agriwelfare.gov.in - Question and Answer Dataset
Economy · 1,013 rows
Higher Education in India - Question and Answer Dataset
Economy · 1,594 rows
Gender Policy of NABARD - Question and Answer Dataset
Economy · 14 rows
ICAR Annual Report 2025-26 - Question and Answer Dataset
Economy · 1,101 rows
First 10 of 127 rows
| question | answer | context | answer_start | question_type | knowledge_quality_score | source_quote | source_page | source_location | confidence | validation_status |
|---|---|---|---|---|---|---|---|---|---|---|
| What is the purpose of the Post Enumeration Survey? | a sample survey conducted immediately after the census in order to assess the coverage and quality of the census enumeration | A massive operation like the population census is no exception, where some amount of error is inevitable considering the fact that a large number of enumerators and supervisors are engaged in the collection of data, in spite of the best of the intentions and efforts to collect the accurate data. Post Enumeration Survey (PES) is a sample survey conducted immediately after the census in order to assess the coverage and quality of the census enumeration. A large number of countries carry out a Post Enumeration Survey (PES) after the completion of the census to scientifically measure the degree of accuracy. | 330 | definition | 1.000 | Post Enumeration Survey (PES) is a sample survey conducted immediately after the census in order to assess the coverage and quality of the census enumeration. | 5 | page=5,block=2 | 0.850 | valid |
| What is the basis of the dual-system estimation procedure? | based on the case-by-case matching of two different and independent sources describing the same event | The dual-system estimation procedure is based on the case-by-case matching of two different and independent sources describing the same event. | 40 | definition | 1.000 | The dual-system estimation procedure is based on the case-by-case matching of two different and independent sources describing the same event. | 24 | page=24,block=11 | 0.850 | valid |
| What is the main goal of the Post Enumeration Survey? | to estimate the magnitude of omissions (under-count) and duplications (over-count) of individuals in the Census | The Post Enumeration Check (PEC), as it used to be called earlier and renamed as Post Enumeration Survey (PES) in the 2001 Census, has become an integral part of the census operations in India since 1951 Census. 1. The primary objective of the PES is to estimate the magnitude of omissions (under-count) and duplications (over-count) of individuals in the Census, or in other words, to determine the coverage error. Omission or under-count included omission of individual persons in enumerated households as well as omission of households and consequently persons in those households. Duplication or over-count included erroneous inclusion of persons in the enumerated households as well as erroneous inclusion of households and consequently persons in those households. | 251 | definition | 1.000 | The primary objective of the PES is to estimate the magnitude of omissions (under-count) and duplications (over-count) of individuals in the Census, or in other words, to determine the coverage error. | 5 | page=5,block=2 | 0.700 | |
| What do Type I and Type II errors refer to in the Post Enumeration Survey? | Omission or duplication of persons due to omission or duplication of households. (ii) Omission or duplication of individuals in enumerated households. | 1. As mentioned above, the coverage error investigated in the PES, consists of two components: (i) Omission or duplication of persons due to omission or duplication of households. (ii) Omission or duplication of individuals in enumerated households. These are called Type I and Type II errors respectively. | 99 | definition | 1.000 | As mentioned above, the coverage error investigated in the PES, consists of two components: (i) Omission or duplication of persons due to omission or duplication of households. (ii) Omission or duplication of individuals in enumerated households. These are called Type I and Type II errors respectively. | 5 | page=5,block=6 | 0.700 | valid |
| What was the objective of PES Schedule I? | to identify the households, which have been omitted or duplicated; in other words, to determine Type I error | Similarly the houseless households may not be available at the place, where they were enumerated. Hence these two types of households were excluded from the scope of the PES. 2. Three main schedules were canvassed in this survey. These are PES Schedule I, PES Schedule IV and PES Schedule VI. The first two schedules relate to coverage error and the last one relates to content error. The basic purposes of canvassing the three schedules were given below: - 1. Schedule I: to identify the households, which have been omitted or duplicated; in other words, to determine Type I error 2. Schedule IV: to find out persons omitted or duplicated in households which have been enumerated in the census; in other words, to determine Total error 3. Schedule VI: to determine content error in selected questions. This schedule was to be canvassed only in a sub-sample of the PES EBs. 2. | 473 | definition | 1.000 | Schedule I: to identify the households, which have been omitted or duplicated; in other words, to determine Type I error 2. | 20 | page=20,block=0 | 0.700 | |
| What was the purpose of PES Schedule VI? | to determine content error in selected questions | Similarly the houseless households may not be available at the place, where they were enumerated. Hence these two types of households were excluded from the scope of the PES. 2. Three main schedules were canvassed in this survey. These are PES Schedule I, PES Schedule IV and PES Schedule VI. The first two schedules relate to coverage error and the last one relates to content error. The basic purposes of canvassing the three schedules were given below: - 1. Schedule I: to identify the households, which have been omitted or duplicated; in other words, to determine Type I error 2. Schedule IV: to find out persons omitted or duplicated in households which have been enumerated in the census; in other words, to determine Total error 3. Schedule VI: to determine content error in selected questions. This schedule was to be canvassed only in a sub-sample of the PES EBs. 2. | 753 | definition | 1.000 | Schedule VI: to determine content error in selected questions. | 20 | page=20,block=0 | 0.700 | valid |
| What is the foundational assumption of the PES estimation methodology? | the assumption of independence between the actual census and PES operations. | As mentioned above, the PES estimation methodology is based on the assumption of independence between the actual census and PES operations. To maintain independence, in addition to the officers and staff of the DCOs, the officers and staff of the Directorate of Economics and Statistics (DES) of the respective State Governments and other government organizations were involved in most of the States and union territories. The enumerators were of the ranks of Assistant Compilers, Compilers, Statistical Assistants, etc. Their work were supervised by the officers of the rank of Senior Statistical Assistants, Investigators, etc. Overall responsibility of conducting the PES in a State/UT was assigned to a senior officer of the DCO, especially appointed as a nodal officer for that purpose. | 63 | definition | 1.000 | As mentioned above, the PES estimation methodology is based on the assumption of independence between the actual census and PES operations. To maintain independence, in addition to the officers and staff of the DCOs, the officers and staff of the Directorate of Economics and Statistics (DES) of the respective State Governments and other government organizations were involved in most of the States and union territories. | ||||
| How are Type II errors determined based on sex and residence? | as the difference of Total and Type I errors, by sex and residence. | The Type I, Type II and total errors by sex and residence at national level annexed at Table 3.1 are higher in urban areas than those in rural areas. Type I error in case of males and females are almost the same in total and rural areas. It is lower in case of females in urban areas compared to the males. There is some variation in case of the Type II errors by residence, which have been arrived as the difference of Total and Type I errors, by sex and residence. | 399 | definition | 1.000 | There is some variation in case of the Type II errors by residence, which have been arrived as the difference of Total and Type I errors, by sex and residence. | 27 | page=27,block=16 | 0.700 | valid |
| What is one purpose of the Post Enumeration Survey (PES)? | an assessment of the quality of the particulars recorded in the census | Content error has been estimated only for matched persons and for selected variables. One of the objectives of the Post Enumeration Survey (PES) is an assessment of the quality of the particulars recorded in the census for the individuals who were enumerated. The following questions were canvassed in the PES for assessing the quality of the particulars collected in census. 1. Name of the person 2. Relationship to the head 3. Sex 4. Age last birthday 5. Current marital status 6. Literacy status 7. Highest educational level attained 8. Type of disability 9. Characteristics of workers and non-workers 10. Economic activity of the main or marginal workers and 11. Fertility particulars 4. Detailed instructions issued to fill up the Schedule VI are given in Appendix I. 4. The data collection for assessing the extent of content error in census was done through Schedule VI. | 148 | definition | 1.000 | One of the objectives of the Post Enumeration Survey (PES) is an assessment of the quality of the particulars recorded in the census for the individuals who were enumerated. | 35 | page=35,block=2 | 0.700 | |
| What factors contribute to response variance in a statistic? | factors that would tend to average out through compensating errors in a large number of repetitions of the experience | The response bias of a statistic for an area is that part of the response error, which would not tend to average out over the work of many interviewers who might be assigned to the area or the many conceivable responses of the respondents in the area. It may and often will differ between types of areas or between one survey or one census and the next, although it may tend to be consistent in direction and to a considerable degree in amount. 4. The response variance of a statistic arises from factors that would tend to average out through compensating errors in a large number of repetitions of the experience, but that may in a particular limited set of measurements have a significant effect on the accuracy of the result. | 497 | definition | 1.000 | The response variance of a statistic arises from factors that would tend to average out through compensating errors in a large number of repetitions of the experience, but that may in a particular limited set of measurements have a significant effect on the accuracy of the result. | 35 | page=35,block=2 |
Read straight from the file — download or use the API URL for the full dataset.
| valid |
| valid |
| 22 |
| page=22,block=0 |
| 0.700 |
| valid |
| valid |
| 0.700 |
| valid |