23 extractive question-and-answer pairs built from UCI Machine Learning Repository, published by archive.ics.uci.edu. Every answer is a verbatim span of text the source prints, and each row carries the passage it sits in, its offset in that passage, the source quote, the page and the location in the document, so any row can be checked against the original. 18 of the 23 pairs (78.3%) are explanatory questions and 5 restate a figure. 100.00% of rows pass the corpus quality gate.
Use the API URL with your free DD token in notebooks, scripts, and pipelines.
https://www.desidata.in/api/datasets/predict-students-dropout-and-academic-success-question-and-answer-dataset/downloadDataset downloads are free. For Python or API downloads, sign in once and create a free DD token; set it as DD_TOKEN or save it in your notebook's secrets. Requests are linked to your account so your download history and counts stay accurate.
# One-time install: pip install desidata
# Set DD_TOKEN in your environment first (create a free token in Profile & settings).
import desidata
df = desidata.load("predict-students-dropout-and-academic-success-question-and-answer-dataset")
df.head()Sign in with Google to download.
Usable for analysis, but expect some cleaning before you rely on it.
Annual Report 2015-16 by agriwelfare.gov.in - Question and Answer Dataset
Economy · 1,013 rows
Higher Education in India - Question and Answer Dataset
Economy · 1,594 rows
Gender Policy of NABARD - Question and Answer Dataset
Economy · 14 rows
ICAR Annual Report 2025-26 - Question and Answer Dataset
Economy · 1,101 rows
First 10 of 23 rows
| question | answer | context | answer_start | question_type | knowledge_quality_score | source_quote | source_page | source_location | confidence | validation_status |
|---|---|---|---|---|---|---|---|---|---|---|
| What data split was used in the project? | used, in our project, with a data split of 80% for training and 20% for test | What do the instances in this dataset represent? Each instance is a student Are there recommended data splits? The dataset was used, in our project, with a data split of 80% for training and 20% for test. Was there any data preprocessing performed? We performed a rigorous data preprocessing to handle data from anomalies, unexplainable outliers, and missing values. | 127 | definition | 1.000 | The dataset was used, in our project, with a data split of 80% for training and 20% for test. | p[7] | 0.850 | valid | |
| How is the problem in the dataset structured? | a three category classification task | A dataset created from a higher education institution (acquired from several disjoint databases) related to students enrolled in different undergraduate degrees, such as agronomy, design, education, nursing, journalism, management, social service, and technologies. The dataset includes information known at the time of student enrollment (academic path, demographics, and social-economic factors) and the students' academic performance at the end of the first and second semesters. The data is used to build classification models to predict students' dropout and academic sucess. The problem is formulated as a three category classification task, in which there is a strong imbalance towards one of the classes. Tabular Social Science Classification Real, Categorical, Integer | 610 | definition | 1.000 | The problem is formulated as a three category classification task, in which there is a strong imbalance towards one of the classes. | p[0] | 0.700 | valid | |
| What permissions does the CC BY 4.0 license grant? | the sharing and adaptation of the datasets for any purpose, provided that the appropriate credit is given | Instituto Politécnico de Portalegre This dataset is licensed under aCreative Commons Attribution 4. International(CC BY 4.0) license. This allows for the sharing and adaptation of the datasets for any purpose, provided that the appropriate credit is given. By using the UCI Machine Learning Repository, you acknowledge and accept the cookies and privacy practices used by the UCI Machine Learning Repository. APAMLAChicagoVancouverIEEEBibTeX | 150 | definition | 1.000 | This allows for the sharing and adaptation of the datasets for any purpose, provided that the appropriate credit is given. | p[34] | 0.700 | valid | |
| How is the classification task formulated in the dataset? | a three category classification task | A dataset created from a higher education institution (acquired from several disjoint databases) related to students enrolled in different undergraduate degrees, such as agronomy, design, education, nursing, journalism, management, social service, and technologies. The dataset includes information known at the time of student enrollment (academic path, demographics, and social-economic factors) and the students' academic performance at the end of the first and second semesters. The data is used to build classification models to predict students' dropout and academic sucess. The problem is formulated as a three category classification task, in which there is a strong imbalance towards one of the classes. Tabular Social Science Classification Real, Categorical, Integer | 610 | definition | 1.000 | The problem is formulated as a three category classification task, in which there is a strong imbalance towards one of the classes. | p[0] | 0.700 | valid | |
| What does each instance in the dataset represent? | Each instance is a student | Each instance is a student Are there recommended data splits? The dataset was used, in our project, with a data split of 80% for training and 20% for test. Was there any data preprocessing performed? We performed a rigorous data preprocessing to handle data from anomalies, unexplainable outliers, and missing values. Has Missing Values? No By Mónica V. Martins, Daniel Tolledo, Jorge Machado, Luís M. T. Baptista, and Valentim Realinho. 2021 Published in Trends and Applications in Information Systems and Technologies | 0 | definition | 1.000 | Each instance is a student Are there recommended data splits? The dataset was used, in our project, with a data split of 80% for training and 20% for test. | p[7] | 0.700 | valid | |
| What does the license allow users to do with the dataset? | the sharing and adaptation of the datasets for any purpose, provided that the appropriate credit is given. | Instituto Politécnico de Portalegre This dataset is licensed under aCreative Commons Attribution 4. International(CC BY 4.0) license. This allows for the sharing and adaptation of the datasets for any purpose, provided that the appropriate credit is given. By using the UCI Machine Learning Repository, you acknowledge and accept the cookies and privacy practices used by the UCI Machine Learning Repository. APAMLAChicagoVancouverIEEEBibTeX | 150 | definition | 1.000 | This allows for the sharing and adaptation of the datasets for any purpose, provided that the appropriate credit is given. By using the UCI Machine Learning Repository, you acknowledge and accept the cookies and privacy practices used by the UCI Machine Learning Repository. | p[34] | 0.700 | valid | |
| How is the classification problem formulated in the dataset? | as a three category classification task | A dataset created from a higher education institution (acquired from several disjoint databases) related to students enrolled in different undergraduate degrees, such as agronomy, design, education, nursing, journalism, management, social service, and technologies. The dataset includes information known at the time of student enrollment (academic path, demographics, and social-economic factors) and the students' academic performance at the end of the first and second semesters. The data is used to build classification models to predict students' dropout and academic sucess. The problem is formulated as a three category classification task, in which there is a strong imbalance towards one of the classes. Tabular Social Science Classification Real, Categorical, Integer | 607 | definition | 1.000 | The problem is formulated as a three category classification task, in which there is a strong imbalance towards one of the classes. | p[0] | 0.700 | valid | |
| What machine learning techniques are used in the project related to the dataset? | machine learning techniques to identify students at risk at an early stage of their academic path | For what purpose was the dataset created? The dataset was created in a project that aims to contribute to the reduction of academic dropout and failure in higher education, by using machine learning techniques to identify students at risk at an early stage of their academic path, so that strategies to support them can be put into place. The dataset includes information known at the time of student enrollment – academic path, demographics, and social-economic factors. The problem is formulated as a three category classification task (dropout, enrolled, and graduate) at the end of the normal duration of the course. Who funded the creation of the dataset? This dataset is supported by program SATDAP - Capacitação da Administração Pública under grant POCI-05-5762-FSE-000191, Portugal. What do the instances in this dataset represent? | 182 | definition | 1.000 | For what purpose was the dataset created? The dataset was created in a project that aims to contribute to the reduction of academic dropout and failure in higher education, by using machine learning techniques to identify students at risk at an early stage of their academic path, so that strategies to support them can be put into place. | ||||
| What outcomes are the classification models designed to predict? | students' dropout and academic sucess | A dataset created from a higher education institution (acquired from several disjoint databases) related to students enrolled in different undergraduate degrees, such as agronomy, design, education, nursing, journalism, management, social service, and technologies. The dataset includes information known at the time of student enrollment (academic path, demographics, and social-economic factors) and the students' academic performance at the end of the first and second semesters. The data is used to build classification models to predict students' dropout and academic sucess. The problem is formulated as a three category classification task, in which there is a strong imbalance towards one of the classes. Tabular Social Science Classification Real, Categorical, Integer | 542 | summary | 0.970 | The data is used to build classification models to predict students' dropout and academic sucess. | p[0] | 0.700 | valid | |
| What is the goal of creating the dataset? | to contribute to the reduction of academic dropout and failure in higher education | For what purpose was the dataset created? The dataset was created in a project that aims to contribute to the reduction of academic dropout and failure in higher education, by using machine learning techniques to identify students at risk at an early stage of their academic path, so that strategies to support them can be put into place. The dataset includes information known at the time of student enrollment – academic path, demographics, and social-economic factors. The problem is formulated as a three category classification task (dropout, enrolled, and graduate) at the end of the normal duration of the course. Who funded the creation of the dataset? This dataset is supported by program SATDAP - Capacitação da Administração Pública under grant POCI-05-5762-FSE-000191, Portugal. What do the instances in this dataset represent? | 89 | summary | 0.970 | For what purpose was the dataset created? The dataset was created in a project that aims to contribute to the reduction of academic dropout and failure in higher education, by using machine learning techniques to identify students at risk at an early stage of their academic path, so that strategies to support them can be put into place. |
Read straight from the file — download or use the API URL for the full dataset.
| p[7] |
| 0.700 |
| valid |
| p[7] |
| 0.700 |
| valid |