Explore high-quality datasets for your AI and machine learning projects.
The GHPR dataset is used for empirical research and evaluation of software defect prediction. It is built from GitHub Pull Requests (PRs) and identifies 3,026 defect‑fix records. Each fix is treated as a record, yielding 6,052 learning instances (3,026 defective and 3,026 non‑defective). The dataset is provided in CSV and SQL formats and includes 16 features such as project name, project owner, project description, tags, programming language, pre‑ and post‑fix version IDs, defective code, commit description, commit time, pre‑ and post‑fix file contents, file‑path changes, PR title and description, etc.