Explore high-quality datasets for your AI and machine learning projects.
The dataset contains Chinese and English math word problems with answers, suitable for text generation tasks, especially the generation of math application problems. It is split into a training set (7,473 samples) and a test set (1,319 samples). Features include question, answer, Chinese question, and answer‑only fields.
GSM8K (Grade School Math 8K) is a dataset of 8.5 K high‑quality, linguistically diverse elementary mathematics word problems. It supports question‑answering tasks that require multi‑step reasoning, typically involving 2–8 steps of basic arithmetic (+, –, ×, ÷). Problems are of middle‑school difficulty and most can be solved without explicitly defining variables. Solutions are provided in natural language rather than pure mathematical notation. The dataset offers two configurations, "main" and "socratic," each with different answer formats.