Datasets | JuheAPI

cimec/lambada

Natural Language Processing

Text Understanding

The LAMBADA dataset is used to evaluate computational models' text‑understanding ability, specifically testing whether a model can handle long‑range dependencies via a word‑prediction task. The dataset consists of narrative passages extracted from BookCorpus, split into development and test sets, with training data covering the full text of 2,662 novels. Its structure includes text and label fields, and it is partitioned into training, development, and test sets. The dataset was created to assess whether language models can retain long‑term contextual memory. Annotation involved paid crowdworkers ensuring that the target word could only be guessed by reading the entire passage. The language is English and the license is CC BY 4.0.

hugging_face

View Details

Dataset Hub

Browse by Category

allenai/winogrande

cimec/lambada