JUHE API Marketplace
High Quality Data

Dataset Hub

Explore high-quality datasets for your AI and machine learning projects.

Sort:

Browse by Category

CodeFeedback-Python105K

Python Programming
Question Answering

This dataset is a subset extracted from the `m-a-p/CodeFeedback-Filtered-Instruction` dataset, specifically selecting 104,848 samples written in Python. The dataset includes two main features: 'query' and 'response', both of string type. It is divided into a training set containing 104,848 samples. The dataset is suitable for question‑answering tasks, in English, with a sample size between 10,000 and 100,000.

huggingface
View Details

asure22/python_obfuscated_small

Code Obfuscation
Python Programming

This dataset is primarily intended for code analysis and processing, and includes multiple code-related features such as repository name, file path, function name, original string, programming language, code, code tokens, docstring, docstring tokens, SHA value, URL, partition, summary, obfuscated code, code length, and obfuscated code length. The dataset is divided into a training split containing 30,000 samples with a total size of 442,939,709.61477566 bytes. The download size of the dataset is 115,314,164 bytes.

hugging_face
View Details

google-research-datasets/mbpp

Python Programming
Code Generation

The Mostly Basic Python Problems (MBPP) dataset contains about 1,000 Python programming problems generated by crowdsourcing and experts, intended for evaluating code generation models. Each problem includes a task description, a code solution, and three automated test cases. The dataset is provided in two versions: full and sanitized, each comprising training, test, validation, and prompt partitions. It was created to assess code generation capabilities and was developed and annotated internally at Google through crowdsourcing efforts.

hugging_face
View Details