JUHE API Marketplace
API CatalogDatasetsDocsBlog
API CatalogDatasetsDocsBlog

Dataset Catalog

Browse trusted datasets for evaluation, enrichment, and production use.

Category index
Showing 1 of 1 datasets
Category: Gene‑Disease Association

DFKI-SLT/GDA

Gene‑Disease AssociationBiomedical Text Mining

The GDA dataset is a sentence-level evaluation dataset for extracting gene‑disease associations, developed by Nourani and Reshadata (2020). Built on the DisGeNET and PubTator databases, it contains 8,000 sentences covering 1,904 diseases and 3,635 genes. The dataset is split into training, validation, and test sets, each instance providing multiple fields such as gene ID, disease name, association type, etc. Construction involved extracting relevant sentences from PubMed abstracts and applying systematic filtering to ensure high‑quality negative samples.

Source hugging_faceUpdated Jun 22, 2024251 viewsLinked
Inspect dataset
JUHE API Marketplace

Accelerate development and ship production-grade integrations with APIs, MCP services, and AI-first infrastructure workflows.

For Developers

ConsoleDocumentation

Product

Browse APIsTemp Mail APIGlobal SMS

Company

What's NewContact SupportTerms Of ServicePrivacy Policy
Copyright © 2026 JUHEDATA HK LIMITED - All rights reserved