Mengqi97/chinese-medical-dataset
Chinese Medical Dataset 致力于详细整理现有尽可能多的中文医学数据集,包括详细的数据汇总、数据示例、下载链接等。
What it solves
This repository provides a comprehensive collection and organization of open-source Chinese medical datasets. It solves the problem of fragmented medical data by aggregating resources for various NLP tasks, including medical question answering, entity recognition, terminology standardization, and knowledge graph construction, making it easier for researchers to find high-quality data for training medical AI.
How it works
The project acts as a curated directory that categorizes various medical datasets by their primary task. It provides detailed summaries for each dataset, including:
- Task Type: Such as classification, medical QA, entity recognition, or relationship extraction.
- Data Volume: The number of training, validation, and test samples.
- Download Links: Direct links to sources like Hugging Face, GitHub, or cloud drives.
- Data Examples: JSON or table-based snippets showing the actual structure of the questions, answers, and labels.
Who it’s for
- AI Researchers: Those developing Large Language Models (LLMs) or specialized models for the healthcare domain.
- ML Engineers: Developers building medical chatbots, intelligent diagnostic systems, or medical information extraction tools.
- Data Scientists: Professionals needing structured Chinese medical text for training or evaluating medical AI models.
Highlights
- Diverse Task Coverage: Includes datasets for medical QA (e.g., Huatuo-26M), entity recognition (e.g., Yidu-S4K), and terminology standardization (e.g., Yidu-N7K).
- Massive Scale: Features the Huatuo-26M dataset, containing over 26 million high-quality medical QA pairs.
- Benchmark Integration: Includes the CBLUE benchmark for evaluating Chinese medical information processing.
- Specialized Domains: Covers specific areas such as Traditional Chinese Medicine (TCM), diabetes research, and autism spectrum disorder (AsdKB).
Related
- Project
- Project
- Project
- Project