








DrivenData Benchmarks provide a rigorous baseline for comparing different approaches to the same problem. Submit multiple different models, see how your methods perform, and contribute to a public body of knowledge about what works.
Much of the world's healthcare data is stored in free-text documents, such as clinical notes taken by doctors. Because this information is unstructured, it can be difficult to analyze and extract meaningful insights. However, by applying a standardized terminology like SNOMED CT, healthcare organizations can convert this free-text data into a structured format that computers can readily analyze, helping stimulating the development of new medicines, treatment pathways, and better patient outcomes.
One way to analyze clinical notes is to identify and label the portions of each note that correspond to specific medical concepts. This process is called entity linking because it involves identifying candidate spans in the unstructured text (the entities) and linking them to a particular concept in a knowledge base of medical terminology.
Clinical entity linking is inherently challenging. Medical notes are often rife with abbreviations (some of them context-dependent) and assumed knowledge. Furthermore, the target knowledge bases can easily include hundreds of thousands of concepts, many of which occur infrequently leading to a “long tail” effect in the distribution of concepts.

A synthetic example of medical text and labeled concepts. The concepts are highlighted in green, and the indicated concept IDs, names, and categories are shown.
The objective of this benchmark is to compare how well different approaches can structure the unstructured data in clinical notes for meaningful use and analysis, using the SNOMED CT clinical terminology. Participants will train models based on real-world doctor's notes which have been de-identified and annotated with SNOMED CT concepts by medically trained professionals. This is the largest publicly available dataset of labelled clinical notes! By submitting to this benchmark, you can see how well your model performs and compare it against other approaches.
The benchmark rules are in place to promote fair benchmarking and useful solutions. If you are ever unsure whether your solution meets the benchmark rules, ask the challenge organizers in the forum or send an email to info@drivendata.org.
Image courtesy of SNOMED International
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。