Medical ontology mapping
Research
Applications
Ontology
Knowledge Extraction
An LLM benchmark on mapping 108 clinical terms to SNOMED CT: GPT-4o reaches an F1 of 96%, far above BERTMap, while open-source models over-generate mappings.

The setup: 108 clinical terms, mapped to SNOMED CT concepts, evaluated across GPT-4o, Claude 3.5 Sonnet v2, Gemini 1.5 Pro, Llama 3.3 70B, DeepSeek R1, and BERTMap as baseline.
The numbers are surprising. GPT-4o achieves 93.75% precision and an F1 of 96.26%, roughly 38 percentage points above BERTMap’s 57.93%. Claude 3.5 Sonnet v2 and Gemini cluster around 60-63%, while the open-source models (Llama, DeepSeek) fall well below, with high false positive rates suggesting they over-generate candidate mappings rather than discriminate cleanly.