DIAアカウントをお持ちの場合、サインインしてください。

サインイン

ユーザーIDをお忘れですか? or パスワードをお忘れですか?

メニュー 戻る Poster-Presentations-Details

P250: Assessing the Impact of Quality Issues on Vocabulary Mappings in Downstream Data





Poster Presenter

      Niamh Catherine McGuinness

      • Director, Pharma Solutions, IQVIA Applied AI Science
      • IQVIA
        United States

Objectives

In studies utilizing real-world data sources, such as electronic health records and insurance claims, vocabulary mappings are crucial for accurate and compliant data source integration. This work quantifies the impact of vocabulary mappings on quality of downstream data in integration processes.

Method

We compared 1,747 vocabulary mappings relating an International Classification of Diseases, Tenth Revision (ICD10) term to a single Systematized Nomenclature of Medicine Clinical Terms (SNOMED CT) concept, as defined by Observational Health Data Sciences and Informatics (OHDSI) Standardized Vocabularies 2024-08-30 and SNOMED CT International 2024-02-01. Only OHDSI active mappings and properly classified SNOMED map source concepts were considered.

Results

We summarize here the results detailed in the article 'Quality Issues on Mappings Between ICD10 and SNOMED CT' presented at the 16th International Semantic Web Applications and Tools for Health Care and Life Sciences Conference. 1,747 ICD10 codes are mapped to a single SNOMED CT code by OHDSI and SNOMED. Of them, 1,266 ICD10 codes are mapped to the same SNOMED CT code by both systems, i.e.: the mappings match, while the remaining 481 ICD10 are mapped to different SNOMED CT codes: i.e.: 27.5% of mappings mismatch. Our analysis found that 27.5% of the mappings do not match due to differences on the level of abstraction of the mappings (47%), slightly variations on the semantics of the terms, i.e.: mapping targets are siblings (10%), evolution of the vocabularies (4%), and (d) plain errors on the release of mappings (2%). Identification of the causes for the remaining mapping mismatches (37%) will be tackled in future works. 147 of the mapping mismatches (30.6%) are due to OHDSI mapping targets which are direct ancestors of the SNOMED counterpart. 71 of the mapping mismatches (14.8%) are due to OHDSI mapping targets which are transitive ancestors of the SNOMED counterpart. 5 of the mapping mismatches (1%) are due to SNOMED mapping targets which are direct ancestors of the OHDSI counterpart. 1 of the mapping mismatches (0.2%) are due to SNOMED mapping targets which are transitive ancestors of the OHDSI counterpart. 48 of the mapping mismatches (10%) are due to OHDSI and SNOMED mapping targets with a direct common ancestor, i.e.: the mapping targets are siblings. 13 of the mapping mismatches (2,7%) are due to mapping target codes that do not exist on the SNOMED CT release used by OHDSI. 6 of the mapping mismatches (1,2%) are due to mapping target codes that do not exist on the SNOMED CT release used by SNOMED. 10 of the mapping mismatches (2,1%) are due to what could be considered as plain mapping errors by OHDSI, since the proposed SNOMED target codes are invalid.

Conclusion

We focused on mappings between two ontologies that are widely used in healthcare, derived from the two of the most relevant resources at the same time of availability. Neglecting the provenance and versions of codes implies that both mappings could, at times, be used arbitrarily within a complex data pipeline involving multiple actors. Consequently, mismatches between mappings serve as indicators of data quality issues, revealing hidden semantic inconsistencies within the resulting dataset. Despite our conservative approach focused on 1:1 mappings, we found that the number of mismatches (30%) is significant. This complements anecdotal evidence of some cases found in previous works. Of these 30%, about a third can be reconducted to different levels of abstraction that are reasonable, in that the two sources of mappings address different levels of abstraction (clinical vs population studies). A surprising 4% of mismatches are due to update or removal of codes, that is a surprising number given the two sources of mappings are synchronous in terms of release, while they present a discrepancy in terms of ontology versions of only a few months. The number of errors (10) is also surprising, given the very conservative focus on simple, widely used mappings. These findings suggest that the lack of provenance and versioning of mappings can have a significant impact (even in the order of 20%), on the quality of downstream data in integration processes. When conducting studies using Real-World Data, accurate and consistent data mapping is essential for ensuring reliability of study outcomes, which directly impacts drug safety and efficacy assessments. More disciplined and robust mapping will enhance data quality, and the overall quality of evidence generated, leading to better-informed decisions and improved patient outcomes. In future works, we intend to inspect possible root causes of such mismatches, as well as extend the analysis to other mappings.

最新情報や機会を逃さないで

DIAのメールを購読すれば、常に最新の業界情報やイベント情報を得ることができます。