Search arXiv⌕ Search

arXiv · 2407.08874

Implications of mappings between ICD clinical diagnosis codes and Human Phenotype Ontology terms

Abstract

Objective: Integrating EHR data with other resources is essential in rare disease research due to low disease prevalence. Such integration is dependent on the alignment of ontologies used for data annotation. The International Classification of Diseases (ICD) is used to annotate clinical diagnoses; the Human Phenotype Ontology (HPO) to annotate phenotypes. Although these ontologies overlap in biomedical entities described, the extent to which they are interoperable is unknown. We investigate how well aligned these ontologies are and whether such alignments facilitate EHR data integration. Materials and Methods: We conducted an empirical analysis of the coverage of mappings between ICD and HPO. We interpret this mapping coverage as a proxy for how easily clinical data can be integrated with research ontologies such as HPO. We quantify how exhaustively ICD codes are mapped to HPO by analyzing mappings in the UMLS Metathesaurus. We analyze the proportion of ICD codes mapped to HPO within a real-world EHR dataset. Results and Discussion: Our analysis revealed that only 2.2% of ICD codes have direct mappings to HPO in UMLS. Within our EHR dataset, less than 50% of ICD codes have mappings to HPO terms. ICD codes that are used frequently in EHR data tend to have mappings to HPO; ICD codes that represent rarer medical conditions are seldom mapped. Conclusion: We find that interoperability between ICD and HPO via UMLS is limited. While other mapping sources could be incorporated, there are no established conventions for what resources should be used to complement UMLS.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Amelia LM Tan, Rafael S Gonçalves, William Yuan, Gabriel A Brat, The Consortium for Clinical Characterization of COVID-19 by EHR, Robert Gentleman, Isaac S Kohane. 2024-07-11. Implications of mappings between ICD clinical diagnosis codes and Human Phenotype Ontology terms. https://arxiv.org/abs/2407.08874

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

MLSkip: Data Skipping for ML Filters via Lightweight Metadata

Database vendors recently released AI functions that can be used in filter predicates. As such functions often rely on costly, black-box ML models, they unveil new data management challenges. Concretely, traditional data skipping techniques for integer and string data fail to be applicable to the new filter type. Indeed, there is no known mechanism for pruning non-qualifying row groups, e.g., when reading files from blob storage. In this work, we initiate the study of data skipping techniques for ML filters. We make the case that Parquet's default min-max metadata is enough to enable pruning. To this end, we draw connections to two lines of research: (i) the recently proposed query language for ML models and (ii) neural network verification. Our preliminary results on ReLU architectures show that on tables from TPC-H and TPC-DS, the average pruning effectiveness for filters of selectivity below 0.1% amounts to 27.4%. Finally, inspired by research on spatial joins, we propose an enhanced metadata structure: a size-bounded 2D convex hull that verification tools can make better use of, increasing the pruning effectiveness to 38.31%, while occupying at most 45 bytes per row group and column pair. We observe an end-to-end speedup of 1.07$\times$ over PyTorch in DuckDB.

cs.DB↗

IBLTs Measure Before They Decode: Self-Sizing Set Reconciliation for Database Consistency Verification

Cross-system data replication pipelines cannot confirm end-to-end consistency from the local guarantees of each hop, so the two endpoints must be compared directly on a periodic basis. Once the rows of a fixed snapshot are normalized into fingerprints, the task reduces to finding the symmetric difference of the two sets. An Invertible Bloom Lookup Table (IBLT) reconciles the sets with communication that grows only with the difference cardinality $d$, independent of table size, but its capacity must be fixed while $d$ is still unknown. Across 41,603 production reconciliations over 90 days, nonzero $d$ spans about seven orders of magnitude, and no reliable empirical constant exists. We show that the count array of an IBLT has already measured $d$ before decoding. The measurement is in-band: it is carried by the recovery sketch itself and adds no bytes dedicated to estimation. A mapping-aware theorem extends the construction to Irregular, Rateless, and MET IBLTs. The protocol reads the estimate only after a decoding failure; we prove that the failure-conditioned lower quantile bounds the risk of underestimation, which gives the second-round capacity a configurable success-probability guarantee. The resulting self-sizing protocol attempts recovery with a small first-round sketch and stops on success; on failure it reads $d$, sizes the second round, and completes reconciliation in at most two rounds. Against a controlled oracle, communication is 1.29-1.47 times that of a scheme given $d$ in advance. Production workload characterization, relational-database replay, and a cross-city KV deployment confirm the end-to-end mechanism. In production on an Oracle-MySQL link, all completed runs succeeded within two rounds, 90.7% on the 1-RTT fast path with a single 16 KB sketch.

cs.DB↗

Towards Effective Orchestration of AI x DB Workloads

AI-driven analytics are increasingly crucial to data-centric decision-making. Executing relational and AI operators in separate runtimes prevents the database optimizer and runtime from coordinating operator ordering, model placement, batching, and state reuse. Integrating AI operators into database engines enables such coordination but raises challenges in jointly optimizing query processing and model execution, scheduling under resource contention, and reusing relational intermediates and AI artifacts. This paper formalizes AIxDB workloads as iterative, concurrent, and shareable executions that interleave relational and AI operators. We then advocate database-native orchestration as a paradigm for redesigning database engines for these workloads and distill two design principles: holistic AIxDB co-optimization and unified AIxDB cache management. We present NeurEngine as a proof-of-concept prototype and report preliminary results illustrating the performance benefits of database-native orchestration

cs.DB↗