Developing a Portable Natural Language Processing Based Phenotyping System

  • 2018-07-17 19:40:28
  • Himanshu Sharma, Chengsheng Mao, Yizhen Zhang, Haleh Vatani, Liang Yao, Yizhen Zhong, Luke Rasmussen, Guoqian Jiang, Jyotishman Pathak, Yuan Luo
  • 5

Abstract

This paper presents a portable phenotyping system that is capable ofintegrating both rule-based and statistical machine learning based approaches.Our system utilizes UMLS to extract clinically relevant features from theunstructured text and then facilitates portability across differentinstitutions and data systems by incorporating OHDSI's OMOP Common Data Model(CDM) to standardize necessary data elements. Our system can also store the keycomponents of rule-based systems (e.g., regular expression matches) in theformat of OMOP CDM, thus enabling the reuse, adaptation and extension of manyexisting rule-based clinical NLP systems. We experimented with our system onthe corpus from i2b2's Obesity Challenge as a pilot study. Our systemfacilitates portable phenotyping of obesity and its 15 comorbidities based onthe unstructured patient discharge summaries, while achieving a performancethat often ranked among the top 10 of the challenge participants. Thisstandardization enables a consistent application of numerous rule-based andmachine learning based classification techniques downstream.

 

Quick Read (beta)

loading the full paper ...