Abstract
The objective of BioCreative8 Track 3 is to extract phenotypic key medicalfindings embedded within EHR texts and subsequently normalize these findings totheir Human Phenotype Ontology (HPO) terms. However, the presence of diversesurface forms in phenotypic findings makes it challenging to accuratelynormalize them to the correct HPO terms. To address this challenge, we exploredvarious models for named entity recognition and implemented data augmentationtechniques such as synonym marginalization to enhance the normalization step.Our pipeline resulted in an exact extraction and normalization F1 score 2.6\%higher than the mean score of all submissions received in response to thechallenge. Furthermore, in terms of the normalization F1 score, our approachsurpassed the average performance by 1.9\%. These findings contribute to theadvancement of automated medical data extraction and normalization techniques,showcasing potential pathways for future research and application in thebiomedical domain.