Abstract
Background: Although LOINC is widely adopted, clinical laboratory test names are frequently recorded in non-standardized free-formats, often involving mixed Persian-English terminology, in Iran. This study presents an intelligent service for automated mapping of mixed-language clinical laboratory test names to LOINC.
Methods: A dataset of 1,400 local laboratory test names and codes was extracted from a hospital information system. All records were manually mapped to LOINC by two experts as the gold standard. The data were split into training (70%) and testing (30%) sets. After data preprocessing, an N-Gram-based similarity approach with adjustable thresholds was developed for automated mapping. System outputs were evaluated against expert-assigned LOINC codes.
Results: At low similarity thresholds (<0.4), the system achieved near-complete coverage but low accuracy. Performance improved substantially with higher thresholds. At similarity thresholds between 0.7 and 0.8, approximately 78% of records were correctly mapped automatically. At a threshold of 0.8, agreement with the expert coder reached 0.84, with both precision and sensitivity of 0.86.Conclusion: The proposed approach demonstrates that accurate automated LOINC mapping is feasible for Persian-English mixed-language laboratory data. This work addresses a critical gap in multilingual laboratory data interoperability in real-world health information systems.