lab2clean: a novel algorithm for automated cleaning of retrospective clinical laboratory results data for secondary uses
- PMID: 39227951
- PMCID: PMC11370074
- DOI: 10.1186/s12911-024-02652-7
lab2clean: a novel algorithm for automated cleaning of retrospective clinical laboratory results data for secondary uses
Abstract
Background: The integrity of clinical research and machine learning models in healthcare heavily relies on the quality of underlying clinical laboratory data. However, the preprocessing of this data to ensure its reliability and accuracy remains a significant challenge due to variations in data recording and reporting standards.
Methods: We developed lab2clean, a novel algorithm aimed at automating and standardizing the cleaning of retrospective clinical laboratory results data. lab2clean was implemented as two R functions specifically designed to enhance data conformance and plausibility by standardizing result formats and validating result values. The functionality and performance of the algorithm were evaluated using two extensive electronic medical record (EMR) databases, encompassing various clinical settings.
Results: lab2clean effectively reduced the variability of laboratory results and identified potentially erroneous records. Upon deployment, it demonstrated effective and fast standardization and validation of substantial laboratory data records. The evaluation highlighted significant improvements in the conformance and plausibility of lab results, confirming the algorithm's efficacy in handling large-scale data sets.
Conclusions: lab2clean addresses the challenge of preprocessing and cleaning clinical laboratory data, a critical step in ensuring high-quality data for research outcomes. It offers a straightforward, efficient tool for researchers, improving the quality of clinical laboratory data, a major portion of healthcare data. Thereby, enhancing the reliability and reproducibility of clinical research outcomes and clinical machine learning models. Future developments aim to broaden its functionality and accessibility, solidifying its vital role in healthcare data management.
Keywords: Algorithms; Clinical laboratories; Data integrity; Data preprocessing; Electronic medical records.
© 2024. The Author(s).
Conflict of interest statement
The authors declare no competing interests.
Figures
References
-
- Garbage. in, garbage out. In: Wikipedia. 2023. Available from: https://en.wikipedia.org/wiki/Garbage_in,_garbage_out. Cited 2024 Feb 12.
-
- Kandel S, Heer J, Plaisant C, Kennedy J, van Ham F, Riche NH, et al. Research directions in data wrangling: visualizations and transformations for usable and credible data. Inform Visual. 2011;10(4):271–88.
MeSH terms
LinkOut - more resources
Full Text Sources
Miscellaneous
