. 2022 Sep:165:104808.

doi: 10.1016/j.ijmedinf.2022.104808. Epub 2022 Jun 10.

Inaccurate recording of routinely collected data items influences identification of COVID-19 patients

Eva S Klappe¹, Ronald Cornet², Dave A Dongelmans³, Nicolette F de Keizer²

Affiliations

¹ Amsterdam UMC, University of Amsterdam, Department of Medical Informatics, Amsterdam Public Health Research Institute, Amsterdam, Netherlands. Electronic address: e.s.klappe@amsterdamumc.nl.
² Amsterdam UMC, University of Amsterdam, Department of Medical Informatics, Amsterdam Public Health Research Institute, Amsterdam, Netherlands.
³ Amsterdam UMC, University of Amsterdam, Department of Intensive Care Medicine, Amsterdam, Netherlands.

PMID: 35767912
PMCID: PMC9186787
DOI: 10.1016/j.ijmedinf.2022.104808

Inaccurate recording of routinely collected data items influences identification of COVID-19 patients

Eva S Klappe et al. Int J Med Inform. 2022 Sep.

. 2022 Sep:165:104808.

doi: 10.1016/j.ijmedinf.2022.104808. Epub 2022 Jun 10.

Authors

Eva S Klappe¹, Ronald Cornet², Dave A Dongelmans³, Nicolette F de Keizer²

Affiliations

¹ Amsterdam UMC, University of Amsterdam, Department of Medical Informatics, Amsterdam Public Health Research Institute, Amsterdam, Netherlands. Electronic address: e.s.klappe@amsterdamumc.nl.
² Amsterdam UMC, University of Amsterdam, Department of Medical Informatics, Amsterdam Public Health Research Institute, Amsterdam, Netherlands.
³ Amsterdam UMC, University of Amsterdam, Department of Intensive Care Medicine, Amsterdam, Netherlands.

PMID: 35767912
PMCID: PMC9186787
DOI: 10.1016/j.ijmedinf.2022.104808

Abstract

Background: During the Coronavirus disease 2019 (COVID-19) pandemic it became apparent that it is difficult to extract standardized Electronic Health Record (EHR) data for secondary purposes like public health decision-making. Accurate recording of, for example, standardized diagnosis codes and test results is required to identify all COVID-19 patients. This study aimed to investigate if specific combinations of routinely collected data items for COVID-19 can be used to identify an accurate set of intensive care unit (ICU)-admitted COVID-19 patients.

Methods: The following routinely collected EHR data items to identify COVID-19 patients were evaluated: positive reverse transcription polymerase chain reaction (RT-PCR) test results; problem list codes for COVID-19 registered by healthcare professionals and COVID-19 infection labels. COVID-19 codes registered by clinical coders retrospectively after discharge were also evaluated. A gold standard dataset was created by evaluating two datasets of suspected and confirmed COVID-19-patients admitted to the ICU at a Dutch university hospital between February 2020 and December 2020, of which one set was manually maintained by intensivists and one set was extracted from the EHR by a research data management department. Patients were labeled 'COVID-19' if their EHR record showed diagnosing COVID-19 during or right before an ICU-admission. Patients were labeled 'non-COVID-19' if the record indicated no COVID-19, exclusion or only suspicion during or right before an ICU-admission or if COVID-19 was diagnosed and cured during non-ICU episodes of the hospitalization in which an ICU-admission took place. Performance was determined for 37 queries including real-time and retrospective data items. We used the F₁ score, which is the harmonic mean between precision and recall. The gold standard dataset was split into one subset including admissions between February and April and one subset including admissions between May and December to determine accuracy differences.

Results: The total dataset consisted of 402 patients: 196 'COVID-19' and 206 'non-COVID-19' patients. F₁ scores of search queries including EHR data items that can be extracted real-time ranged between 0.68 and 0.97 and for search queries including the data item that was retrospectively registered by clinical coders F₁ scores ranged between 0.73 and 0.99. F₁ scores showed no clear pattern in variability between the two time periods.

Conclusions: Our study showed that one cannot rely on individual routinely collected data items such as coded COVID-19 on problem lists to identify all COVID-19 patients. If information is not required real-time, medical coding from clinical coders is most reliable. Researchers should be transparent about their methods used to extract data. To maximize the ability to completely identify all COVID-19 cases alerts for inconsistent data and policies for standardized data capture could enable reliable data reuse.

Keywords: COVID-19; Data accuracy; Electronic Health Records; Problem list; Real-time data extraction; Routinely collected data.

PubMed Disclaimer

Conflict of interest statement

This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors. This study was funded by Amsterdam UMC 2019-AMC-JK-7. Amsterdam UMC did not have any role in the study design, collection, analysis, interpretation of the data, writing the report and the decision to submit the report for publication.

Figures

**Fig. 1**
Flow chart to annotate a patient with a COVID-19 or non-COVID-19 label.

**Fig. 2**
Dataset inclusion and exclusion and final gold standard dataset (n = 402) with 196 COVID-19 labeled patients and 206 non-COVID-19 labeled patients.

**Fig. 3**
Search queries applied to the gold standard dataset (n = 402). The numbers indicate the search queries, shown in the legend.

**Fig. 4**
Search queries applied to the subset including admissions between February and April (n = 208). The numbers indicate the search queries, shown in the legend.

**Fig. 5**
Search queries applied to the subset including admissions between May and December (n = 194). The numbers indicate the search queries, shown in the legend.

See this image and copyright information in PMC

Cited by

Reusing routine electronic health record data for nationwide COVID-19 surveillance in nursing homes: barriers, facilitators, and lessons learned.
Wieland-Jorna Y, Verheij RA, Francke AL, Coppen R, de Greeff SC, Elffers A, Oosterveld-Vlug MG. Wieland-Jorna Y, et al. BMC Med Inform Decis Mak. 2024 Dec 27;24(1):408. doi: 10.1186/s12911-024-02818-3. BMC Med Inform Decis Mak. 2024. PMID: 39731119 Free PMC article.
Factors Affecting the Medical Coding Errors of COVID-19 Hospital Records: A Cross-Sectional Study in East Iran.
Garavand A, Esmaeeli E, Bahador F, Sabahi A. Garavand A, et al. Tanaffos. 2024 Feb;23(2):189-197. Tanaffos. 2024. PMID: 39959802 Free PMC article.
Utilization of Computable Phenotypes in Electronic Health Record Research: A Review and Case Study in Atopic Dermatitis.
Masison J, Lehmann HP, Wan J. Masison J, et al. J Invest Dermatol. 2025 May;145(5):1008-1016. doi: 10.1016/j.jid.2024.08.025. Epub 2024 Nov 1. J Invest Dermatol. 2025. PMID: 39488781 Review.

References

1. Reeves J.J., Hollandsworth H.M., Torriani F.J., Taplitz R., Abeles S., Tai-Seale M., et al. Rapid response to COVID-19: health informatics support for outbreak management in an academic health system. J Am Med Inform Assoc. 2020;27(6):853–859. - PMC - PubMed
1. Narayan V., Hoong P.B., Chuin S. Innovative Use of Health Informatics to Augment Contact Tracing during the COVID19 Pandemic in an Acute Hospital. J Am Med Inform Assoc. 2020 - PMC - PubMed
1. Moore J.H., Barnett I., Boland M.R., Chen Y., Demiris G., Gonzalez-Hernandez G., Herman D.S., Himes B.E., Hubbard R.A., Kim D., Morris J.S., Mowery D.L., Ritchie M.D., Shen L.i., Urbanowicz R., Holmes J.H. Ideas for how informaticians can get involved with COVID-19 research. BioData Mining. 2020;13(1) - PMC - PubMed
1. Holmgren A.J., Apathy N.C., Adler-Milstein J. Barriers to Hospital Electronic Public Health Reporting and Implications for the COVID-19 Pandemic. J Am Med Inform Assoc. 2020 - PMC - PubMed
1. Wolkewitz M., Puljak L. Methodological challenges of analysing COVID-19 data during the pandemic. BMC Med Res Methodol. 2020;20(1) - PMC - PubMed

MeSH terms

Actions
Actions
Actions
Actions
Actions
Actions
Actions

LinkOut - more resources

Full Text Sources
Medical
- MedlinePlus Health Information
Miscellaneous
- NCI CPTAC Assay Portal

Save citation to file

Email citation

Add to Collections

Add to My Bibliography

Your saved search

Create a file for external citation management software

Your RSS Feed

Inaccurate recording of routinely collected data items influences identification of COVID-19 patients

Affiliations

Inaccurate recording of routinely collected data items influences identification of COVID-19 patients

Authors

Affiliations

Abstract

Conflict of interest statement

Figures

Similar articles

Cited by

References

MeSH terms

LinkOut - more resources

Full Text Sources

Medical

Miscellaneous

Abstract

Conflict of interest statement

Figures

Similar articles

Cited by

References

MeSH terms

Related information

LinkOut - more resources

Full Text Sources

Medical

Miscellaneous