Skip to main page content
U.S. flag

An official website of the United States government

Dot gov

The .gov means it’s official.
Federal government websites often end in .gov or .mil. Before sharing sensitive information, make sure you’re on a federal government site.

Https

The site is secure.
The https:// ensures that you are connecting to the official website and that any information you provide is encrypted and transmitted securely.

Access keys NCBI Homepage MyNCBI Homepage Main Content Main Navigation
. 2025 Aug 7:329:1808-1809.
doi: 10.3233/SHTI251225.

Using ChatGPT to Structure Phenotypical Entities from Disease Texts

Affiliations

Using ChatGPT to Structure Phenotypical Entities from Disease Texts

Marcel Luis Palacios-Jaén et al. Stud Health Technol Inform. .

Abstract

Despite the great advancements in biomedical text mining, accurate extraction of phenotypical entities from unstructured texts is still a critical issue. This paper explores the extraction of phenotypical data from disease texts using ChatGPT 3.5 Turbo with various prompt engineering techniques. One hundred Wikipedia disease texts have been processed with multiple prompts to extract and structure phenotypical manifestations. The results compared with UMLS MetaMap, using metrics as precision, recall, and F1 score. Simpler prompts focused on extracting "Phenotypical Manifestations" performed best, achieving an F1 score of 0.60, though still below the MetaMap benchmark. In contrast, complex prompts using entity dictionaries or explanation steps led to lower precision and recall. Including the keyword "manifestations" improved recall, while techniques like one-shot examples and prompt chaining had little effect. Overall, straightforward prompts were shown to be more effective in extracting phenotypical data entities.

Keywords: ChatGPT; Large Language Models (LLMs); Phenotypical Entities.

PubMed Disclaimer

LinkOut - more resources