. 2020 Aug 28;20(1):204.

doi: 10.1186/s12911-020-01216-9.

PASCAL: a pseudo cascade learning framework for breast cancer treatment entity normalization in Chinese clinical text

Yang An¹, Jianlin Wang², Liang Zhang³, Hanyu Zhao⁴, Zhan Gao⁵, Haitao Huang⁶, Zhenguang Du⁶, Zengtao Jiao⁷, Jun Yan⁷, Xiaopeng Wei¹, Bo Jin⁸

Affiliations

¹ School of Computer Science and Technology, Dalian University of Technology, No.2 Linggong Road, Ganjingzi District, Dalian, Liaoning, 116024, China.
² First Hospital of Lanzhou University, 1 Donggang W Rd, Chengguan District, Lanzhou, Gansu, 730000, China.
³ International Bussiness College, Dongbei University of Finance and Economics, No.20 Jianshan Street, Shahekou District, Dalian, Liaoning, 116025, China. liang.zhang@dufe.edu.cn.
⁴ Dalian University, No.10 Xuefu Street, Economic and Technological Development Zone, Dalian, Liaoning, 116622, China.
⁵ BeiJing Haoyisheng Cloud Hospital Management Technology Ltd., No.10 Dewai Street, Xicheng District, Beijing, 100088, China.
⁶ The People's Hospital of Liaoning Province, No.33 Shenhe District, Shenyang, Liaoning, 110016, China.
⁷ AI Lab, Yidu Cloud, No.35 of Huayuan North Road, Haidian District, Beijing, 100191, China.
⁸ School of Innovation and Entrepreneurship, Dalian University of Technology, No.2 Linggong Road, Ganjingzi District, Dalian, Liaoning, 116024, China.

PMID: 32859189
PMCID: PMC7456389
DOI: 10.1186/s12911-020-01216-9

PASCAL: a pseudo cascade learning framework for breast cancer treatment entity normalization in Chinese clinical text

Yang An et al. BMC Med Inform Decis Mak. 2020.

. 2020 Aug 28;20(1):204.

doi: 10.1186/s12911-020-01216-9.

Authors

Yang An¹, Jianlin Wang², Liang Zhang³, Hanyu Zhao⁴, Zhan Gao⁵, Haitao Huang⁶, Zhenguang Du⁶, Zengtao Jiao⁷, Jun Yan⁷, Xiaopeng Wei¹, Bo Jin⁸

Affiliations

¹ School of Computer Science and Technology, Dalian University of Technology, No.2 Linggong Road, Ganjingzi District, Dalian, Liaoning, 116024, China.
² First Hospital of Lanzhou University, 1 Donggang W Rd, Chengguan District, Lanzhou, Gansu, 730000, China.
³ International Bussiness College, Dongbei University of Finance and Economics, No.20 Jianshan Street, Shahekou District, Dalian, Liaoning, 116025, China. liang.zhang@dufe.edu.cn.
⁴ Dalian University, No.10 Xuefu Street, Economic and Technological Development Zone, Dalian, Liaoning, 116622, China.
⁵ BeiJing Haoyisheng Cloud Hospital Management Technology Ltd., No.10 Dewai Street, Xicheng District, Beijing, 100088, China.
⁶ The People's Hospital of Liaoning Province, No.33 Shenhe District, Shenyang, Liaoning, 110016, China.
⁷ AI Lab, Yidu Cloud, No.35 of Huayuan North Road, Haidian District, Beijing, 100191, China.
⁸ School of Innovation and Entrepreneurship, Dalian University of Technology, No.2 Linggong Road, Ganjingzi District, Dalian, Liaoning, 116024, China.

PMID: 32859189
PMCID: PMC7456389
DOI: 10.1186/s12911-020-01216-9

Abstract

Backgrounds: Knowledge discovery from breast cancer treatment records has promoted downstream clinical studies such as careflow mining and therapy analysis. However, the clinical treatment text from electronic health data might be recorded by different doctors under their hospital guidelines, making the final data rich in author- and domain-specific idiosyncrasies. Therefore, breast cancer treatment entity normalization becomes an essential task for the above downstream clinical studies. The latest studies have demonstrated the superiority of deep learning methods in named entity normalization tasks. Fundamentally, most existing approaches adopt pipeline implementations that treat it as an independent process after named entity recognition, which can propagate errors to later tasks. In addition, despite its importance in clinical and translational research, few studies directly deal with the normalization task in Chinese clinical text due to the complexity of composition forms.

Methods: To address these issues, we propose PASCAL, an end-to-end and accurate framework for breast cancer treatment entity normalization (TEN). PASCAL leverages a gated convolutional neural network to obtain a representation vector that can capture contextual features and long-term dependencies. Additionally, it treats treatment entity recognition (TER) as an auxiliary task that can provide meaningful information to the primary TEN task and as a particular regularization to further optimize the shared parameters. Finally, by concatenating the context-aware vector and probabilistic distribution vector from TEN, we utilize the conditional random field layer (CRF) to model the normalization sequence and predict the TEN sequential results.

Results: To evaluate the effectiveness of the proposed framework, we employ the three latest sequential models as baselines and build the model in single- and multitask on a real-world database. Experimental results show that our method achieves better accuracy and efficiency than state-of-the-art approaches.

Conclusions: The effectiveness and efficiency of the presented pseudo cascade learning framework were validated for breast cancer treatment normalization in clinical text. We believe the predominant performance lies in its ability to extract valuable information from unstructured text data, which will significantly contribute to downstream tasks, such as treatment recommendations, breast cancer staging and careflow mining.

Keywords: Breast cancer; Cascade learning; Chinese clinical text mining; Treatment entity normalization.

PubMed Disclaimer

Conflict of interest statement

The authors declare that they have no competing interests.

Figures

**Fig. 1**
Illustration of clinical text, normalization examples and possible applications. a EHR data; b Clinical text from EHRs : an example of real clinical text and translated version; c Real-world data and standard entity; d Applied scenarios

**Fig. 2**
Main architecture of PASCAL model. PASCAL consists of four modules: character embedding module, encoder module (containing a gated convolutional neural network to learn the shared representation with temporal relationship), pseudo cascade structure module (including the enhanced primary task TEN and an auxiliary task TER)

**Fig. 3**
Detailed structure of encoder module: gated convolutional neural network (GCNN). GCNN consists of three key parts: convolutional block, gating block and residual connection

**Fig. 4**
Accuracy comparison between PASCAL and Feedback [17]

**Fig. 5**
Computational efficiency comparison of *PASCAL* with different encoders

See this image and copyright information in PMC

References

1. Marklund L, Hammarstedt L. Impact of hpv in oropharyngeal cancer. J Oncol. 2011;2011(1687-8450):509036. - PMC - PubMed
1. What Is Breast Cancer?https://www.imaginis.com/general-information-on-breast-cancer/what-is-br.... Accessed 11 June 2008.
1. Dagliati A, Sacchi L, Zambelli A, Tibollo V, Pavesi L, Holmes JH, Bellazzi R. Temporal electronic phenotyping by mining careflows of breast cancer patients. J Biomed Inform; 66:136–47. 10.1016/j.jbi.2016.12.012. - PubMed
1. Yadav R, Khan Z, Saxena H. Chemotherapy prediction of cancer patient by using data mining techniques. Int J Comput Appl. 2014;76(10):28–31.
1. Wang XH, Zheng B, Good WF, King JL, Chang Y-H. Computer-assisted diagnosis of breast cancer using a data-driven bayesian belief network. Int J Med Inform; 54(2):115–26. 10.1016/S1386-5056(98)00174-9. - PubMed

Publication types

Actions

MeSH terms

Actions
Actions
Actions
Actions
Actions
Actions
Actions

LinkOut - more resources

Full Text Sources
Medical
- MedlinePlus Health Information

Save citation to file

Email citation

Add to Collections

Add to My Bibliography

Your saved search

Create a file for external citation management software

Your RSS Feed

PASCAL: a pseudo cascade learning framework for breast cancer treatment entity normalization in Chinese clinical text

Affiliations

PASCAL: a pseudo cascade learning framework for breast cancer treatment entity normalization in Chinese clinical text

Authors

Affiliations

Abstract

Conflict of interest statement

Figures

References

Publication types

MeSH terms

LinkOut - more resources

Full Text Sources

Medical