arXiv · 2108.07971
De-identification of Unstructured Clinical Texts from Sequence to Sequence Perspective
Abstract
In this work, we propose a novel problem formulation for de-identification of unstructured clinical text. We formulate the de-identification problem as a sequence to sequence learning problem instead of a token classification problem. Our approach is inspired by the recent state-of -the-art performance of sequence to sequence learning models for named entity recognition. Early experimentation of our proposed approach achieved 98.91% recall rate on i2b2 dataset. This performance is comparable to current state-of-the-art models for unstructured clinical text de-identification.
Explore related subjects
Keep this discovery
Md Monowar Anjum, Noman Mohammed, Xiaoqian Jiang. 2021-08-18. De-identification of Unstructured Clinical Texts from Sequence to Sequence Perspective. https://doi.org/10.1145/3460120.3485354
Cite the original work for its findings. Save a collection to share your selection of sources.