Title: Context-aware sensitive information detection in unstructured text using BERT

Authors: Longjam Velentina Devi; Navanath Saharia

Addresses: Department of Computer Science and Engineering, Indian Institute of Information Technology, Manipur, India ' Department of Computer Science and Engineering, Indian Institute of Information Technology, Manipur, India

Abstract: This paper proposes a BERT-based token classification model specifically designed for detecting sensitive information inside unstructured text data. Unlike existing sentence level or context level classification methods, our methodology allows for finer granularity by examining individual words as well as their impact on one another. Previously available methods often misclassify non-sensitive information such as customer care or toll-free numbers as sensitive data, but our methodology effectively distinguish between genuinely sensitive information and non-sensitive public contact numbers, improving precision and reducing false positives. Our model surpasses previous techniques on key criteria, with an accuracy of 98%, precision of 0.98 and F1-score of 0.99 performing better than the existing model by a considerable margin. This study highlights the effectiveness of classification for sensitive data detection and establishes a new benchmark for token-level analysis contributing to more secure and effective sensitive content management.

Keywords: bidirectional encoder representations from transformers; BERT; privacy; machine learning; security; privacy; deep learning; CRF; sensitive information; personally identifiable information; PII.

DOI: 10.1504/IJICS.2026.154854

International Journal of Information and Computer Security, 2026 Vol.30 No.3, pp.291 - 307

Received: 03 Mar 2025
Accepted: 27 Aug 2025

Published online: 16 Jul 2026 *

Full-text access for editors Full-text access for subscribers Purchase this article Comment on this article