Category: Text Anonymization on Sensitive Data

SIESTA Concepts #5: Named Entity Recognition
Sharing cyber incident reports aids threat detection, but they contain personal data that must be anonymised. Since the Named Entity Recognition models behind anonymisation are usually trained in English, EOSC SIESTA researchers tested them on Spanish and found that multilingual models trained on target-language reports work best.

SIESTA Tool: Text Anonymization on Sensitive Data
EOSC-SIESTA has developed a text anonymization tool prototype to securely share sensitive cyber incident reports. Created by Universidad de León, the tool uses AI and four anonymization techniques to mask confidential data, enabling safe data reuse for research, machine learning, and collaborative cyber defence.

SIESTA Video Series: Text Anonymization on Sensitive Data
The SIESTA Video Series brings together a set of short videos in which several of the project’s use cases present their work and main objectives. In particular, the use cases on Medical Imaging, Energy Domain, Text Anonymization on sensitive data, and Demography have contributed to this series, each explaining their respective activities within the project.…





