Tag: Text Anonymization

SIESTA Concepts #5: Named Entity Recognition
Sharing cyber incident reports aids threat detection, but they contain personal data that must be anonymised. Since the Named Entity Recognition models behind anonymisation are usually trained in English, EOSC SIESTA researchers tested them on Spanish and found that multilingual models trained on target-language reports work best.

SIESTA Tool: Text Anonymization on Sensitive Data
EOSC-SIESTA has developed a text anonymization tool prototype to securely share sensitive cyber incident reports. Created by Universidad de León, the tool uses AI and four anonymization techniques to mask confidential data, enabling safe data reuse for research, machine learning, and collaborative cyber defence.





