SIESTA Video Series: Text Anonymization on Sensitive Data

The SIESTA Video Series brings together a set of short videos in which several of the project’s use cases present their work and main objectives. In particular, the use cases on Medical Imaging, Energy Domain, Text Anonymization on sensitive data, and Demography have contributed to this series, each explaining their respective activities within the project. These videos are structured to provide a clear and concise explanation of each use case, allowing viewers to understand both the technical aspects and the broader context in which each activity is developed.

The series is designed as an accessible way to showcase how SIESTA addresses challenges related to sensitive data, privacy, and secure data analysis across different scientific domains. Available in video format, it allows viewers to get a clear overview of the scope and diversity of the project, while also highlighting the practical tools and approaches being developed. In addition, the format helps to communicate complex concepts in a more visual and understandable way, making it easier for a wider audience to engage with the project’s objectives. Overall, the SIESTA Video Series serves as a concise introduction to the project’s use cases and their role within the broader SIESTA ecosystem.

Text anonymization on sensitive data

The University of León presents Use Case 4 of the EOSC-SIESTA project, developed under Work Package 17, which focuses on text anonymization of sensitive data. The use case addresses the challenge faced by cyber emergency response teams, which generate large volumes of incident reports containing highly sensitive information that cannot be publicly shared, limiting the development of effective classification models. This limitation has a direct impact on the ability to train robust machine learning systems in the cybersecurity domain.

To overcome this barrier, the University of León is developing a tool capable of automatically and securely anonymizing these texts while preserving their usefulness for training machine learning models. In addition, due to the lack of access to real data, they have created a synthetic dataset of 10,000 cyber incident reports generated using advanced language models. Preliminary results based on named entity recognition techniques have shown promising performance, with further improvements expected once the dataset is fully annotated. This work aims to enable secure data sharing and support the development of next-generation cybersecurity solutions within the EOSC ecosystem.