Data Anonymization.pdf

430962_1_En_13_Chapter.pdf
Preview of Data Anonymization
🔗 Source: inria.hal.science
📊 Size: 181 KB
👤 Author: Yoan Miche, Ian Oliver, Silke Holtmanns, Aapo Kalliola, Anton Akusok, Amaury Lendasse, Kaj-Mikael Bj
⬇️ Downloads: 412

Summary

Data Anonymization as a Vector Quantization Problem:
Control Over Privacy for Health Data is a research paper that tackles data anonymization from a vector quantization point of view. The goal is to anonymize data to avoid single individual or group re-identification while maintaining data integrity and structure.

The paper proposes using vector quantization to capture the structure of the data through clustering, and then anonymizing the data by "pushing back outliers in the crowd" and modifying the data to blend in individuals while preserving cluster-level statistics and structure.

The approach involves considering data fields as separate entities and building clusters of records based on metric proximity. The assumption is that if a group is large enough and the records inside have been "stirred" enough, identification of a single individual becomes impossible.

The paper uses an example of health-related information to illustrate the concept, where relationships between non-sensitive fields can make it easy to identify individuals. The proposed approach seeks to "blend in" individuals with the rest of the records and make the data anonymized by modifying the data values to realistic ones while preserving some of the information.

The high-level motivation for data privacy is to distort the data in a known manner to increase privacy or entropy, without changing the format of the data or the underlying space. The paper aims to propose practical solutions to the identification problem by defining notations and assumptions on the structure of the data and tackling the problem of moving single samples or groups back into bigger clusters.

Key points:
- Data anonymization through vector quantization
- Clustering to capture data structure
- Anonymizing data by "pushing back outliers" and modifying data to blend in individuals
- Preserving cluster-level statistics and structure
- Example of health-related information to illustrate the concept
- High-level motivation for data privacy to distort data without changing format or underlying space.

Description

Data anonymization is viewed as a vector quantization problem. It controls privacy for health data. Published under Creative Commons Attribution 4.0 International License.

Technical Information

  • File Format: PDF
  • File Size: 181 KB
  • Pages: 11
  • Language: EN
  • Author: Yoan Miche, Ian Oliver, Silke Holtmanns, Aapo Kalliola, Anton Akusok, Amaury Lendasse, Kaj-Mikael Bj
  • Total Downloads: 412
  • Last Updated: 2 hours ago

Document Overview

This PDF document about Data Anonymization provides comprehensive information and guidance. Whether you're a beginner or advanced user, this resource offers valuable insights into Data Anonymization.

Related Topics

If you're interested in Data Anonymization, you might also want to explore:

Download Data Anonymization eBooks for free and learn more about Data Anonymization. These books contain exercises and tutorials to improve your practical skills, at all levels!

Not satisfied with this document? We have related documents to Data Anonymization, try searching with similar keywords: Data Anonymization, Data Anonymization?page=2, Anonymization Eazydoc Com, Lightning Anonymization, Anonymization, Anonymization Pdf, Anonymization Tutorial, Anonymization Reference

You can download PDF versions of the user's guide, manuals and ebooks about Data Anonymization, you can also find and download for free A free online manual (notices) with beginner and intermediate, Downloads Documentation, You can download PDF files (or DOC and PPT) about Data Anonymization for free, but please respect copyrighted ebooks.