MOSAIC: Differentially private representation of density-based clustering results

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

While the expanded role of Artificial Intelligence necessitates utilizing Big Data, ensuring the privacy of individuals remains paramount. A key difficulty lies in balancing the utility of data analysis methods, such as clustering, with the fundamental requirements of privacy preservation. In particular, density-based clustering can handle arbitrary dataset shapes and detect noise (unlike centroid-based methods), but releasing its outcomes while preserving data privacy is especially demanding. In privacy preservation, the non-interactive setting presents unique challenges compared to interactive settings. In the latter, results are released after query-based interactions, while in the former, results require stronger privacy guarantees from the output itself. This paper proposes MOSAIC, a differentially private representation designed for the non-interactive release of density-based clustering results. For MOSAIC generation, a data space is first divided into a grid, which facilitates the differentially private publication of a noise-infused, grid-based histogram of the clustering result. To enhance interpretability, a representative cluster within each grid cell is then selected. Furthermore, we propose two refinement algorithms utilizing weaving and convolution to improve MOSAIC quality. Comprehensive experimental evaluation on real and synthetic datasets demonstrates the performance of our method.

키워드

Density-based clustering; Privacy preservation; Differential privacy; Non-interactive setting; Data privacy
제목
MOSAIC: Differentially private representation of density-based clustering results
저자
Kim, Namil; Baek, Incheol; Shim, Changbeom; Chung, Yon Dohn
DOI
10.1016/j.ins.2026.123524
발행일
2026-09-05
유형
Article
저널명
Information Sciences
권
749