Information Leakage in Encrypted Deduplication via Frequency Analysis: Attacks and Defenses

Li, Jingwei; Lee, Patrick P. C.; Tan, Chufeng; Qin, Chuan; Zhang, Xiaosong

Computer Science > Cryptography and Security

arXiv:1904.05736 (cs)

[Submitted on 11 Apr 2019 (v1), last revised 9 Oct 2019 (this version, v2)]

Title:Information Leakage in Encrypted Deduplication via Frequency Analysis: Attacks and Defenses

Authors:Jingwei Li, Patrick P. C. Lee, Chufeng Tan, Chuan Qin, Xiaosong Zhang

View PDF

Abstract:Encrypted deduplication combines encryption and deduplication to simultaneously achieve both data security and storage efficiency. State-of-the-art encrypted deduplication systems mainly build on deterministic encryption to preserve deduplication effectiveness. However, such deterministic encryption reveals the underlying frequency distribution of the original plaintext chunks. This allows an adversary to launch frequency analysis against the ciphertext chunks and infer the content of the original plaintext chunks. In this paper, we study how frequency analysis affects information leakage in encrypted deduplication storage, from both attack and defense perspectives. Specifically, we target backup workloads, and propose a new inference attack that exploits chunk locality to increase the coverage of inferred chunks. We further combine the new inference attack with the knowledge of chunk sizes and show its attack effectiveness against variable-size chunks. We conduct trace-driven evaluation on both real-world and synthetic datasets and show that our proposed attacks infer a significant fraction of plaintext chunks under backup workloads. To defend against frequency analysis, we present two defense approaches, namely MinHash encryption and scrambling. Our trace-driven evaluation shows that our combined MinHash encryption and scrambling scheme effectively mitigates the severity of the inference attacks, while maintaining high storage efficiency and incurring limited metadata access overhead.

Comments:	31 pages, Accepted by ACM Transactions on Storage
Subjects:	Cryptography and Security (cs.CR); Distributed, Parallel, and Cluster Computing (cs.DC)
Cite as:	arXiv:1904.05736 [cs.CR]
	(or arXiv:1904.05736v2 [cs.CR] for this version)
	https://doi.org/10.48550/arXiv.1904.05736

Submission history

From: Jingwei Li [view email]
[v1] Thu, 11 Apr 2019 14:49:34 UTC (399 KB)
[v2] Wed, 9 Oct 2019 14:07:27 UTC (734 KB)

Computer Science > Cryptography and Security

Title:Information Leakage in Encrypted Deduplication via Frequency Analysis: Attacks and Defenses

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Cryptography and Security

Title:Information Leakage in Encrypted Deduplication via Frequency Analysis: Attacks and Defenses

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators