EFFICIENT AND PRIVACY-PRESERVING BIG DATA STORAGE AND COMPUTATION: A MATHEMATICAL AND HYPOTHESIS-DRIVEN APPROACH
Loading...
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
International Journal of Applied Mathematics
Abstract
This work addresses the long-standing tension between privacy guarantees and system performance in big-data storage and computation. We present a mathematical, hypothesis-driven framework that formalizes the mapping from input parameters—dataset size DD, compression ratio CrC_r, encryption strength EsE_s, privacy budget ϵ\epsilon, and resources RR—to output metrics—storage efficiency SeS_e, computation time TcT_c, utility QQ, and privacy guarantee PP. Closed-form relations for SeS_e, TcT_c, and QQ enable testable hypotheses, parameter estimation, and reproducible evaluation across workloads. The framework integrates compressed encryption for storage, differential privacy and homomorphic encryption for computation, and fine-grained access control, providing a unified basis for reasoning about privacy–performance trade-offs. Analytical validation demonstrates 68.8% storage efficiency, ≈21% reduction in computation time relative to AES-128 + Gzip, and ≤1.8% utility loss while satisfying ϵ≤1.0\epsilon \le 1.0 and 256-bit security. These results indicate that privacy preservation need not be at odds with performance when design choices are guided by a unified model. Eventually, we framed the issue as a multi-objective optimization, revealing a Pareto frontier over privacy, utility, storage efficiency, and latency, and allowing for automatic tuning in different deployment contexts. The proposed formulation provides a system-agnostic, reproducible foundation for designing, analyzing, and improving privacy-preserving big-data systems.
