![]()
Certificate: View Certificate
Published Paper PDF: Download PDF
Confirmation Letter: View
DOI: https://doi.org/10.63345/ijrhs.net.v10.i1.1
Sarvesh Kumar Gupta
Consulting Member of Technical Staff
Oracle
Saint Peters, Missouri -63376 , USA
ORCID: 0009-0008-7460-4874
Abstract— Cloud-native data warehouses increasingly rely on columnar storage to support large-scale analytical workloads, but the performance benefits of columnar design depend on the optimization techniques applied at the storage and query-processing layers. This paper presents a conceptual benchmarking and comparative analysis of major columnar storage optimization techniques used in cloud-native warehouse environments. The study examines compression, dictionary encoding, run-length encoding, delta encoding, partitioning, clustering, predicate pushdown, metadata indexing, and data pruning in relation to query latency, storage efficiency, scan reduction, throughput, and resource utilization. Rather than claiming live execution across specific commercial platforms, the paper uses an illustrative benchmarking framework based on a generic cloud-native columnar warehouse model and representative OLAP-style analytical workloads. The analysis indicates that data pruning and predicate pushdown provide the strongest query-performance improvements by reducing unnecessary data scans, while compression and encoding techniques substantially reduce storage footprint and improve I/O efficiency. The findings suggest that no single technique is sufficient for optimal performance; instead, combined optimization strategies offer the greatest benefits for analytical query execution, storage reduction, and cost efficiency. The paper concludes that effective columnar storage optimization in cloud-native warehouses requires an integrated approach combining storage layout design, compression-aware processing, metadata-based filtering, and workload-sensitive query execution.
Keywords— Cloud-Native Data Warehouses; Columnar Storage; Storage Optimization; Data Compression; Partitioning; Query Performance; Benchmarking; Data Pruning; Cloud Analytics; Distributed Data Processing.
References
- Stonebraker, M., Abadi, D. J., Batkin, A., Chen, X., Cherniack, M., Ferreira, M., Lau, E., Lin, A., Madden, S., O’Neil, E., O’Neil, P., Rasin, A., Tran, N., & Zdonik, S. (2005). C-Store: A column-oriented DBMS. Proceedings of VLDB, 553–564.
- Abadi, D. J., Madden, S. R., & Ferreira, M. C. (2006). Integrating compression and execution in column-oriented database systems. Proceedings of SIGMOD, 671–682. https://doi.org/10.1145/1142473.1142548
- Abadi, D. J., Madden, S. R., & Hachem, N. (2008). Column-stores vs. row-stores: How different are they really? Proceedings of SIGMOD, 967–980. https://doi.org/10.1145/1376616.1376712
- Abadi, D., Boncz, P., Harizopoulos, S., Idreos, S., & Madden, S. (2013). The design and implementation of modern column-oriented database systems. Foundations and Trends in Databases, 5(3), 197–280.
- Boncz, P. A., Zukowski, M., & Nes, N. (2005). MonetDB/X100: Hyper-pipelining query execution. CIDR.
- Floratou, A., Patel, J. M., Shekita, E. J., & Tata, S. (2011). Column-oriented storage techniques for MapReduce. Proceedings of the VLDB Endowment, 4(7), 419–429. https://doi.org/10.14778/1988776.1988778
- Melnik, S., Gubarev, A., Long, J. J., Romer, G., Shivakumar, S., Tolton, M., & Vassilakis, T. (2010). Dremel: Interactive analysis of web-scale datasets. Proceedings of the VLDB Endowment, 3(1–2), 330–339. https://doi.org/10.14778/1920841.1920886
- Melnik, S., et al. (2020). Dremel: A decade of interactive SQL analysis at web scale. Proceedings of the VLDB Endowment, 13(12), 3461–3472.
- Dageville, B., Cruanes, T., Zukowski, M., Antonov, V., Avanes, A., Bock, J., Claybaugh, J., Engovatov, D., Hentschel, M., Huang, J., Lee, A. W., Motivala, A., Munir, A. Q., Pelley, S., Povinec, P., Rahn, G., Triantafyllis, S., & Unterbrunner, P. (2016). The Snowflake Elastic Data Warehouse. Proceedings of SIGMOD, 215–226. https://doi.org/10.1145/2882903.2903741
- Gupta, A., Agarwal, D., Tan, D., Kulesza, J., Pathak, R., Stefani, S., & Srinivasan, V. (2015). Amazon Redshift and the case for simpler data warehouses. Proceedings of SIGMOD, 1917–1923. https://doi.org/10.1145/2723372.2742795