A Trustworthy and Scalable AI Infrastructure for Large-Scale E-Commerce Systems
DOI:
https://doi.org/10.54536/ajise.v5i2.6461Keywords:
Data integrity, E-commerce, Large language models, Machine learning, Trustworthy AIAbstract
Large-scale e-commerce platforms generate vast amounts of data that demand reliable, real-time analytics and intelligent decision-making. We propose a trustworthy and scalable AI infrastructure that integrates big data pipelines, machine learning models, and large language models (LLMs) within a secure data management framework to enhance operational efficiency and data integrity in e-commerce. The system architecture combines a distributed streaming data pipeline with robust data validation, state-of-the-art ML algorithms for prediction and anomaly detection, and LLM-driven services for unstructured data analysis, all under strict access control and privacy safeguards. Open-source e-commerce datasets were used to simulate high-volume transactions and user interactions. Experimental results under heavy load (up to tens of thousands of events per second) demonstrate that the proposed architecture maintains reliable performance and near-perfect data integrity (>98% valid data) even as throughput increases. Additionally, sensitive customer information is protected through encryption and differential privacy without significant impact on latency. The outcomes show that our architecture can uphold consistency, privacy, and trustworthy AI behavior at scale. This work provides a practical foundation for deploying AI-driven services in e-commerce settings that require both scalability and trustworthiness, bridging the gap between big data processing and responsible AI deployment.
Downloads
References
Alrumiah, S., & Hadwan, M. (2021). Implementing big data analytics in e-commerce: Vendor and customer view. IEEE Access, 9, 37281–37286. https://doi.org/10.1109/ACCESS.2021.3063517
Amershi, S., Begel, A., Bird, C., Zimmermann, T., Nurvitadhi, E., et al. (2019). Software engineering for machine learning: A case study. In Proceedings of the 41st International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP) (pp. 291–300). IEEE. https://doi.org/10.1109/ICSE-SEIP.2019.00042
Atluri, R. P., Taylor, A., & Prudhvi, R. (2025). Data contracts in the wild: An approach for redefining trust and accountability in modern data pipelines. World Journal of Advanced Research and Reviews, 26(3), 290–299. https://doi.org/10.30574/wjarr.2025.26.3.0410
Baylor, D., Koc, L., Koo, C. Y., Jain, V., Schleier-Smith, J., et al. (2017). TFX: A TensorFlow-based production-scale machine learning platform. In Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 1387–1395). ACM. https://doi.org/10.1145/3097983.3098021
Bollikonda, M., & Bollikonda, T. (2025). Secure pipelines, smarter AI: LLM-powered data engineering for threat detection and compliance. Preprints, 2025041365. https://doi.org/10.20944/preprints202504.1365.v1
Carlini, N., Liu, C., Erlingsson, Ú., Kos, J., & Song, D. (2019). The secret sharer: Evaluating and testing unintended memorization in neural networks. In Proceedings of the 28th USENIX Security Symposium (pp. 267–284). USENIX Association.
Davenport, T. H., Guha, A., Grewal, D., & Bressgott, T. (2020). How artificial intelligence will change the future of marketing. Journal of the Academy of Marketing Science, 48(1), 24–42. https://doi.org/10.1007/s11747-019-00696-0
Desai, P., & Ganatra, K. (2022). Artificial intelligence in strengthening the operations of e-commerce-based business. In Proceedings of the 2022 International Conference on Data Analytics for Business and Industry (ICDABI) (pp. 275–280). IEEE. https://doi.org/10.1109/ICDABI56852.2022.10080483
Dwork, C., & Roth, A. (2014). The algorithmic foundations of differential privacy. Foundations and Trends in Theoretical Computer Science, 9(3–4), 211–407. https://doi.org/10.1561/0400000042
European Commission. (2019). Ethics guidelines for trustworthy AI. Publications Office of the European Union. https://ec.europa.eu/newsroom/dae/document.cfm?doc_id=60419
Feretzakis, G., & Verykios, V. S. (2024). Trustworthy AI: Securing sensitive data in large language models. AI, 5(4), 2773–2800. https://doi.org/10.3390/ai5040134
Kairouz, P., McMahan, H. B., Avent, B., Bellet, A., Bennis, M., et al. (2019). Advances and open problems in federated learning (arXiv:1912.04977). arXiv. https://arxiv.org/abs/1912.04977
Munshi, A., Alhindi, A., Qadah, T. M., & Alqurashi, A. (2023). An electronic commerce big data analytics architecture and platform. Applied Sciences, 13(19), Article 10962. https://doi.org/10.3390/app131910962
Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., et al. (2015). Hidden technical debt in machine learning systems. In Advances in Neural Information Processing Systems (Vol. 28, pp. 2503–2511).
Sura, R. (2025). Scalable AI-powered data pipelines for enterprise analytics. International Journal of Innovative Research in Technology, 12(1), 1094–1101.
Xin, Q. (2025a). A deep reinforcement learning approach to optimizing cloud workload migration. American Journal of Interdisciplinary Research and Innovation, 4(3), 10–15.
Xin, Q. (2025b). Hybrid cloud architecture for efficient and cost-effective large language model deployment. Journal of Information Systems and Informatics, 7(3), 2182–2195.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Daren Zheng, Boning Zhang, Xinzhuo Sun

This work is licensed under a Creative Commons Attribution 4.0 International License.