Architectural Foundations for Latency-Aware Scalability in AI-Enhanced Enterprise Systems

Authors

  • Daniil Sergeevich MARTYNOV Chief Software Solutions Architect, Digital Performance Labs LLC, Kazan, Russia

DOI:

https://doi.org/10.54536/ajiri.v5i3.8311

Keywords:

AI-Enhanced Systems, Enterprise Scalability, Latency-Aware Architecture, Zero-Trust Security

Abstract

Enterprise software now routes access-control decisions and code-level performance tuning through machine learning components sitting directly inside the request path. The literature has not kept pace. Distributed-systems scholarship on scalability and the newer body of work on AI-driven architecture have grown along separate tracks, and few sources spell out how the two concerns should share a single latency budget. Point solutions attack individual pieces of the problem, and the problem as a whole goes unaddressed. Static policy engines evaluate rules they did not generate from behavioral evidence. Canary-analysis platforms decide whether to promote a change without asking what that change should contain. Compiler-level optimizers raise throughput without governing how a proposed transformation earns the right to reach production traffic. This paper examines two production-documented architectures that fold detection, decision, verification, and staged deployment into one closed loop, one built for continuous security enforcement and the other for continuous performance optimization, and asks whether the integration amounts to a genuine architectural advance over the fragmented tooling it responds to. Drawing on architectural decompositions, three enterprise deployment case studies, and a controlled ablation experiment, the analysis identifies a governance pattern shared by both systems: every learned proposal answers to deterministic verification and staged production exposure before it reaches a live user. The pattern is then checked against established findings on tail latency, consistency-availability trade-offs, and microservice decomposition overhead.

Downloads

Download data is not yet available.

References

Ahmed, M. (2020). Introducing policy as code: The Open Policy Agent (OPA). Cloud Native Computing Foundation Blog.

Armbrust, M., Fox, A., Griffith, R., Joseph, A. D., Katz, R., Konwinski, A., Lee, G., Patterson, D., Rabkin, A., Stoica, I., & Zaharia, M. (2010). A view of cloud computing. Communications of the ACM, 53(4), 50–58. https://doi.org/10.1145/1721654.1721672

Blinowski, G., Ojdowska, A., & Przybylek, A. (2022). Monolithic vs. microservice architecture: A performance and scalability evaluation. IEEE Access, 10, 20357–20374. https://doi.org/10.1109/ACCESS.2022.3152803

Chen, T., Moreau, T., Jiang, Z., Zheng, L., Yan, E., Shen, H., Cowan, M., Wang, L., Hu, Y., Ceze, L., Guestrin, C., & Krishnamurthy, A. (2018). TVM: An automated end-to-end optimizing compiler for deep learning. In Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18) (pp. 578–594).

Crankshaw, D., Wang, X., Zhou, G., Franklin, M. J., Gonzalez, J. E., & Stoica, I. (2017). Clipper: A low-latency online prediction serving system. In Proceedings of the 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17) (pp. 613–627).

Dean, J., & Barroso, L. A. (2013). The tail at scale. Communications of the ACM, 56(2), 74–80. https://doi.org/10.1145/2408776.2408794

Faustino, D., Goncalves, N., Portela, M., & Silva, A. R. (2022). Stepwise migration of a monolith to a microservices architecture: Performance and migration effort evaluation. arXiv. https://doi.org/10.48550/arXiv.2201.07226

Gilbert, S., & Lynch, N. (2002). Brewer’s conjecture and the feasibility of consistent, available, partition-tolerant web services. ACM SIGACT News, 33(2), 51–59. https://doi.org/10.1145/564585.564601

Graff, M., & Sanden, C. (2018). Automated canary analysis at Netflix with Kayenta. Netflix Technology Blog.

Gujarati, A., Karimi, R., Alzayat, S., Hao, W., Kaufmann, A., Vigfusson, Y., & Mace, J. (2020). Serving DNNs like clockwork: Performance predictability from the bottom up. In Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20) (pp. 443–462).

Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., & Stoica, I. (2023). Efficient memory management for large language model serving with PagedAttention. In Proceedings of the 29th Symposium on Operating Systems Principles (SOSP ‘23) (pp. 611–626). https://doi.org/10.1145/3600006.3613165

Li, W., Lemieux, Y., Gao, J., Zhao, Z., & Han, Y. (2019). Service mesh: Challenges, state of the art, and future research opportunities. In Proceedings of the 2019 IEEE International Conference on Service-Oriented System Engineering (SOSE) (pp. 122–125). https://doi.org/10.1109/SOSE.2019.00026

Oiun, D. (2026a). AI-enhanced software architecture for modern applications. LAP LAMBERT Academic Publishing.

Oiun, D. (2026b). Neural code optimization as a new paradigm in software engineering. In Proceedings of the XXII International Scientific and Practical Conference “Topical Issues of Modern Scientific Research.”

Oiun, D. (2026c). Neural optimization and performance engineering of enterprise software. LAP LAMBERT Academic Publishing.

Panchenko, M., Auler, R., Nell, B., & Ottoni, G. (2018). BOLT: A practical binary optimizer for data centers and beyond. arXiv. https://doi.org/10.48550/arXiv.1807.06735

Shi, W., Cao, J., Zhang, Q., Li, Y., & Xu, L. (2016). Edge computing: Vision and challenges. IEEE Internet of Things Journal, 3(5), 637–646. https://doi.org/10.1109/JIOT.2016.2579198

Ward, R., & Beyer, B. (2014). BeyondCorp: A new approach to enterprise security. Login: The Usenix Magazine, 39(6), 6–11.

Yang, C., Zhao, Z., Xie, Z., Li, H., & Zhang, L. (2025). KNighter: Transforming static analysis with LLM-synthesized checkers. arXiv. https://doi.org/10.48550/arXiv.2503.09002

Downloads

Published

2026-09-16

How to Cite

MARTYNOV, D. S. (2026). Architectural Foundations for Latency-Aware Scalability in AI-Enhanced Enterprise Systems. American Journal of Interdisciplinary Research and Innovation, 5(3), 183-188. https://doi.org/10.54536/ajiri.v5i3.8311

Similar Articles

41-50 of 64

You may also start an advanced similarity search for this article.