Stateful Mediation and Selective Auditing for Financial Language-Model Actions: A Paired Controlled Simulator Study
DOI:
https://doi.org/10.54536/ajfti.v4i1.8543Keywords:
Condition-Blind Proposal, Financial Agent, Post-State Harm, Selective Audit, Stateful MediationAbstract
Language-model agents can propose financial actions based on observations that become stale before execution. This study measures how five execution arrangements translate the same model proposals into post-state harm and correct completion in a controlled synthetic financial workflow. A protocol was internally frozen before generation. Four fixed quantized local model builds and two seeds converted 240 visible contexts into 1,920 context-level proposals. Mapping those proposals to 480 paired episode variants produced 3,840 proposal-by-episode units; replay through five arms yielded 19,200 transitions. State-bound checking produced 24 harmful transitions among 3,840 units, compared with 1,126 under direct execution (paired risk difference −0.2870; separate one-sided 95% generator-sensitivity bounds −0.2901 to −0.2836). On hazardous interpositions, its paired difference from snapshot checking was −0.8167; on paired benign interpositions, its correct-completion difference from coarse full-state checking was +0.8167. Those contrasts are algebraic sign mirrors over shared contexts, not independent confirmations. State-bound checking retained 81.67% eligible completion but did not eliminate harms visible at observation. All three contrast directions were structurally constrained by the nested mediators; the resampling bounds address nonzero magnitude under the synthetic generator, not discovery of an empirically contestable sign. In 60,000 selective-audit trajectories, no policy crossed either prespecified stopping threshold by 480 observations; all threshold times were censored at 481. Binding execution to current control state changed consequences in this fixed suite, while the conservative audit procedure did not yield an operational stopping certificate.
References
Chen, H.-H. (2026). Insuring every action: An authority frontier framework for runtime actuarial control of autonomous AI agents [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2605.25632
Chen, Z., Chen, J., Chen, J., & Sra, M. (2025). Standard benchmarks fail—Auditing LLM agents in finance must prioritize risk [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2502.15865
Chen, Z., Chen, J., Chen, J., & Sra, M. (2026). From tasks to teams: A risk-first evaluation framework for multi-agent LLM systems in finance. In M. Liakata, V. P. Moreira, J. Zhang, & D. Jurgens (Eds.), Findings of the Association for Computational Linguistics: ACL 2026 (pp. 38819–38857). Association for Computational Linguistics. https://doi.org/10.18653/v1/2026.findings-acl.1934
Debenedetti, E., Zhang, J., Balunovic, M., Beurer-Kellner, L., Fischer, M., & Tramèr, F. (2024). AgentDojo: A dynamic environment to evaluate prompt injection attacks and defenses for LLM agents. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, & C. Zhang (Eds.), Advances in neural information processing systems (Vol. 37, pp. 82895–82920). Curran Associates. https://doi.org/10.52202/079017-2636
Feng, Y., Lin, R., Wen, M., He, Q., Guo, Y., Ding, Y., Wu, Y., Chen, J., Xu, Z., Du, X., Ma, J., Chen, Z., Ma, X., Chen, Y., & Deng, X. (2026). Safety testing LLM agents at scale: From risk discovery to evidence-grounded verification [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2607.01793
Horvitz, D. G., & Thompson, D. J. (1952). A generalization of sampling without replacement from a finite universe. Journal of the American Statistical Association, 47(260), 663–685. https://doi.org/10.1080/01621459.1952.10483446
Hou, Y., Jiang, Y., Xie, Y., Yang, J., Zhang, L., Huang, H., Chen, G., & Chen, Y. (2026). FinSafetyBench: Evaluating LLM safety in real-world financial scenarios. In M. Liakata, V. P. Moreira, J. Zhang, & D. Jurgens (Eds.), Findings of the Association for Computational Linguistics: ACL 2026 (pp. 14181–14208). Association for Computational Linguistics. https://doi.org/10.18653/v1/2026.findings-acl.694
Howard, S. R., Ramdas, A., McAuliffe, J., & Sekhon, J. (2021). Time-uniform, nonparametric, nonasymptotic confidence sequences. The Annals of Statistics, 49(2), 1055–1080. https://doi.org/10.1214/20-AOS1991
Huang, D., Chua, J. K., & Wang, Z. (2026). Beyond task success: Measuring workflow fidelity in LLM-based agentic payment systems [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2605.06457
Li, J., & Zhu, J. (2026). Auditing self-evolution in financial agents: Capability gains, security drift, and execution-interface mismatch [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2608.17684
Lu, J., Wang, K., Wang, Y., Tang, Q., Zeng, H., Chen, X., Pi, J., Deng, S., Chen, L., Fu, Y., Yang, K., & Sun, X. (2026). FinToolBench: Evaluating LLM agents for real-world financial tool use [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2603.08262
Lu, Z., Zuo, Y., Xu, H., Chen, W., He, X., Guo, J., & Jin, S. (2026). WirelessOpsAgent: A benchmark and agent design for action assurance in wireless networks [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2608.08277
Luo, M., Chen, C., Cao, A., Huang, Z., & Dai, W. (2026). STAGE: Stateful translation to agentic graph execution with policy-scoped context and deterministic control [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2608.22538
Luo, Y., Jiang, Y., Xie, Q., Lan, L., Cong, L. W., Rao, A., & Song, Y. (2026). ReguSim: Evaluating LLM agent rule grounding in financial compliance [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2608.19974
Mao, Q., Wang, J., Liu, Y., Zhu, L., Ma, C., & Yan, J. (2026). SoK: Security of autonomous LLM agents in agentic commerce [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2604.15367
Meng, F., Du, L., Wu, Z., Chen, G., Liu, X., Liao, J., Jiang, C., Wan, Z., Gu, J., Zhou, P., Huang, R., Zhao, Z., Ding, S., Yu, A., Peng, B., Xia, B., Sun, H., Liang, H., Xie, J., . . . Shieh, M. Q. (2026). ClawMark: A living-world benchmark for multi-turn, multi-day, multimodal coworker agents [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2604.23781
Mohammed, Y., & Lund, B. (2026). Cultural-historical activity theory and AI: Innovating and optimizing financial data retrieval. American Journal of Financial Technology and Innovation, 4(1), 75–89. https://doi.org/10.54536/ajfti.v4i1.4187
Mou, Y., Xue, Z., Li, L., Liu, P., Zhang, S., Ye, W., & Shao, J. (2026). ToolSafe: Enhancing tool invocation safety of LLM-based agents via proactive step-level guardrail and feedback. In M. Liakata, V. P. Moreira, J. Zhang, & D. Jurgens (Eds.), Findings of the Association for Computational Linguistics: ACL 2026 (pp. 37125–37153). Association for Computational Linguistics. https://doi.org/10.18653/v1/2026.findings-acl.1850
Pai, D. M., & Xian, L. (2026). FraudBench: Stress-testing policy-grounded banking agents against adaptive fraud [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2608.18136
Peng, Y., & Wu, X. (2026). Stateful governance for concurrent agentic systems [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2608.02764
Plaat, A., van Duijn, M., van Stein, N., Preuss, M., van der Putten, P., & Batenburg, K. J. (2025). Agentic large language models, a survey. Journal of Artificial Intelligence Research, 84, Article 29. https://doi.org/10.1613/jair.1.18675
Raidah, F. Z., Jobair, M., Halimuzzaman, M., Sharma, J., & Ahmed, S. S. (2026). Artificial intelligence and the transformation of internal audit functions. American Journal of Financial Technology and Innovation, 4(1), 178–184. https://doi.org/10.54536/ajfti.v4i1.8001
Ruan, Y., Dong, H., Wang, A., Pitis, S., Zhou, Y., Ba, J., Dubois, Y., Maddison, C. J., & Hashimoto, T. (2024). Identifying the risks of LM agents with an LM-emulated sandbox. In International Conference on Learning Representations. https://proceedings.iclr.cc/paper_files/paper/2024/hash/7274ed909a312d4d869cc328ad1c5f04-Abstract-Conference.html
Santos-Grueiro, I. (2026). Temporary authority, permanent effects: Commit-time authorization for LLM agents [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2607.10487
Shekhar, S., Xu, Z., Lipton, Z., Liang, P., & Ramdas, A. (2023). Risk-limiting financial audits via weighted sampling without replacement. In R. J. Evans & I. Shpitser (Eds.), Proceedings of the Thirty-Ninth Conference on Uncertainty in Artificial Intelligence (Proceedings of Machine Learning Research, Vol. 216, pp. 1932–1941). PMLR. https://proceedings.mlr.press/v216/shekhar23a.html
Shi, T., Mo, Y., Liu, Y., Hao, Z., Wang, Y., Hu, W., Yu, N., Zhou, M., & Yu, J. (2026). Organizational control layer: Governance infrastructure at the execution boundary of LLM agent systems [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2606.04306
Tang, R., Liu, Q., Zhang, Y., Yang, Y., Chen, X., & Dong, C. (2026). Context is not authority: Structured runtime governance for financial market agents [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2608.09025
Tong, Y., Dai, L., & Guo, S. (2026). AID-Guard: Stateful authorization for delegated agent effects [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2608.21159
Wang, H., Poskitt, C. M., & Sun, J. (2025). AgentSpec: Customizable runtime enforcement for safe and reliable LLM agents [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2503.18666
Xia, H., Wang, H., Liu, Z., Yu, Q., Guo, Y., & Wang, H. (2025). SafeToolBench: Pioneering a prospective benchmark to evaluating tool utilization safety in LLMs. In C. Christodoulopoulos, T. Chakraborty, C. Rose, & V. Peng (Eds.), Findings of the Association for Computational Linguistics: EMNLP 2025 (pp. 17643–17660). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.findings-emnlp.958
Xu, J., Fan, L., Wang, Z., Li, X., & Liu, H. (2026). Beyond single-use tokens: Durable authorization state for replay-resistant LLM agent actions [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2608.01710
Yang, Z., Li, R., Qiang, Q., Wang, J., Lou, F., Li, M., Cheng, D., Xu, R., Lian, H., Zhang, S., Liang, X., Huang, X., Wei, Z., Liu, Z., Guo, X., Wang, H., Chen, R., & Zhang, L. (2026). FinVault: Benchmarking financial agent safety in execution-grounded environments [Preprint, version 1]. arXiv. https://arxiv.org/abs/2601.07853v1
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2023). ReAct: Synergizing reasoning and acting in language models. In International Conference on Learning Representations. https://openreview.net/forum?id=WE_vluYUL-X
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Ayush Ojha

This work is licensed under a Creative Commons Attribution 4.0 International License.