Presentation Order and Structured Diagnosis in a Synthetic Payment Workflow Benchmark
DOI:
https://doi.org/10.54536/ajise.v5i3.8542Keywords:
Financial Technology, Local Language Model, Payment Workflow, Root-Cause Analysis, Synthetic BenchmarkAbstract
We compare three orderings of identical synthetic payment-event records across 96 cases, four local model builds, two seeds, and 2,304 calls under a fixed 256-token output limit. On 80 labeled cases, trace ordering increased exact diagnosis from 187/640 (29.22%) to 248/640 (38.75%), a 9.53-percentage-point gain. Safe abstention on 16 underdetermined cases rose from 21/128 (16.41%) to 23/128 (17.97%), satisfying the registered five-percentage-point relative margin and numerical joint gate. However, three builds never safely abstained, and structurally valid commitments increased from 50/128 to 68/128. All 803 length-terminated outputs were invalid; 47 of the net 61 additional correct outputs crossed the validity boundary. Post-hoc structural diagnostics identified fixed templates, a visible marker separating every underdetermined case, and decoy timestamps that place all distractors last under trace ordering. The protocol’s intended exhaustive two-alternative ambiguity construction was not established, although all 16 observations remained non-identifying. These results demonstrate end-to-end presentation sensitivity within the recorded suite and response budget. They do not isolate temporal reasoning, establish general ambiguity recognition, or demonstrate operational safety.
Downloads
References
Bendinelli, T., Dox, A., & Holz, C. (2026). TraceBench: Controlled evaluation of LLM agents for time-series root-cause attribution [Preprint]. arXiv. https://arxiv.org/abs/2608.27182
Fang, A., Yang, Y., Shang, J., Lu, Q., Xu, J., Wang, R., Zhang, S., Zhang, Y., Yu, B., & He, P. (2026). OpenRCA 2.0: From outcome labels to causal process supervision [Preprint]. arXiv. https://arxiv.org/abs/2606.27154
Ghawate, P. (2026). FinRCA-Bench: Benchmarking evidence retrieval and reasoning for financial AI systems [Preprint]. arXiv. https://arxiv.org/abs/2608.18534
Gong, A., Choi, K., Agarwal, A., Schechner, J., Huang, R., Agrawal, R., Agarwal, A., & Dwivedi, R. (2026). ORCA-bench: How ready are language model agents for oncall? [Preprint]. arXiv. https://arxiv.org/abs/2607.28545v2
Gopal, A., & Krishnan, A. (2026). How far can root cause analysis go on real-world telemetry data? [Preprint]. arXiv. https://arxiv.org/abs/2607.13548
Han, Y., Lan, M., & Kilicoglu, H. (2026). When evidence conflicts: Uncertainty and order effects in retrieval-augmented biomedical question answering. In D. Demner-Fushman, S. Ananiadou, K. Roberts, & J. Tsujii (Eds.), BioNLP 2026 (pp. 630–643). Association for Computational Linguistics. https://doi.org/10.18653/v1/2026.bionlp-1.50
Kim, T., Park, W., Yun, H., & Lee, K. (2026). Why do AI agents systematically fail at cloud root cause analysis? [Preprint]. arXiv. https://arxiv.org/abs/2602.09937
Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2024). Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12, 157–173. https://doi.org/10.1162/tacl_a_00638
Lu, Q., Fang, A., Xu, J., Shang, J., Zhang, S., Yang, Y., Yan, X., & He, P. (2026). Beyond fault localization: A trajectory-level study of LLM agents for microservice root cause analysis [Preprint]. arXiv. https://arxiv.org/abs/2608.21310
Mohammed, Y., & Lund, B. (2026). Cultural-historical activity theory and AI: Innovating and optimizing financial data retrieval. American Journal of Financial Technology and Innovation, 4(1), 75–89. https://doi.org/10.54536/ajfti.v4i1.4187
OpenTelemetry. (2026, January 14). Traces. https://opentelemetry.io/docs/concepts/signals/traces/
Presacan, O., Grama, A., Irimină, L., Nik, A., Ojha, J., Thambawita, V., Băcilă, C. I., Ionescu, B., & Riegler, M. A. (2026). Ask before you diagnose: Safe-Psych, a sequential evaluation benchmark for LLMs in psychiatry [Preprint]. arXiv. https://arxiv.org/abs/2607.13036
Riddell, E., Riddell, J., Sun, G., Antkiewicz, M., & Czarnecki, K. (2026). Stalled, biased, and confused: Uncovering reasoning failures in LLMs for cloud-based root cause analysis. In Proceedings of the 2026 IEEE/ACM Third International Conference on AI Foundation Models and Software Engineering (pp. 172–183). Association for Computing Machinery. https://doi.org/10.1145/3793655.3793732
Serifat, O. A., Igah, R. C., Balogun, K. M., Mensah, G. R., & Odai, E. N. (2025). AI-driven fraud detection in digital banking: ML approach for secure and transparent financial transactions. American Journal of Financial Technology and Innovation, 3(1), 177–187. https://doi.org/10.54536/ajfti.v3i1.5168
Xu, R., Li, J., Chen, P., & Xie, Z. (2026). TELLER: Non-intrusive cross-layer root-cause analysis for LLM inference [Preprint]. arXiv. https://arxiv.org/abs/2608.01975
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Ayush Ojha

This work is licensed under a Creative Commons Attribution 4.0 International License.