Presentation Order and Structured Diagnosis in a Synthetic Payment Workflow Benchmark

Authors

  • Ayush Ojha Georgia Institute of Technology, Atlanta, Georgia, United States

DOI:

https://doi.org/10.54536/ajise.v5i3.8542

Keywords:

Financial Technology, Local Language Model, Payment Workflow, Root-Cause Analysis, Synthetic Benchmark

Abstract

We compare three orderings of identical synthetic payment-event records across 96 cases, four local model builds, two seeds, and 2,304 calls under a fixed 256-token output limit. On 80 labeled cases, trace ordering increased exact diagnosis from 187/640 (29.22%) to 248/640 (38.75%), a 9.53-percentage-point gain. Safe abstention on 16 underdetermined cases rose from 21/128 (16.41%) to 23/128 (17.97%), satisfying the registered five-percentage-point relative margin and numerical joint gate. However, three builds never safely abstained, and structurally valid commitments increased from 50/128 to 68/128. All 803 length-terminated outputs were invalid; 47 of the net 61 additional correct outputs crossed the validity boundary. Post-hoc structural diagnostics identified fixed templates, a visible marker separating every underdetermined case, and decoy timestamps that place all distractors last under trace ordering. The protocol’s intended exhaustive two-alternative ambiguity construction was not established, although all 16 observations remained non-identifying. These results demonstrate end-to-end presentation sensitivity within the recorded suite and response budget. They do not isolate temporal reasoning, establish general ambiguity recognition, or demonstrate operational safety.

Downloads

Download data is not yet available.

Author Biography

  • Ayush Ojha, Georgia Institute of Technology, Atlanta, Georgia, United States

    Student, Georgia Institute of Technology.

References

Bendinelli, T., Dox, A., & Holz, C. (2026). TraceBench: Controlled evaluation of LLM agents for time-series root-cause attribution [Preprint]. arXiv. https://arxiv.org/abs/2608.27182

Fang, A., Yang, Y., Shang, J., Lu, Q., Xu, J., Wang, R., Zhang, S., Zhang, Y., Yu, B., & He, P. (2026). OpenRCA 2.0: From outcome labels to causal process supervision [Preprint]. arXiv. https://arxiv.org/abs/2606.27154

Ghawate, P. (2026). FinRCA-Bench: Benchmarking evidence retrieval and reasoning for financial AI systems [Preprint]. arXiv. https://arxiv.org/abs/2608.18534

Gong, A., Choi, K., Agarwal, A., Schechner, J., Huang, R., Agrawal, R., Agarwal, A., & Dwivedi, R. (2026). ORCA-bench: How ready are language model agents for oncall? [Preprint]. arXiv. https://arxiv.org/abs/2607.28545v2

Gopal, A., & Krishnan, A. (2026). How far can root cause analysis go on real-world telemetry data? [Preprint]. arXiv. https://arxiv.org/abs/2607.13548

Han, Y., Lan, M., & Kilicoglu, H. (2026). When evidence conflicts: Uncertainty and order effects in retrieval-augmented biomedical question answering. In D. Demner-Fushman, S. Ananiadou, K. Roberts, & J. Tsujii (Eds.), BioNLP 2026 (pp. 630–643). Association for Computational Linguistics. https://doi.org/10.18653/v1/2026.bionlp-1.50

Kim, T., Park, W., Yun, H., & Lee, K. (2026). Why do AI agents systematically fail at cloud root cause analysis? [Preprint]. arXiv. https://arxiv.org/abs/2602.09937

Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2024). Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12, 157–173. https://doi.org/10.1162/tacl_a_00638

Lu, Q., Fang, A., Xu, J., Shang, J., Zhang, S., Yang, Y., Yan, X., & He, P. (2026). Beyond fault localization: A trajectory-level study of LLM agents for microservice root cause analysis [Preprint]. arXiv. https://arxiv.org/abs/2608.21310

Mohammed, Y., & Lund, B. (2026). Cultural-historical activity theory and AI: Innovating and optimizing financial data retrieval. American Journal of Financial Technology and Innovation, 4(1), 75–89. https://doi.org/10.54536/ajfti.v4i1.4187

OpenTelemetry. (2026, January 14). Traces. https://opentelemetry.io/docs/concepts/signals/traces/

Presacan, O., Grama, A., Irimină, L., Nik, A., Ojha, J., Thambawita, V., Băcilă, C. I., Ionescu, B., & Riegler, M. A. (2026). Ask before you diagnose: Safe-Psych, a sequential evaluation benchmark for LLMs in psychiatry [Preprint]. arXiv. https://arxiv.org/abs/2607.13036

Riddell, E., Riddell, J., Sun, G., Antkiewicz, M., & Czarnecki, K. (2026). Stalled, biased, and confused: Uncovering reasoning failures in LLMs for cloud-based root cause analysis. In Proceedings of the 2026 IEEE/ACM Third International Conference on AI Foundation Models and Software Engineering (pp. 172–183). Association for Computing Machinery. https://doi.org/10.1145/3793655.3793732

Serifat, O. A., Igah, R. C., Balogun, K. M., Mensah, G. R., & Odai, E. N. (2025). AI-driven fraud detection in digital banking: ML approach for secure and transparent financial transactions. American Journal of Financial Technology and Innovation, 3(1), 177–187. https://doi.org/10.54536/ajfti.v3i1.5168

Xu, R., Li, J., Chen, P., & Xie, Z. (2026). TELLER: Non-intrusive cross-layer root-cause analysis for LLM inference [Preprint]. arXiv. https://arxiv.org/abs/2608.01975

Downloads

Published

2026-09-21

How to Cite

Ojha, A. . (2026). Presentation Order and Structured Diagnosis in a Synthetic Payment Workflow Benchmark. American Journal of Innovation in Science and Engineering , 5(3), 19-30. https://doi.org/10.54536/ajise.v5i3.8542

Similar Articles

1-10 of 83

You may also start an advanced similarity search for this article.