Typed Field Support with a Value-Only Deterministic Executor for Synthetic Financial Records: A Multi-Seed Controlled Evaluation

Authors

  • Ayush Ojha 595 Pacific Ave, San Francisco, California, United States

DOI:

https://doi.org/10.54536/ajfti.v4i1.8551

Keywords:

Financial Reasoning, Language Model, Numerical Reliability, Selective Prediction, Tool-Augmented Reasoning

Abstract

We evaluated whether providing deterministic numerical values to a language model improves supported financial answers without increasing unsupported answer attempts. The internally frozen design crossed 96 synthetic financial bases, four paired evidence conditions, four fixed local model builds, three decoding seeds and three interfaces. All 13,824 scheduled calls completed, yielding 55,296 field decisions. The two principal interfaces shared the same typed support certificate and post-generation gate; only one received a value-only projection from a separately implemented Decimal executor. Relative to TRACE alone, TRACE+Tool increased raw unsupported attempts by 10.1 percentage points (one-sided 97.5% upper bound, 12.2), supported coverage by 16.0 points (lower bound, 14.0), and sound completeness by 86.6 points (lower bound, 85.2). Bounds used stratified resampling of the 96 logical bases. The primary intersection-union test failed because the unsupported-attempt safety margin was exceeded. The broader positive-tool and near-oracle claims also failed. Both TRACE interfaces prevented unsupported emissions by construction, which does not erase the increase in raw attempts. The result identifies a trade-off in this fixed synthetic suite: deterministic value provision improved retained numerical completion while increasing unsupported attempted answers. It does not establish reliability on real financial records or other model populations.

Author Biography

  • Ayush Ojha, 595 Pacific Ave, San Francisco, California, United States

    Student at the Georgia Institute of Technology.

References

Abdallah, A., Elmadany, A. A., Al Natour, S., Cavusoglu, H., Jatowt, A., & Abdul-Mageed, M. (2026). MoCA-Agent: A market-of-claims code agent for financial and numerical reasoning [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2606.11537

Baldazzi, T., Bellomarini, L., Coletta, A., Iezzi, M., Maple, C., Pesare, A., & Sallinger, E. (2026). VADAOrchestra: Neurosymbolic orchestration of adaptive reasoning workflows. In R. Wassermann, M.-L. Mugnier, & F. Baader (Eds.), Proceedings of the 23rd International Conference on Principles of Knowledge Representation and Reasoning (pp. 815–825). IJCAI Organization. https://doi.org/10.24963/kr.2026/77

Berger, R. L., & Hsu, J. C. (1996). Bioequivalence trials, intersection-union tests and equivalence confidence sets. Statistical Science, 11(4), 283–319. https://doi.org/10.1214/ss/1032280304

Buchmann, J., Liu, X., & Gurevych, I. (2024). Attribute or abstain: large language models as long document assistants. In Y. Al-Onaizan, M. Bansal, & Y.-N. Chen (Eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (pp. 8113–8140). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.emnlp-main.463

Cai, X., Guo, M., Liu, J., Han, J., Guo, B., Long, Y., Chen, Y., Wu, B., Metaxas, D. N., & Li, R. (2026). OpenPM: Auditable point-in-time evaluation for LLM portfolio-management agents [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2608.09988

Chen, H., Tang, X., Liu, Q., Shi, W., Li, S., Lyu, F., Luo, W., Du, X., & He, X. (2026). Fighting numerical hallucinations via data-centric compilation for online financial QA. In Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 (pp. 7082–7093). Association for Computing Machinery. https://doi.org/10.1145/3770855.3818407

Chen, W., Ma, X., Wang, X., & Cohen, W. W. (2023). Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks. Transactions on Machine Learning Research. https://openreview.net/forum?id=YfZ4ZPt8zd

Davison, A. C., & Hinkley, D. V. (1997). Bootstrap methods and their application. Cambridge University Press. https://doi.org/10.1017/CBO9780511802843

Dong, Y., Zhao, E., & Chen, E. (2026). Gold-guided programmatic distillation for financial reasoning over hybrid tables and text [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2607.14709

Efron, B. (1979). Bootstrap methods: another look at the jackknife. The Annals of Statistics, 7(1), 1–26. https://doi.org/10.1214/aos/1176344552

Faria, F. T. J., Moin, M. B., Mahmud, J. A., Mridha, M. F., & Hossain, M. A. (2026). CLAIR-Fin: An adversarial multi-agent framework for claim-level verification and adaptive debate in cross-modal financial QA [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2608.13706

Ferguson, N., Pennington, J., Beghian, N., Mohan, A., Kiela, D., Agrawal, S., & Nguyen, T. H. (2026). ExtractBench: A benchmark and evaluation methodology for complex structured extraction [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2602.12247

Franzmeyer, T., Sravankumar, A., Liu, L., Mao, Y., Hou, R., Wang, S., Foerster, J. N., Zettlemoyer, L., & Khabsa, M. (2026). High accuracy, less talk (HALT): reliable LLMs through capability-aligned finetuning. In The Fourteenth International Conference on Learning Representations. https://openreview.net/forum?id=LYqBnNVaXD

Gao, L., Madaan, A., Zhou, S., Alon, U., Liu, P., Yang, Y., Callan, J., & Neubig, G. (2023). PAL: program aided language models. In A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, & J. Scarlett (Eds.), Proceedings of the 40th International Conference on Machine Learning (Vol. 202, pp. 10764–10799). PMLR. https://proceedings.mlr.press/v202/gao23f.html

Ghawate, P. (2026). FinRCA-Bench: Benchmarking evidence retrieval and reasoning for financial AI systems [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2608.18534

Gu, F., Ren, X., Jiang, Z., Zhang, Z., García-Fernández, Á. F., Stefanidis, A., Zhou, M., Li, H., & Su, J. (2026). EvidenceLens: A claim-evidence matrix for auditing financial question answering [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2606.23724

Holm, S. (1979). A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics, 6(2), 65–70. https://www.jstor.org/stable/4615733

Hu, M., Hu, S., Wang, B., Sa, Y., Wang, X., Guo, X., Zha, D., & Xiao, J. (2026). ECPO: Evidence-coupled policy optimization for evidence-certified candidate ranking [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2605.21993

Huang, T., Xu, S., Dang, J. T., Yan, S., & Yin, K. (2026). Answer only as precisely as justified: Calibrated claim-level specificity control for agentic systems [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2604.17487

Kan, S. (2026). Claim-selective certification for high-risk medical retrieval-augmented generation [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2605.21949

Khatchadourian, R. (2026). Replayable financial agents: A determinism-faithfulness assurance harness for tool-using LLM agents [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2601.15322

Kim, H. J., Kim, Y., Lee, S.-G., & Kim, T. (2025). When to speak, when to abstain: contrastive decoding with abstention. In W. Che, J. Nabende, E. Shutova, & M. T. Pilehvar (Eds.), Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 9710–9730). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.acl-long.479

Lee, Y., Kim, S., Kwak, Y., & Choo, J. (2026). BankMathBench: a benchmark for numerical reasoning in banking scenarios. In S. Piperidis, N. Bel, H. van den Heuvel, N. Ide, S. Krek, & A. Toral (Eds.), Proceedings of the Fifteenth Language Resources and Evaluation Conference (pp. 11010–11027). ELRA Language Resource Association. https://doi.org/10.63317/3uxnd7yxsmsb

Lu, J., Wang, K., Wang, Y., Tang, Q., Zeng, H., Chen, X., Pi, J., Deng, S., Chen, L., Fu, Y., Yang, K., & Sun, X. (2026). FinToolBench: Evaluating LLM agents for real-world financial tool use [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2603.08262

Min, B., Edemacu, K., Cho, S.-H., Choi, Y., Jang, B., & Kim, J. W. (2026). When absence is evidence: Evaluating completeness-sensitive negative reasoning in large language models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2608.04591

Mohammed, Y., & Lund, B. (2026). Cultural-historical activity theory and AI: innovating and optimizing financial data retrieval. American Journal of Financial Technology and Innovation, 4(1), 75–89. https://doi.org/10.54536/ajfti.v4i1.4187

Ogunsusi, K. L., Ayisi, R. K. K., Offei, A. N. A., Giami, K. L., & Ogunsusi, K. B. (2026). Generative AI and advanced analytics for financial modeling, valuation and strategic decision-making. American Journal of Applied Statistics and Economics, 5(1), 97–112. https://doi.org/10.54536/ajase.v5i1.6986

Parekh, K., Tiwari, A. K., & Saxena, D. (2026). CIFQA: A deterministic tool-grounded multi-agent LLM framework for financial query answering [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2608.26114

Usman, R. M. (2026). PhantomFill: When the form demands an answer, language models invent one [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2607.20492

Wang, Q. (2026). AWARE-FX: An auditable knowledge-guided AI system for measuring corporate foreign-exchange hedging disclosure [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2607.27611

Wang, Y., Ai, X., Patel, J., Peng, X., Mo, F., Cao, Y., Li, H., Cao, M., Qian, L., & Gutiérrez-Basulto, V. (2026). AUDITFLOW: Executable symbolic environments for structured financial reporting verification [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2606.03031

Zhong, S., Zhu, J., Xu, Q., Sun, L., Wang, Y., Sun, Q., Chen, S., & Zhang, T. (2026). FinRiskAtlas: Decision-aligned evaluation of large language models for financial risk review [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2608.25325

Downloads

Published

2026-09-30

How to Cite

Ojha, A. . (2026). Typed Field Support with a Value-Only Deterministic Executor for Synthetic Financial Records: A Multi-Seed Controlled Evaluation. American Journal of Financial Technology and Innovation, 4(1), 185-197. https://doi.org/10.54536/ajfti.v4i1.8551

Similar Articles

11-20 of 41

You may also start an advanced similarity search for this article.