Typed Field Support with a Value-Only Deterministic Executor for Synthetic Financial Records: A Multi-Seed Controlled Evaluation
DOI:
https://doi.org/10.54536/ajfti.v4i1.8551Keywords:
Financial Reasoning, Language Model, Numerical Reliability, Selective Prediction, Tool-Augmented ReasoningAbstract
We evaluated whether providing deterministic numerical values to a language model improves supported financial answers without increasing unsupported answer attempts. The internally frozen design crossed 96 synthetic financial bases, four paired evidence conditions, four fixed local model builds, three decoding seeds and three interfaces. All 13,824 scheduled calls completed, yielding 55,296 field decisions. The two principal interfaces shared the same typed support certificate and post-generation gate; only one received a value-only projection from a separately implemented Decimal executor. Relative to TRACE alone, TRACE+Tool increased raw unsupported attempts by 10.1 percentage points (one-sided 97.5% upper bound, 12.2), supported coverage by 16.0 points (lower bound, 14.0), and sound completeness by 86.6 points (lower bound, 85.2). Bounds used stratified resampling of the 96 logical bases. The primary intersection-union test failed because the unsupported-attempt safety margin was exceeded. The broader positive-tool and near-oracle claims also failed. Both TRACE interfaces prevented unsupported emissions by construction, which does not erase the increase in raw attempts. The result identifies a trade-off in this fixed synthetic suite: deterministic value provision improved retained numerical completion while increasing unsupported attempted answers. It does not establish reliability on real financial records or other model populations.
References
Abdallah, A., Elmadany, A. A., Al Natour, S., Cavusoglu, H., Jatowt, A., & Abdul-Mageed, M. (2026). MoCA-Agent: A market-of-claims code agent for financial and numerical reasoning [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2606.11537
Baldazzi, T., Bellomarini, L., Coletta, A., Iezzi, M., Maple, C., Pesare, A., & Sallinger, E. (2026). VADAOrchestra: Neurosymbolic orchestration of adaptive reasoning workflows. In R. Wassermann, M.-L. Mugnier, & F. Baader (Eds.), Proceedings of the 23rd International Conference on Principles of Knowledge Representation and Reasoning (pp. 815–825). IJCAI Organization. https://doi.org/10.24963/kr.2026/77
Berger, R. L., & Hsu, J. C. (1996). Bioequivalence trials, intersection-union tests and equivalence confidence sets. Statistical Science, 11(4), 283–319. https://doi.org/10.1214/ss/1032280304
Buchmann, J., Liu, X., & Gurevych, I. (2024). Attribute or abstain: large language models as long document assistants. In Y. Al-Onaizan, M. Bansal, & Y.-N. Chen (Eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (pp. 8113–8140). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.emnlp-main.463
Cai, X., Guo, M., Liu, J., Han, J., Guo, B., Long, Y., Chen, Y., Wu, B., Metaxas, D. N., & Li, R. (2026). OpenPM: Auditable point-in-time evaluation for LLM portfolio-management agents [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2608.09988
Chen, H., Tang, X., Liu, Q., Shi, W., Li, S., Lyu, F., Luo, W., Du, X., & He, X. (2026). Fighting numerical hallucinations via data-centric compilation for online financial QA. In Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 (pp. 7082–7093). Association for Computing Machinery. https://doi.org/10.1145/3770855.3818407
Chen, W., Ma, X., Wang, X., & Cohen, W. W. (2023). Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks. Transactions on Machine Learning Research. https://openreview.net/forum?id=YfZ4ZPt8zd
Davison, A. C., & Hinkley, D. V. (1997). Bootstrap methods and their application. Cambridge University Press. https://doi.org/10.1017/CBO9780511802843
Dong, Y., Zhao, E., & Chen, E. (2026). Gold-guided programmatic distillation for financial reasoning over hybrid tables and text [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2607.14709
Efron, B. (1979). Bootstrap methods: another look at the jackknife. The Annals of Statistics, 7(1), 1–26. https://doi.org/10.1214/aos/1176344552
Faria, F. T. J., Moin, M. B., Mahmud, J. A., Mridha, M. F., & Hossain, M. A. (2026). CLAIR-Fin: An adversarial multi-agent framework for claim-level verification and adaptive debate in cross-modal financial QA [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2608.13706
Ferguson, N., Pennington, J., Beghian, N., Mohan, A., Kiela, D., Agrawal, S., & Nguyen, T. H. (2026). ExtractBench: A benchmark and evaluation methodology for complex structured extraction [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2602.12247
Franzmeyer, T., Sravankumar, A., Liu, L., Mao, Y., Hou, R., Wang, S., Foerster, J. N., Zettlemoyer, L., & Khabsa, M. (2026). High accuracy, less talk (HALT): reliable LLMs through capability-aligned finetuning. In The Fourteenth International Conference on Learning Representations. https://openreview.net/forum?id=LYqBnNVaXD
Gao, L., Madaan, A., Zhou, S., Alon, U., Liu, P., Yang, Y., Callan, J., & Neubig, G. (2023). PAL: program aided language models. In A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, & J. Scarlett (Eds.), Proceedings of the 40th International Conference on Machine Learning (Vol. 202, pp. 10764–10799). PMLR. https://proceedings.mlr.press/v202/gao23f.html
Ghawate, P. (2026). FinRCA-Bench: Benchmarking evidence retrieval and reasoning for financial AI systems [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2608.18534
Gu, F., Ren, X., Jiang, Z., Zhang, Z., García-Fernández, Á. F., Stefanidis, A., Zhou, M., Li, H., & Su, J. (2026). EvidenceLens: A claim-evidence matrix for auditing financial question answering [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2606.23724
Holm, S. (1979). A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics, 6(2), 65–70. https://www.jstor.org/stable/4615733
Hu, M., Hu, S., Wang, B., Sa, Y., Wang, X., Guo, X., Zha, D., & Xiao, J. (2026). ECPO: Evidence-coupled policy optimization for evidence-certified candidate ranking [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2605.21993
Huang, T., Xu, S., Dang, J. T., Yan, S., & Yin, K. (2026). Answer only as precisely as justified: Calibrated claim-level specificity control for agentic systems [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2604.17487
Kan, S. (2026). Claim-selective certification for high-risk medical retrieval-augmented generation [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2605.21949
Khatchadourian, R. (2026). Replayable financial agents: A determinism-faithfulness assurance harness for tool-using LLM agents [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2601.15322
Kim, H. J., Kim, Y., Lee, S.-G., & Kim, T. (2025). When to speak, when to abstain: contrastive decoding with abstention. In W. Che, J. Nabende, E. Shutova, & M. T. Pilehvar (Eds.), Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 9710–9730). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.acl-long.479
Lee, Y., Kim, S., Kwak, Y., & Choo, J. (2026). BankMathBench: a benchmark for numerical reasoning in banking scenarios. In S. Piperidis, N. Bel, H. van den Heuvel, N. Ide, S. Krek, & A. Toral (Eds.), Proceedings of the Fifteenth Language Resources and Evaluation Conference (pp. 11010–11027). ELRA Language Resource Association. https://doi.org/10.63317/3uxnd7yxsmsb
Lu, J., Wang, K., Wang, Y., Tang, Q., Zeng, H., Chen, X., Pi, J., Deng, S., Chen, L., Fu, Y., Yang, K., & Sun, X. (2026). FinToolBench: Evaluating LLM agents for real-world financial tool use [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2603.08262
Min, B., Edemacu, K., Cho, S.-H., Choi, Y., Jang, B., & Kim, J. W. (2026). When absence is evidence: Evaluating completeness-sensitive negative reasoning in large language models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2608.04591
Mohammed, Y., & Lund, B. (2026). Cultural-historical activity theory and AI: innovating and optimizing financial data retrieval. American Journal of Financial Technology and Innovation, 4(1), 75–89. https://doi.org/10.54536/ajfti.v4i1.4187
Ogunsusi, K. L., Ayisi, R. K. K., Offei, A. N. A., Giami, K. L., & Ogunsusi, K. B. (2026). Generative AI and advanced analytics for financial modeling, valuation and strategic decision-making. American Journal of Applied Statistics and Economics, 5(1), 97–112. https://doi.org/10.54536/ajase.v5i1.6986
Parekh, K., Tiwari, A. K., & Saxena, D. (2026). CIFQA: A deterministic tool-grounded multi-agent LLM framework for financial query answering [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2608.26114
Usman, R. M. (2026). PhantomFill: When the form demands an answer, language models invent one [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2607.20492
Wang, Q. (2026). AWARE-FX: An auditable knowledge-guided AI system for measuring corporate foreign-exchange hedging disclosure [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2607.27611
Wang, Y., Ai, X., Patel, J., Peng, X., Mo, F., Cao, Y., Li, H., Cao, M., Qian, L., & Gutiérrez-Basulto, V. (2026). AUDITFLOW: Executable symbolic environments for structured financial reporting verification [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2606.03031
Zhong, S., Zhu, J., Xu, Q., Sun, L., Wang, Y., Sun, Q., Chen, S., & Zhang, T. (2026). FinRiskAtlas: Decision-aligned evaluation of large language models for financial risk review [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2608.25325
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Ayush Ojha

This work is licensed under a Creative Commons Attribution 4.0 International License.