Budgeted Multi-Hop Retrieval Agents with Evidence-Chain Reliability and Cost-Aware Abstention

Authors

DOI:

https://doi.org/10.60087/jklst.vol5.n2.006

Keywords:

multi-hop retrieval, retrieval-augmented generation, reciprocal rank fusion, dynamic stopping, calibration, selective prediction, abstention, evidence chains

Abstract

Fixed top-k retrieval spends the same context budget on every question without indicating whether a multi-document evidence chain is complete. This study evaluates a budgeted agent on MultiHop-RAG using all 609 documents and 2,556 questions. The 1,534/511/511 train/validation/test split kept each evidence group in one partition. The pipeline compared BM25, 256-dimensional latent-semantic retrieval, equal and weighted reciprocal rank fusion, supervised reranking, fixed depth, and calibrated dynamic stopping. A sparse reader, null detector, and isotonic calibration supported answer, evidence, citation, cost, and abstention evaluation. On 451 answerable test questions, reranking achieved 0.9069 Recall@10, 0.8064 nDCG@10, and 0.7428 complete-chain retrieval. The dynamic agent used a mean of 544.16 tokens and obtained 0.6184 F1, compared with 666.29 tokens and 0.6243 F1 for the null-gated Reranker-8 policy. Its paired F1 difference was -0.0059 with a 95% cluster-bootstrap interval of [-0.0246, 0.0115], while the 122.13-token reduction remained significant after Holm adjustment. Isotonic calibration reduced expected calibration error from 0.2671 to 0.0321. Selective answering operated at 0.3326 coverage and 0.2200 exact-match risk. Snippet retention and citation precision, rather than candidate generation alone, limited joint reliability. The results establish a measured quality-cost trade-off and identify where deeper retrieval still fails to improve answers.

Downloads

Download data is not yet available.

References

[1] W. Pi, “Efficient Information Retrieval and Response Generation with Retrieval-Augmented Generation (RAG),” 2024. doi: 10.59350/r9dj1-zkx52.

[2] M. Kang, S. Lee, J. Baek, K. Kawaguchi, and S. J. Hwang, “Knowledge-Augmented Reasoning Distillation for Small Language Models in Knowledge-Intensive Tasks,” Advances in Neural Information Processing Systems 36, pp. 48573-48602, 2023. doi: 10.52202/075280-2109.

[3] J. Jin, “Evidence-Chain Reliable RAG: Hallucination Detection, Source Attribution, and Deterministic Provenance Explanations,” Journal of Technology Informatics and Engineering, vol. 4, no. 2, pp. 520-533, 2025. doi: 10.51903/jtie.v4i2.535.

[4] Z. S. Zhong, J. Chen, E. Zhong, and X. Sun, “Evidence-Calibrated RAG for Unanswerable Question Answering: Retrieval Coverage, Abstention Calibration, and Hallucination-Proxy Analysis on SQuAD 2.0,” Journal of Technology Informatics and Engineering, vol. 4, no. 2, pp. 502-520, 2025. doi: 10.51903/jtie.v4i2.536.

[5] Y. Chen, and H. Xu, “Trust-Calibrated Multilingual RAG for Humanitarian Information Platforms: Empirical Evaluation on OMoS-QA for Migration Information Access,” International Journal of Graphic Design, vol. 4, no. 1, pp. 141-164, 2026. doi: 10.51903/ijgd.v4i1.3552.

[6] W. Su, S. Chen, and C. Zhao, “Budgeted Multi-Hop Retrieval Agent for Compositional Question Answering: A Retrieval-Policy Evaluation on the Official MultiHop-RAG Benchmark,” Journal of Technology Informatics and Engineering, vol. 4, no. 3, pp. 649-662, 2025. doi: 10.51903/jtie.v4i3.543.

[7] S. Jeong, J. Baek, S. Cho, S. J. Hwang, and J. Park, “Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity,” Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 7036-7050, 2024. doi: 10.18653/v1/2024.naacl-long.389.

[8] F. ding, “How to Precisely Update Large Language Models Knowledge While Avoiding Catastrophic Forgetting,” 2024. doi: 10.55415/deep-2024-0007.v1.

[9] Ziliang Samuel Zhong, Chenyu Li, and Hengning Rao, “Trajectory Reliability Prediction for Generalist AI Agents: Tool-Use Failure Analysis and Success Forecasting on ZClawBench,” Journal of Technology Informatics and Engineering, vol. 5, no. 1, pp. 341-360, 2026. doi: 10.51903/jtie.v5i1.539.

[10] J. Jin, “LLM-style evidence cards for scientific search interfaces: A UI/UX design framework for retrieval transparency, ranking trust, and visual evidence hierarchy,” Int. J. Graph. Des., vol. 3, no. 2, pp. 397–414, 2025, doi: 10.51903/ijgd.v3i2.3698.

[11] Z. Li, Kai Zhang, and A. Wong, “Numerical-Reasoning Guardrails for a Quant Research Assistant: A Compact Reproducible Benchmark Using SEC and FRED Data,” Journal of Technology Informatics and Engineering, vol. 5, no. 2, pp. 75-90, 2026. doi: 10.51903/jtie.v5i2.541.

[12] Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. Cohen, R. Salakhutdinov, et al., “HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering,” Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp. 2369-2380, 2018. doi: 10.18653/v1/d18-1259.

[13] X. Ho, A. K. Duong Nguyen, S. Sugawara, and A. Aizawa, “Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps,” Proceedings of the 28th International Conference on Computational Linguistics, pp. 6609-6625, 2020. doi: 10.18653/v1/2020.coling-main.580.

[14] H. Trivedi, N. Balasubramanian, T. Khot, and A. Sabharwal, “♫ MuSiQue: Multihop Questions via Single-hop Question Composition,” Transactions of the Association for Computational Linguistics, vol. 10, pp. 539-554, 2022. doi: 10.1162/tacl_a_00475.

[15] Z. S. Zhong, X. Pan, and Q. Lei, “Bridging domains with approximately shared features,” in Proc. 28th Int. Conf. Artificial Intelligence and Statistics (AISTATS), PMLR, vol. 258, 2025, pp. 559–567.

[16] Y. Zhang, P. Nie, A. Ramamurthy, and L. Song, “Answering Any-hop Open-domain Questions with Iterative Document Reranking,” Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 481-490, 2021. doi: 10.1145/3404835.3462853.

[17] H. Trivedi, N. Balasubramanian, T. Khot, and A. Sabharwal, “Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions,” Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 10014-10037, 2023. doi: 10.18653/v1/2023.acl-long.557.

[18] Z. Zhuang, Z. Zhang, S. Cheng, F. Yang, J. Liu, S. Huang, et al., “EfficientRAG: Efficient Retriever for Multi-Hop Question Answering,” Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 3392-3411, 2024. doi: 10.18653/v1/2024.emnlp-main.199.

[19] H. Lee, S. Yang, H. Oh, and M. Seo, “Generative Multi-hop Retrieval,” Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 1417-1436, 2022. doi: 10.18653/v1/2022.emnlp-main.92.

[20] Z. Shao, Y. Gong, Y. Shen, M. Huang, N. Duan, and W. Chen, “Enhancing Retrieval-Augmented Large Language Models with Iterative Retrieval-Generation Synergy,” Findings of the Association for Computational Linguistics: EMNLP 2023, pp. 9248-9274, 2023. doi: 10.18653/v1/2023.findings-emnlp.620.

[21] R. Zhang, Z. Wen, C. Wang, C. Tang, P. Xu, and Y. Jiang, “Quality analysis and evaluation prediction of RAG retrieval based on machine learning algorithms,” 2025. doi: 10.48550/arxiv.2511.19481.

[22] S. Robertson, and H. Zaragoza, “The Probabilistic Relevance Framework: BM25 and Beyond,” Foundations and Trends® in Information Retrieval, vol. 4, no. 1-2, pp. 1-174, 2009. doi: 10.1561/1500000019.

[23] V. Karpukhin, B. Oguz, S. Min, P. Lewis, L. Wu, S. Edunov, et al., “Dense Passage Retrieval for Open-Domain Question Answering,” Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 6769-6781, 2020. doi: 10.18653/v1/2020.emnlp-main.550.

[24] S. Deerwester, S. T. Dumais, G. W. Furnas, T. K. Landauer, and R. Harshman, “Indexing by latent semantic analysis,” Journal of the American Society for Information Science, vol. 41, no. 6, pp. 391-407, 1990. doi: 10.1002/(sici)1097-4571(199009)41:6<391::aid-asi1>3.0.co;2-9.

[25] K. Wojtasik, K. Wołowiec, V. Shishkin, A. Janz, and M. Piasecki, “BEIR-PL: Zero Shot Information Retrieval Benchmark for the Polish Language,” Proceedings of the Language Resources and Evaluation Conference, pp. 2149-2160, 2024. doi: 10.63317/3tsv3vgn8kns.

[26] G. V. Cormack, C. L. A. Clarke, and S. Buettcher, “Reciprocal rank fusion outperforms condorcet and individual rank learning methods,” Proceedings of the 32nd international ACM SIGIR conference on Research and development in information retrieval, pp. 758-759, 2009. doi: 10.1145/1571941.1572114.

[27] Z. Samuel Zhong, and S. Ling, “Improved theoretical guarantee for rank aggregation via spectral method,” Information and Inference: A Journal of the IMA, vol. 13, no. 3, 2024. doi: 10.1093/imaiai/iaae020.

[28] D. Zhao, and H. Fang, “Higway BERT for Passage Ranking,” 2019. doi: 10.6028/nist.sp.1250.deep-udel_fang.

[29] S. Meng, J. Chen, and I. Zheng, “LLM-Inspired Offline Reranking for Financial Search: Query Rewriting, Hybrid Retrieval, and Listwise Relevance Ranking on FiQA,” Journal of Technology Informatics and Engineering, vol. 5, no. 1, pp. 361-378, 2026. doi: 10.51903/jtie.v5i1.537.

[30] A. Koo, “Self-Reflection Through Visual Art: Social Emotional Learning Strategy and Application,” AERA 2024, 2024. doi: 10.3102/ip.24.2106302.

[31] Z. Jiang, F. Xu, L. Gao, Z. Sun, Q. Liu, J. Dwivedi-Yu, et al., “Active Retrieval Augmented Generation,” Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 7969-7992, 2023. doi: 10.18653/v1/2023.emnlp-main.495.

[32] W. Su, Y. Tang, Q. Ai, Z. Wu, and Y. Liu, “DRAGIN: Dynamic Retrieval Augmented Generation based on the Real-time Information Needs of Large Language Models,” Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 12991-13013, 2024. doi: 10.18653/v1/2024.acl-long.702.

[33] C. Li, G. Liu, and Z. Zhao, “Cost-Aware LLM-Style Routing for AIOps Log Analysis: Log Parsing, Anomaly Detection, Fault Diagnosis, and Incident Summarization on LogEval Task Files,” Journal of Technology Informatics and Engineering, vol. 5, no. 2, pp. 91-103, 2026. doi: 10.51903/jtie.v5i2.538.

[34] M. Młyńczak, and G. Cybulski, “Flow Parameters Derived from Impedance Pneumography after Nonlinear Calibration based on Neural Networks,” Proceedings of the 10th International Joint Conference on Biomedical Engineering Systems and Technologies, pp. 70-77, 2017. doi: 10.5220/0006146800700077.

[35] Y. Wiener, and R. El-Yaniv, “Agnostic Pointwise-Competitive Selective Classification,” Journal of Artificial Intelligence Research, vol. 52, pp. 171-201, 2015. doi: 10.1613/jair.4439.

[36] G. Bar-Shalom, Y. Geifman, and R. El-Yaniv, “Window-Based Distribution Shift Detection for Deep Neural Networks,” Advances in Neural Information Processing Systems 36, pp. 22978-22998, 2023. doi: 10.52202/075280-0996.

[37] A. Kamath, R. Jia, and P. Liang, “Selective Question Answering under Domain Shift,” Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 5684-5696, 2020. doi: 10.18653/v1/2020.acl-main.503.

[38] Q. Xin, “Uncertainty-Aware Late Fusion for 3D Perception (Confidence Calibration + Fusion Rule Learning),” Journal of Technology Informatics and Engineering, vol. 4, no. 1, pp. 215-238, 2026. doi: 10.51903/jtie.v4i1.485.

[39] J. Jin, “Calibrated Resume-Job Matching for Trustworthy LLM-Assisted Recruiter Screening: Pairwise Matching, Probability Calibration, and Selective Refusal on Two Public Recruitment Datasets,” Journal of Technology Informatics and Engineering, vol. 4, no. 3, pp. 625-648, 2025. doi: 10.51903/jtie.v4i3.529.

[40] Z. S. Zhong, and S. Ling, “Uncertainty Quantification of Spectral Estimator and MLE for Orthogonal Group Synchronization,” 2024. doi: 10.48550/arxiv.2408.05944.

[41] Q. Xin, “<b>Explainable and Fair Credit Risk Scoring with Counterfactual Explanations: A Reproducible Evaluation on the German Credit Dataset (HELOC-Motivated)</b>,” J-INTECH, vol. 14, no. 02, pp. 215-231, 2026. doi: 10.32664/j-intech.v14i02.2228.

[42] K. Järvelin, and J. Kekäläinen, “Cumulated gain-based evaluation of IR techniques,” ACM Transactions on Information Systems, vol. 20, no. 4, pp. 422-446, 2002. doi: 10.1145/582415.582418.

[43] P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang, “SQuAD: 100,000+ Questions for Machine Comprehension of Text,” Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pp. 2383-2392, 2016. doi: 10.18653/v1/d16-1264.

[44] T. Gao, H. Yen, J. Yu, and D. Chen, “Enabling Large Language Models to Generate Text with Citations,” Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 6465-6488, 2023. doi: 10.18653/v1/2023.emnlp-main.398.

[45] S. Es, J. James, L. Espinosa Anke, and S. Schockaert, “RAGAs: Automated Evaluation of Retrieval Augmented Generation,” Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, pp. 150-158, 2024. doi: 10.18653/v1/2024.eacl-demo.16.

[46] J. Saad-Falcon, O. Khattab, C. Potts, and M. Zaharia, “ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems,” Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 338-354, 2024. doi: 10.18653/v1/2024.naacl-long.20.

[47] Y. Chen, S. Zhou, and E. Lin, “Accounting-Aware Evidence Retrieval for Institutional Due Diligence of Tokenized Trade Receivable RWA,” Journal of Technology Informatics and Engineering, vol. 4, no. 3, pp. 649-663, 2025. doi: 10.51903/jtie.v4i3.542.

[48] K. Zhang, Y. Chen, and A. Qian, “Evidence-grounded accounting disclosure review cards: A visual communication framework for LLM-style explanations over SEC financial statements and notes,” Int. J. Graph. Des., vol. 3, no. 2, p. 395, 2025, doi: 10.51903/ijgd.v3i2.3710.

[49] S. Zhou, Y. Chen, and K. Lee, “Accounting-Aware Evidence-Constrained Agents for Disclosure, Settlement, and Secondary-Market Risk Monitoring in Tokenized,” Journal of Technology Informatics and Engineering, vol. 5, no. 2, pp. 60-74, 2026. doi: 10.51903/jtie.v5i2.544.

[50] W. Su, S. Chen, and E. Qian, “Narrative-Aware Scientific Claim Verification Agent with Evidence Ranking for ClimateCheck,” Journal of Technology Informatics and Engineering, vol. 5, no. 1, pp. 327-340, 2026. doi: 10.51903/jtie.v5i1.549.

[51] J. Li, and A. Zhou, “Multi-Regulation RAG for AI Product Counsel: A Legal Governance Framework for Cross-Border Digital Commerces,” Rule of Law Studies Journal, vol. 2, no. 2, pp. 91-108, 2026. doi: 10.64780/rolsj.v2i2.225.

[52] Z. Li, S. Zhou, and Z. Zhou, “Financial risk dashboard design for institutional RWA investors: Visual hierarchy, chart comprehension, and explainability in FinChart-Bench,” Int. J. Graph. Des., vol. 3, no. 1, pp. 196–210, 2025, doi: 10.51903/ijgd.v3i1.3715.

[53] Z. S. Zhong, Q. Wu, and G. Mi, “Uncertainty-aware medical image explanation cards: LLM-generated visual explanations for AI-assisted radiology interfaces,” Int. J. Graph. Des., vol. 3, no. 2, pp. 415–436, 2025, doi: 10.51903/ijgd.v3i2.3616.

[54] B. Zhou, H. Wang, and X. Chang, “Risk-calibrated patient-facing AI safety cards: A UI/UX design framework for rubric-based medical risk communication,” Int. J. Graph. Des., vol. 3, no. 2, pp. 365–380, 2025, doi: 10.51903/ijgd.v3i2.3696.

[55] C. Li, B. Zhou, and K. Gao, “Risk-calibrated patient-facing AI safety cards: A UI/UX benchmark for explainable medical AI response interfaces,” Int. J. Graph. Des., vol. 3, no. 2, pp. 381–394, 2025, doi: 10.51903/ijgd.v3i2.3709.

[56] B. Zhang, Y. Ren, and J. Zou, “LLM-style explainable e-commerce recommendation cards: A UI/UX design framework for trust-calibrated product recommendation,” Int. J. Graph. Des., vol. 3, no. 2, pp. 381–396, 2025, doi: 10.51903/ijgd.v3i2.3697.

[57] Y. Li, S. Lu, and L. Zhao, “LLM-as-design-critic: Aligning AI-generated UI feedback with human graphic design judgment,” Int. J. Graph. Des., vol. 3, no. 1, pp. 196–215, 2025, doi: 10.51903/ijgd.v3i1.3661.

[58] X. Chang, Y. Lu, and Z. S. Zhong, “Review-Grounded Explainable Recommendation with Faithfulness Evaluation on Amazon Reviews,” JEECS (Journal of Electrical Engineering and Computer Sciences), vol. 11, no. 1, pp. 9-22, 2026. doi: 10.54732/jeecs.v11i1.2.

[59] Yifan Zhang, and Hailey Zhang, “A Therapist-Facing Session Copilot for Live Counseling Support: Reasoning-Guided Retrieval and Ranking from Multi-Turn Counseling Dialogues,” Journal of Technology Informatics and Engineering, vol. 4, no. 2, pp. 464-486, 2025. doi: 10.51903/jtie.v4i2.547.

[60] Q. Xin, “Behavior retrieval plus response generation for interpretable conversational personalized recommendation,” IJEEPSE, vol. 9, no. 2, pp. 120–136, 2026, doi: 10.31258/ijeepse.9.2.120-136.

[61] Y. Zhang, and H. Zhang, “Visualizing the Right Counseling Support: Evidence-Linked Recommendation Cards for Explainable Mental Health Intake Interfaces,” International Journal of Graphic Design, vol. 3, no. 1, pp. 214-229, 2025. doi: 10.51903/ijgd.v3i1.3722.

[62] Q. Xin, “Auditable automated essay scoring and formative feedback: A rubric-grounded pipeline for secondary and higher education,” J. Artificial Intelligence Educ., vol. 2, no. 1, 2026, doi: 10.66053/jaaie.v2i1.348.

[63] H. Xu, Y. Chen, and A. Med, “Automatic Detection and Explanation of Dark Patterns from Interface Microcopy: Empirical Comparison of BERT-Style Encoders, RoBERTa-Style Encoders, and LLM-Style Decoders on the ec-darkpattern Dataset,” Journal of Technology Informatics and Engineering, vol. 4, no. 3, pp. 590-612, 2025. doi: 10.51903/jtie.v4i3.491.

[64] Q. Xin, “Early-warning analytics with LLM intervention rationales for student retention decisions: Classroom interaction modeling with xAPI-edu-data and dropout/success prediction,” Interdiscip. J. Pedagog. Res. Media Technol., vol. 2, no. 1, 2026, doi: 10.64268/inspire.v2i1.117.

[65] X. Sun, Ziliang Samuel Zhong, and Q. Wu, “<b>Retrieval-Grounded HDFS Log Anomaly Detection and Deterministic Failure Narrative Generation</b>,” Journal of Computational Systems and Applications, vol. 3, no. 1, pp. 15-30, 2026. doi: 10.64229/j6d7fr94.

[66] Q. Xin, “Explaining OpenStack Failure-Injection Log Anomalies with Retrieved Normal Prototypes,” Emerging Information Science and Technology, vol. 6, no. 2, pp. 125-146, 2025. doi: 10.18196/eist.v6i2.31232.

[67] J. Nie, G. Liu, and L. Zhao, “Evidence-Constrained Incident Visualization Cards for Distributed Cloud Logs: A UI/UX Framework for Turning Hadoop, OpenStack, and ZooKeeper Logs into Actionable SRE Design Interfaces,” International Journal of Graphic Design, vol. 4, no. 1, pp. 179-185, 2026. doi: 10.51903/ijgd.v4i1.3703.

[68] Q. Xin, “Log Anomaly Detection with Conformal Alert Control and Evidence-Grounded Incident Ticket Generation,” Aviation Electronics, Information Technology, Telecommunications, Electricals, and Controls (AVITEC), vol. 8, no. 2, pp. 247, 2026. doi: 10.28989/avitec.v8i2.3974.

[69] J. Bai, S. Chen, D. Zheng, and M.-J. Kuo, “Interpretable attack-chain stage detection from AWS CloudTrail event sequences via linear models and HMM smoothing,” Inf. Electr. Electron. Eng., vol. 6, no. 1, pp. 28–43, 2026, doi: 10.33474/infotron.v6i1.24923.

[70] Q. Xin, “Self-Supervised Log Anomaly Detection with LogBERT-Style Transformers: Full Empirical Evaluation on a Reproducible SynHDFS Benchmark,” JEECS (Journal of Electrical Engineering and Computer Sciences), vol. 11, no. 1, pp. 23-35, 2026. doi: 10.54732/jeecs.v11i1.3.

[71] Y. Li, and S. Lu, “Language-Guided Feature Selection for DDoS and Intrusion Detection on CICIDS2017,” Journal of Technology Informatics and Engineering, vol. 4, no. 1, pp. 284-305, 2025. doi: 10.51903/jtie.v4i1.531.

[72] Q. Xin, Z. Xu, L. Guo, F. Zhao, and B. Wu, “IoT traffic classification and anomaly detection method based on deep autoencoders,” Applied and Computational Engineering, vol. 69, no. 1, pp. 64-70, 2024. doi: 10.54254/2755-2721/69/20241511.

[73] B. Wang, Y. He, Z. Shui, Q. Xin, and H. Lei, “Predictive optimization of DDoS attack mitigation in distributed systems using machine learning,” Applied and Computational Engineering, vol. 64, no. 1, pp. 94-99, 2024. doi: 10.54254/2755-2721/64/20241350.

[74] Q. Xin, “Host-Based Intrusion Detection with System Call Sequences: Window Localization and Forensic Narratives,” Aviation Electronics, Information Technology, Telecommunications, Electricals, and Controls (AVITEC), vol. 8, no. 2, pp. 325, 2026. doi: 10.28989/avitec.v8i2.3973.

[75] W. Su, H. Rao, and E. Ma, “Privacy and Data-Integrity Risk Cards for LLM Agents: A UI/UX Design Framework for Secure Human Oversight under Prompt-Injection Attacks,” International Journal of Graphic Design, vol. 4, no. 1, pp. 186-191, 2026. doi: 10.51903/ijgd.v4i1.3699.

[76] S. Zhao, J. Bai, and D. Roberson, “Multi-Horizon GPU Demand Forecasting with Workload Semantics and Operational Risk Curves: An Empirical Study on Alibaba Clusterdata GPU Trace,” Journal of Technology Informatics and Engineering, vol. 4, no. 3, pp. 544-571, 2025. doi: 10.51903/jtie.v4i3.498.

[77] S. He, C. Li, and H. Rao, “Few-Shot Cold-Start Workload Forecasting for New AI Inference Tenants with Time-Series Foundation Models,” Journal of Technology Informatics and Engineering, vol. 4, no. 1, pp. 306-324, 2025. doi: 10.51903/jtie.v4i1.546.

[78] Q. Xin, “Hybrid Cloud Architecture for Efficient and Cost-Effective Large Language Model Deployment,” Journal of Information Systems and Informatics, vol. 7, no. 3, pp. 2182-2195, 2025. doi: 10.51519/journalisi.v7i3.1170.

[79] Siming Zhao, Y. Ren, and Xiaohan Chang, “Profit-Aware Spot GPU Admission Control with Cost-Sensitive Loss and Evidence-Grounded Policy Memos for AI Workload Supply-Demand Matching,” Journal of Technology Informatics and Engineering, vol. 5, no. 2, pp. 45-59, 2026. doi: 10.51903/jtie.v5i2.545.

[80] S. He, J. Nie, and C. Li, “Power-Aware Inventory Planning for AI Infrastructure Using Job-Level Forecasting and LLM Workload Explanations,” Journal of Technology Informatics and Engineering, vol. 5, no. 1, pp. 341-359, 2026. doi: 10.51903/jtie.v5i1.548.

[81] G. Liu, S. He, and H. Wong, “LLM-Compatible Visual Brief Cards for AI Infrastructure Capacity Dashboards: A UI/UX Framework for Turning Forecast Risk into Graphic Design Decisions,” International Journal of Graphic Design, vol. 3, no. 1, pp. 196-213, 2025. doi: 10.51903/ijgd.v3i1.3723.

[82] C. Wang, Z. Wen, R. Zhang, P. Xu, and Y. Jiang, “GPU Memory Requirement Prediction for Deep Learning Task Based on Bidirectional Gated Recurrent Unit Optimization Transformer,” 2025 5th International Conference on Artificial Intelligence, Virtual Reality and Visualization (AIVRV), pp. 31-35, 2025. doi: 10.1109/aivrv67401.2025.11350369.

[83] J. Mu, T. Ye, and P. Patel, “Offline counterfactual evaluation for advertising and recommendation slot policies: A reproducible study on the Open Bandit Dataset (Small),” J. Technol. Informatics Eng., vol. 4, no. 3, pp. 521–543, 2025, doi: 10.51903/jtie.v4i3.500.

[84] J. Bai, H. Wang, Q. Wu, and B. Zhang, “Privacy-Robust Incrementality Estimation in Cookieless Settings via Uplift Modeling: Reproducible Evidence from the Hillstrom E-Mail Experiment,” Journal of Technology Informatics and Engineering, vol. 5, no. 1, pp. 17-38, 2026. doi: 10.51903/jtie.v5i1.468.

[85] T. Ye, J. Mu, and J. Hunter, “Off-Policy Evaluation and Conservative Policy Selection for Slot-Level Dynamic Bidding and Ranking on the Open Bandit Dataset (Small),” Journal of Technology Informatics and Engineering, vol. 5, no. 1, pp. 178-199, 2026. doi: 10.51903/jtie.v5i1.503.

[86] Y. Lu, H. Zhou, and Y. Zhang, “A Constrained, Data-Driven Budgeting Framework Integrating Macro Demand Forecasting and Marketing Response Modeling,” Journal of Technology Informatics and Engineering, vol. 4, no. 3, pp. 493-520, 2025. doi: 10.51903/jtie.v4i3.466.

[87] J. Bai and Q. Wu, “Privacy-safe marketing mix modeling and budget optimization under identifier loss: A controlled simulation study,” Int. J. Electron. Commun. Syst., vol. 6, no. 1, 2026, doi: 10.24042/ijecs.v6i1.30533.

[88] Q. Wu, S. Meng, and J. Zhao, “Text-Grounded LLM-Assisted Design Rationale Interfaces: Turning Advertising Layout Metadata into Explainable UI/UX Decision Cards,” International Journal of Graphic Design, vol. 3, no. 1, pp. 216-240, 2025. doi: 10.51903/ijgd.v3i1.3713.

[89] H. Tu, S. Zhao, and A. Zhou, “Visual brief cards for advertising design: A structured UI/UX framework for turning creative intentions into graphic design decisions,” Int. J. Graph. Des., vol. 3, no. 1, pp. 210–226, 2025, doi: 10.51903/ijgd.v3i1.3714.

[90] J. Mu, Y. Lu, and E. Hwang, “Structured Visual Brief Interfaces for Advertising Design: A UI/UX Framework for Turning Creative Intentions into Designer-Editable Graphic Design Cards,” International Journal of Graphic Design, vol. 4, no. 1, pp. 192-208, 2026. doi: 10.51903/ijgd.v4i1.3702.

Downloads

Published

25-06-2026

How to Cite

Kong, C. (2026). Budgeted Multi-Hop Retrieval Agents with Evidence-Chain Reliability and Cost-Aware Abstention. Journal of Knowledge Learning and Science Technology ISSN: 2959-6386 (online), 5(2), 79-102. https://doi.org/10.60087/jklst.vol5.n2.006

Similar Articles

31-40 of 74

You may also start an advanced similarity search for this article.