BALANGAY-RAG: Adaptive Pre-Generation Retrieval and Abstention for Filipino-English Code-Switched Queries
DOI:
https://doi.org/10.69569/jip.2026.443xKeywords:
Abstention, Adaptive retrieval, Code-switching, Retrieval-augmented generation, TaglishAbstract
Fixed retrieval depths can waste context or proceed without designated support when language form, evidence availability, and context budgets vary. This controlled pre-generation study evaluates BALANGAY-RAG, combining the BALANGAY-PH benchmark with BALANGAY-RF, a Random Forest controller that selects k = 1, 2, 4, or 8 passages or abstains. The reported experiment contains 44 public-information passages, 147 English, Filipino, and Taglish queries, and 1,764 grouped held-out states. BALANGAY-RF achieved 0.80 feasible-evidence retrieval recall, 0.80 gold-passage precision among proceeded states, a 0.11 gold-absent proceed rate, and 15.49 unconditional mean retrieved words. A separately archived post-review reconstruction and reproduction regenerates 1,764 states and 10,584 policy outcomes from 39 official-source URLs, reproducing the qualitative context-coverage trade-off while yielding different numerical operating points. No downstream generator was evaluated.
Downloads
References
Asai, A., Wu, Z., Wang, Y., Sil, A., & Hajishirzi, H. (2024). Self-RAG: Learning to retrieve, generate, and critique through self-reflection. The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=hSyW5go0v8
Breiman, L. (2001). Random forests. Machine Learning, 45, 5–32. https://doi.org/10.1023/A:1010933404324
Es, S., James, J., Espinosa-Anke, L., & Schockaert, S. (2024). RAGAs: Automated evaluation of retrieval augmented generation. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, 150–158. https://doi.org/10.18653/v1/2024.eacl-demo.16
Herrera, M., Aich, A., & Parde, N. (2022). TweetTaglish: A dataset for investigating Tagalog-English code-switching. Proceedings of the Thirteenth Language Resources and Evaluation Conference, 2090–2097. https://aclanthology.org/2022.lrec-1.225/
Jeong, S., Baek, J., Cho, S., Hwang, S.J., & Park, J. (2024). Adaptive-RAG: Learning to adapt retrieval-augmented large language models through question complexity. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 7036–7050. https://doi.org/10.18653/v1/2024.naacl-long.389
Jiang, Z., Xu, F., Gao, L., Sun, Z., Liu, Q., Dwivedi-Yu, J., Yang, Y., Callan, J., & Neubig, G. (2023). Active retrieval augmented generation. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 7969–7992. https://doi.org/10.18653/v1/2023.emnlp-main.495
Kim, Y., Kim, H.J., Park, C., Park, C., Cho, H., Kim, J., Yoo, K. M., Lee, S.-G., & Kim, T. (2024). Adaptive contrastive decoding in retrieval-augmented generation for handling noisy contexts. Findings of the Association for Computational Linguistics: EMNLP 2024, 2421–2431. https://doi.org/10.18653/v1/2024.findings-emnlp.136
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459–9474.
Liu, W., Trenous, S., Ribeiro, L.F.R., Byrne, B., & Hieber, F. (2025). XRAG: Cross-lingual retrieval-augmented generation. Findings of the Association for Computational Linguistics: EMNLP 2025, 15669–15690. https://doi.org/10.18653/v1/2025.findings-emnlp.849
Manning, C.D., Raghavan, P., & Schütze, H. (2008). Introduction to information retrieval. Cambridge University Press.
Moskvoretskii, V., Marina, M., Salnikov, M., Ivanov, N., Pletenev, S., Galimzianova, D., Krayko, N., Konovalov, V., Nikishina, I., & Panchenko, A. (2025). Adaptive retrieval without self-knowledge? Bringing uncertainty back home. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6355–6384. https://doi.org/10.18653/v1/2025.acl-long.319
Peng, X., Choubey, P.K., Xiong, C., & Wu, C.-S. (2025). Unanswerability evaluation for retrieval augmented generation. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8452–8472. https://doi.org/10.18653/v1/2025.acl-long.415
Tang, X., Gao, Q., Li, J., Du, N., Li, Q., & Xie, S. (2025). MBA-RAG: A bandit approach for adaptive retrieval-augmented generation through question complexity. Proceedings of the 31st International Conference on Computational Linguistics, 3248–3254. https://aclanthology.org/2025.coling-main.218/
Tiopes, J.S., Tiopes, J.B., Orlanda, W., & Desoloc, R. (2026). BALANGAY-RAG Post-Review Reconstruction and Independent Reproduction Package [Dataset]. Zenodo. https://doi.org/10.5281/zenodo.22274900
Yoran, O., Wolfson, T., Ram, O., & Berant, J. (2024). Making retrieval-augmented language models robust to irrelevant context. The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=ZS4m74kZpH
Zeng, Q., Lu, Y., Zhou, Z., Qi, H., Yu, P., Zhao, F., Yanaka, H., Xuan, W., & Yokoya, N. (2026). Code-switching information retrieval: Benchmarks, analysis, and the limits of current retrievers. Findings of the Association for Computational Linguistics: ACL 2026, 13055–13071. https://doi.org/10.18653/v1/2026.findings-acl.636
Zhang, Z., Fang, M., & Chen, L. (2024). RetrievalQA: Assessing adaptive retrieval-augmented generation for short-form open-domain question answering. Findings of the Association for Computational Linguistics: ACL 2024, 6963–6975. https://doi.org/10.18653/v1/2024.findings-acl.415
Downloads
Published
How to Cite
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
JIP is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License. Under the Open Access Policy, appropriate attribution can be provided by simply citing the original article.

