BALANGAY-RAG: Adaptive Pre-Generation Retrieval and Abstention for Filipino-English Code-Switched Queries

Authors

  • John Syte L. Tiopes Parañaque City College, Coastal Road, cor Victor Medina, Parañaque, 1700 Metro Manila, Philippines
  • John Best L. Tiopes National University, Mall of Asia Complex, Pasay City, Metro Manila, Philippines
  • Wilmar S. Orlanda Parañaque City College, Coastal Road, cor Victor Medina, Parañaque, 1700 Metro Manila, Philippines
  • Ronnel A. Desoloc Parañaque City College, Coastal Road, cor Victor Medina, Parañaque, 1700 Metro Manila, Philippines

DOI:

https://doi.org/10.69569/jip.2026.443x

Keywords:

Abstention, Adaptive retrieval, Code-switching, Retrieval-augmented generation, Taglish

Abstract

Fixed retrieval depths can waste context or proceed without designated support when language form, evidence availability, and context budgets vary. This controlled pre-generation study evaluates BALANGAY-RAG, combining the BALANGAY-PH benchmark with BALANGAY-RF, a Random Forest controller that selects k = 1, 2, 4, or 8 passages or abstains. The reported experiment contains 44 public-information passages, 147 English, Filipino, and Taglish queries, and 1,764 grouped held-out states. BALANGAY-RF achieved 0.80 feasible-evidence retrieval recall, 0.80 gold-passage precision among proceeded states, a 0.11 gold-absent proceed rate, and 15.49 unconditional mean retrieved words. A separately archived post-review reconstruction and reproduction regenerates 1,764 states and 10,584 policy outcomes from 39 official-source URLs, reproducing the qualitative context-coverage trade-off while yielding different numerical operating points. No downstream generator was evaluated.

 

Downloads

Download data is not yet available.

References

Asai, A., Wu, Z., Wang, Y., Sil, A., & Hajishirzi, H. (2024). Self-RAG: Learning to retrieve, generate, and critique through self-reflection. The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=hSyW5go0v8

Breiman, L. (2001). Random forests. Machine Learning, 45, 5–32. https://doi.org/10.1023/A:1010933404324

Es, S., James, J., Espinosa-Anke, L., & Schockaert, S. (2024). RAGAs: Automated evaluation of retrieval augmented generation. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, 150–158. https://doi.org/10.18653/v1/2024.eacl-demo.16

Herrera, M., Aich, A., & Parde, N. (2022). TweetTaglish: A dataset for investigating Tagalog-English code-switching. Proceedings of the Thirteenth Language Resources and Evaluation Conference, 2090–2097. https://aclanthology.org/2022.lrec-1.225/

Jeong, S., Baek, J., Cho, S., Hwang, S.J., & Park, J. (2024). Adaptive-RAG: Learning to adapt retrieval-augmented large language models through question complexity. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 7036–7050. https://doi.org/10.18653/v1/2024.naacl-long.389

Jiang, Z., Xu, F., Gao, L., Sun, Z., Liu, Q., Dwivedi-Yu, J., Yang, Y., Callan, J., & Neubig, G. (2023). Active retrieval augmented generation. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 7969–7992. https://doi.org/10.18653/v1/2023.emnlp-main.495

Kim, Y., Kim, H.J., Park, C., Park, C., Cho, H., Kim, J., Yoo, K. M., Lee, S.-G., & Kim, T. (2024). Adaptive contrastive decoding in retrieval-augmented generation for handling noisy contexts. Findings of the Association for Computational Linguistics: EMNLP 2024, 2421–2431. https://doi.org/10.18653/v1/2024.findings-emnlp.136

Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459–9474.

Liu, W., Trenous, S., Ribeiro, L.F.R., Byrne, B., & Hieber, F. (2025). XRAG: Cross-lingual retrieval-augmented generation. Findings of the Association for Computational Linguistics: EMNLP 2025, 15669–15690. https://doi.org/10.18653/v1/2025.findings-emnlp.849

Manning, C.D., Raghavan, P., & Schütze, H. (2008). Introduction to information retrieval. Cambridge University Press.

Moskvoretskii, V., Marina, M., Salnikov, M., Ivanov, N., Pletenev, S., Galimzianova, D., Krayko, N., Konovalov, V., Nikishina, I., & Panchenko, A. (2025). Adaptive retrieval without self-knowledge? Bringing uncertainty back home. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6355–6384. https://doi.org/10.18653/v1/2025.acl-long.319

Peng, X., Choubey, P.K., Xiong, C., & Wu, C.-S. (2025). Unanswerability evaluation for retrieval augmented generation. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8452–8472. https://doi.org/10.18653/v1/2025.acl-long.415

Tang, X., Gao, Q., Li, J., Du, N., Li, Q., & Xie, S. (2025). MBA-RAG: A bandit approach for adaptive retrieval-augmented generation through question complexity. Proceedings of the 31st International Conference on Computational Linguistics, 3248–3254. https://aclanthology.org/2025.coling-main.218/

Tiopes, J.S., Tiopes, J.B., Orlanda, W., & Desoloc, R. (2026). BALANGAY-RAG Post-Review Reconstruction and Independent Reproduction Package [Dataset]. Zenodo. https://doi.org/10.5281/zenodo.22274900

Yoran, O., Wolfson, T., Ram, O., & Berant, J. (2024). Making retrieval-augmented language models robust to irrelevant context. The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=ZS4m74kZpH

Zeng, Q., Lu, Y., Zhou, Z., Qi, H., Yu, P., Zhao, F., Yanaka, H., Xuan, W., & Yokoya, N. (2026). Code-switching information retrieval: Benchmarks, analysis, and the limits of current retrievers. Findings of the Association for Computational Linguistics: ACL 2026, 13055–13071. https://doi.org/10.18653/v1/2026.findings-acl.636

Zhang, Z., Fang, M., & Chen, L. (2024). RetrievalQA: Assessing adaptive retrieval-augmented generation for short-form open-domain question answering. Findings of the Association for Computational Linguistics: ACL 2024, 6963–6975. https://doi.org/10.18653/v1/2024.findings-acl.415

Downloads

Published

2026-09-24

How to Cite

Tiopes, J. S., Tiopes, J. B., Orlanda, W., & Desoloc, R. (2026). BALANGAY-RAG: Adaptive Pre-Generation Retrieval and Abstention for Filipino-English Code-Switched Queries. Journal of Interdisciplinary Perspectives, 4(10), 436–447. https://doi.org/10.69569/jip.2026.443x