Cascading Multi-Agent Architectures for Multilingual IT Support: Integrating Explainability and Dynamic Workload Optimisation

Authors

  • Annette Nayana Nellyet University of Stirling, RAK, United Arab Emirates
  • Syed Muhammad Umar Afnan University of Stirling, RAK, United Arab Emirates

DOI:

https://doi.org/10.68104/ijasit.v1.i3.42

Keywords:

Multi-agent systems large language models, IT helpdesk, Multilingual natural language processing, Explainable AI workload optimisation

Abstract

Modern IT helpdesks increasingly need to support multiple languages while staying transparent and resistant to hallucination, yet most deployed chatbots still rely on single-model architectures that struggle on both counts. This paper presents NexaServe v2, a four-tier cascading multi-agent IT helpdesk system combining retrieval-based response, LLM validation, LLM fallback, and human escalation, alongside a hybrid multilingual classification pipeline, confidence-gated routing, and workload-aware ticket assignment. The system is evaluated on a public Kaggle multilingual helpdesk dataset (n = 28,587 tickets), 89 logged chatbot interactions, and operational logs collected during development and testing. Fine-tuning mBERT on combined English-German data reached 89.84% category accuracy, while an English-trained classical model applied to translated German priority tickets dropped from 77.70% to 57.23% accuracy, a much steeper decline than the equivalent category-classification result. The chatbot cascade maintained a low factual hallucination rate (2.7% on a 75-message held-out batch) despite rejecting 93.3-100% of retrieval-layer responses as insufficiently confident. An AI-based workload supervisor, despite generating fluent and specific justifications for its decisions, performed no better than a simple deterministic heuristic and violated its own explicit numerical rule in every recorded overload flag. These results demonstrate that a transparent, near-zero-marginal-cost multilingual helpdesk is achievable using local LLMs for classification and cascaded response generation, while showing that LLM-based supervision does not reliably outperform deterministic alternatives for tasks with strict, checkable correctness criteria.

Downloads

Download data is not yet available.

References

Bansal, G., Wu, T., Zhou, J., Fok, R., Nushi, B., Kamar, E., Ribeiro, M. T., & Weld, D. S. (2021). Does the whole exceed its parts? The effect of AI explanations on complementary team performance. Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, 1-16.

Bueck, T. (2023). Multilingual customer support tickets [Dataset]. Kaggle. https://www.kaggle.com/datasets/tobiasbueck/multilingual-customer-support-tickets

Cemri, M., Pan, M. Z., Yang, S., et al. (2025). Why do multi-agent LLM systems fail? arXiv.

Chawla, N. V., Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P. (2002). SMOTE: Synthetic minority over-sampling technique. Journal of Artificial Intelligence Research, 16, 321-357.

Chen, L., Zaharia, M., & Zou, J. (2024). FrugalGPT: How to use large language models while reducing cost and improving performance. Transactions on Machine Learning Research. arXiv:2305.05176.

Chen, Y., Sahoo, P., Pujar, S., et al. (2025). STRATUS: A multi-agent system for autonomous reliability engineering of modern clouds. Proceedings of the 39th Conference on Neural Information Processing Systems (NeurIPS 2025).

Conneau, A., Khandelwal, K., Goyal, N., et al. (2020). Unsupervised cross-lingual representation learning at scale. Proceedings of ACL 2020, 8440-8451.

De Koninck, J., Muller, M. N., & Vechev, M. (2024). A unified approach to routing and cascading for LLMs. arXiv:2410.10347.

Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of NAACL-HLT, 4171-4186.

Eden AI. (2025). The 2025 guide to retrieval-augmented generation (RAG) [Industry whitepaper].

Isbister, T., Sahlgren, M., & Kurtz, R. (2021). Why translate? A comparison of translation-based and language-agnostic text classification. Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa).

Joachims, T. (1998). Text categorization with support vector machines: Learning with many relevant features. Proceedings of the 10th European Conference on Machine Learning (ECML), 137-142.

Larsen, A. G., Skjuve, M. B., Folstad, A., & van As, N. (2025). LLM hallucinations in conversational AI for customer service: Framework and end-user perceptions. International Journal of Human-Computer Interaction.

Li, Y., Yang, Y., Song, P., Duan, L., & Ren, R. (2025). An improved SMOTE algorithm for enhanced imbalanced data classification by expanding sample generation space. Scientific Reports, 15, Article 23521. https://doi.org/10.1038/s41598-025-09506-w

Liao, Q. V., Gruen, D., & Miller, S. (2020). Questioning the AI: Informing design practices for explainable AI user experiences. Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (CHI ’20).

Loshchilov, I., & Hutter, F. (2019). Decoupled weight decay regularization. Proceedings of the International Conference on Learning Representations (ICLR).

Mahmoudi, L., Salem, M., & Alharbe, N. R. (2025). Addressing class imbalance in text classification with LLMs: A prompt-based GPT-2 approach. Journal of Information Science.

McKinsey & Company. (2025). Reimagining tech infrastructure for and with agentic AI [Technology insights report].

Mosqueira-Rey, E., Hernandez-Pereira, E., Alonso-Rios, D., Bobes-Bascaran, J., & Fernandez-Leal, A. (2023). Human-in-the-loop machine learning: A state of the art. Artificial Intelligence Review.

Pande, A., & Pande, A. (2021). Automated ticket routing in IT services using natural language processing. Proceedings of the IEEE International Conference on Artificial Intelligence and Knowledge Engineering (AIKE).

Rajammal, K., & Chinnadurai, M. (2025). Dynamic load balancing in cloud computing using predictive graph networks and adaptive neural scheduling. Scientific Reports, 15, Article 22181.

Ramya, C., Paramesh, S. P., & Shreedhara, K. S. (2021). Classifying the unstructured IT service desk tickets using ensemble of classifiers. arXiv:2103.15822.

Silva, S., Pereira, R., & Ribeiro, R. (2018). Machine learning in incident categorization automation. Proceedings of the 2018 13th Iberian Conference on Information Systems and Technologies (CISTI).

Wang, S., Tan, Z., Chen, Z., et al. (2025). AnyMAC: Cascading flexible multi-agent collaboration via next-agent prediction. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP).

Yamsani, N., & Chenna Reddy, P. (2026). SLA aware deep reinforcement learning for adaptive edge-cloud task scheduling. Scientific Reports, 16, Article 10037.

Downloads

Published

2026-09-20

Issue

Section

Articles

How to Cite

Nellyet, Annette Nayana, and Syed Muhammad Umar Afnan. 2026. “Cascading Multi-Agent Architectures for Multilingual IT Support: Integrating Explainability and Dynamic Workload Optimisation”. International Journal of Applied Smart Interdisciplinary Technologies (IJASIT) 1 (3): 117-34. https://doi.org/10.68104/ijasit.v1.i3.42.

Similar Articles

1-10 of 27

You may also start an advanced similarity search for this article.