Sinhala-English Multilingual AI Call Center Bot with Sentiment-Aware Dialogue and Multimodal CSAT Prediction

Abstract

This paper presents a Sri Lanka focused, voice-based call center automation framework supporting Sinhala, English, and Sinhala-English code-mixed conversations in low-resource environments. The system adopts an end-to-end architecture integrating bilingual dataset construction, privacy-preserving speech processing, retrieval-grounded response generation, multimodal sentiment intelligence, and customer satisfaction (CSAT) estimation. A Sinhala-English call-center corpus is created using noise reduction, speaker diarization, and transcript alignment, combined with transcript-aligned PII redaction. During live interaction, language identification routes calls to unified processing pipelines. Real-time sentiment analysis with explainable risk scoring supports escalation decisions, while retrieval-augmented generation ensures factually grounded responses. Emotion-adaptive text-to-speech enhances conversational naturalness. The framework enables interaction-based CSAT estimation without relying solely on post-call surveys, providing scalable and privacy-aware automation tailored to multilingual Sri Lankan call center operations

Description

Keywords

Code-mixed speech processing, Customer satisfaction prediction, Multimodal sentiment analysis, Privacy-preserving speech processing, Retrieval-augmented generation

Citation

T. Kulathunga, B. Amarasinghe, V. Fernando, P. Lakruwani, M. Weerasinghe and D. Kasthurirathna, "Sinhala-English Multilingual AI Call Center Bot with Sentiment-Aware Dialogue and Multimodal CSAT Prediction," 2026 6th International Conference on Electrical, Computer and Energy Technologies (ICECET), Rome, Italy, 2026, pp. 1-6, doi: 10.1109/ICECET65726.2026.11632550.

Endorsement

Review

Supplemented By

Referenced By