Conference papers
Permanent URI for this communityhttps://rda.sliit.lk/handle/123456789/4889
Browse
9 results
Search Results
Item Embargo Adaptive AI-Based Enhancement of Critical External Sounds in Insulated Vehicle Cabins for Improved Safety(Institute of Electrical and Electronics Engineers Inc., 2025-12-09) Rathnayaka D.B; Wickramasuriya L.H.N.Y; Walpalage J.V; Rathnayake, SThe increasing acoustic insulation in modern and electric vehicles improves passenger comfort but unintentionally suppresses critical external sounds such as ambulance sirens, car horns, and train alarms, creating potential safety risks. While existing research has explored sound detection or localization in isolation, few systems integrate both capabilities in a unified framework for real-time vehicular deployment. This research proposes an adaptive AI-based system that detects, classifies, and selectively enhances these critical sounds in real time while providing directional awareness. Using a convolutional recurrent neural network (CRNN) trained on the UrbanSound8K dataset, the system processes incoming audio from external microphones, extracts Mel-frequency cepstral coefficients (MFCCs), and distinguishes safety-relevant cues from non-essential background noise. A dual-microphone setup enables the estimation of sound direction (left or right), providing additional spatial awareness to the driver. Detected signals are isolated through spectral filtering and relayed into the cabin with sub-30 ms latency, ensuring timely driver and passenger awareness without compromising comfort. Experimental results achieved 91.2% classification accuracy and 87.4% directional accuracy,confirming the system's feasibility for enhancing safety in insulated vehicle cabins and supporting future autonomous driving environments.Item Embargo HCLIP: Beyond CLIP for Cost-Effective Multimodal Retrieval in Education(Institute of Electrical and Electronics Engineers Inc., 2025-12-09) Weerasinghe, S; Gunatunga, O; Dewpura, W; Fernando, S; Kasthurirathna, D; Rathnayake, SMultimodal retrieval systems have gained significant attention due to their ability to process and cross-retrieve data containing images and text. However, the factors such as high cost of development, limitation on resources, and the proper addressing of the modality gap, the inherent representational differences between modalities pose a challenge to building effective and efficient retrieval models. In this work, we propose a low-resource, cost-efficient hybrid multimodal retrieval model that integrates Contrastive Language-Image Pre-training (CLIP) and All-MiniLM-L6-v2 to create a shared embedding space while storing raw images in an unstructured database. Our primary contributions include (1) the development of a hybrid model that outperforms CLIP-native retrieval, (2) a novel bidirectional neural network alignment technique that brings textual and visual modalities closer together, and (3) a comprehensive analysis of the modality gap's impact on downstream retrieval performance. Through proper evaluation using transparent techniques such as Mean Reciprocal Rank (MRR) and Cosine-Weighted MRR, our method demonstrates improved retrieval accuracy over baseline approaches. Experimental results exhibit that a lower modality gap does not always prove to be efficient on the downstream retrieval. Our findings pave the way for more efficient, adaptable, and cost-effective multimodal retrieval methodologies in low-resource environments, not limited to the education domain.Item Embargo NSCLC 360 - Leveraging Multi-Omics Data for a Holistic and Explainable Decision Support for Non Small Cell Lung Cancer Management(Institute of Electrical and Electronics Engineers Inc., 2025-07-03) Pirabaharan, A; Irfan, A. A.; Lareef, W; Ahamed, S; Rathnayake, S; Shyamalee, TNon-Small Cell Lung Cancer (NSCLC) remains a leading cause of cancer-related mortality, with existing diagnostic and prognostic models often failing to capture the complexity of tumor biology. This study proposes a holistic and explainable decision support system that integrates multi-omics data - including genomics, transcriptomics, and proteomics - along with advanced machine learning (ML) and deep learning (DL) techniques to enhance NSCLC detection, prognosis prediction, complication forecasting, and recurrence assessment. To address the challenge of interpretability in AI-driven healthcare, we incorporate Explainable AI (XAI) methods such as SHAP and LIME, ensuring model transparency and clinical trust. Additionally, traditional statistical models like Cox proportional hazards regression are combined with ML approaches for robust survival analysis, while modern AI architectures, including Vision Transformers and multi-task learning models, improve tumor localization and TNM classification. By developing an interpretable and clinically meaningful AI-based decision support system, this research aims to advance personalized lung cancer management and improve patient outcomes through seamless integration into clinical workflows.Item Embargo Object Detection Approach for Pure and Cross Chicken Breed Identification(Institute of Electrical and Electronics Engineers Inc., 2025-12-09) Jayarathna, N; Rathnayake, S; Panduwawala, PThe proposed system aims to identify different types of purebred and crossbred chicken breeds across ten categories. Traditional methods such as visual inspection are often subjective and inaccurate, making breed identification challenging. To address this, image processing and deep learning techniques were employed, with the YOLOv5 object detection algorithm trained on a custom data set of 1,310 images. The model achieved strong results, with an overall precision of 83.5%, recall of 81.8%, and mAP@0.5 of 89.3%. The Class-wise evaluation showed particularly high performance for the Brahma and Leghorn breeds. Based on these outcomes, a mobile application was designed to provide farmers with a fast, reliable, and cost-effective tool for breed identification. In addition to classification, the application provides detailed information on breed characteristics and commercial value, helping small-scale farmers improve productivity, efficiency, and animal welfareItem Embargo Multimodal Knowledge Graph for Domain-Specific Intelligence(Institute of Electrical and Electronics Engineers Inc., 2025-06-25) Mohan, K; Munasinghe, M; Bandara, L; Wijesinghe, H; Rathnayake, S; Abeywardhana, LIn the era of information abundance, transforming vast amounts of data into meaningful knowledge remains a critical challenge, especially in domains like medicine, engineering, and education, where visual and multimodal elements play a vital role. Traditional Knowledge Graphs (KGs) excel in organizing structured and textual data but struggle to incorporate multimodal information and implicit relationships, limiting their effectiveness. This paper explores the potential of Multimodal Knowledge Graphs (MMKGs) to address these limitations by integrating text, images, videos, and audio into a unified framework. We investigate how MMKGs enhance knowledge retrieval, comprehension, and interactive learning through advanced techniques, including Natural Language Processing and deep learning. Our findings demonstrate that MMKGs significantly improve knowledge retention and application in specialized fields, offering a foundation for more intuitive and effective domain-specific knowledge ecosystems.Item Embargo Knowledge Graph-Based AI Framework for Predicting Nutritional and Health Impacts of Food Ingredients(Institute of Electrical and Electronics Engineers Inc., 2026-08-04) Dakshina P.D.S.D; Rupasighe W.A.R.K; Waduge N.P; Nimsitha M.V.T; Tissera, W; Rathnayake, S; Krishara, JThe increasing complexity of modern food products and dietary supplements has made it challenging for both consumers and healthcare professionals to interpret nutritional information and assess the potential health risks associated with these products. Modern food labeling schemes provide static and fragmented information and cannot effectively capture the relationships between different ingredients, nutrients and their health effects. In this study, a new AI-based framework named Food Health Risk Analyzer has been proposed that utilizes KGs, GNNs, RAG and a dose-response module based on consumption quantities to perform the dynamic, explainable and evidence-based prediction of food-related health risks. The model uses heterogeneous data in order to analyze the relationships between ingredients and diseases to predict potential health risks while generating scientifically supported explanations as well. The experimental evaluation has shown high prediction accuracy with a micro-F1 score of 0.88 and AUC of 0.85 which shows that the framework surpasses conventional machine learning baseline models. In addition to that, the use of RAG has helped in improving the interpretability of predictions through evidence-based natural language explanations whereas dose-response module improves the practical relevance of risk assessment by considering the consumption quantities of ingredients.Item Embargo Adaptive Voice Communication in Emotion-Aware Digital Companions(Institute of Electrical and Electronics Engineers Inc., 2025) Rathnayake, P; Rathnaweera, C; Jithma, U; Aththanayake, I; Rathnayake, S; Gunaratne, MThis paper presents an adaptive voice communication system for emotion-aware digital companions that dynamically responds to users' affective states through expressive speech and synchronized 3D avatar animation. The system integrates real-time voice input, emotion recognition, and context-aware dialogue generation using GPT-3.5, followed by emotional text-to-speech synthesis via neural TTS. Lip-sync data is generated using phoneme alignment and rendered in sync with the avatar's facial expressions and gestures. To enhance user trust and engagement, the avatar visually mirrors the emotional tone of the speech. A cultural adaptation layer is introduced to align voice output and speech style with Sri Lankan communication norms, including tone, pacing, and formality. Implemented using a Node.js backend and React + Three.js frontend, the system demonstrates strong potential for emotionally intelligent, culturally adaptive AI interactions. This work contributes a modular pipeline for building empathetic voice agents capable of enhancing realism and trust in human-AI communication.Item Embargo Hybrid Motion Prediction for Autonomous Vehicles using GNN-Transformer Architecture(Institute of Electrical and Electronics Engineers Inc., 2025) Akalanka, A; Athukorala, D; Ganepola, N; Tharindu, I; Rathnayake, SAccurate perception and scene understanding are pivotal in enabling autonomous vehicles to navigate safely and intelligently. This paper presents an integrated perception module comprising three core subcomponents: real-time object detection using YOLOv5, lane-keeping using a CNN-based steering predictor, and a novel motion prediction architecture based on a hybrid Graph Neural Network (GNN) and Transformer design. The system is deployed and validated within the CARLA simulation environment, with custom data generation pipelines designed to mimic real-world behavioral patterns of nearby agents. The novelty lies in the hybrid GNN-Transformer model, which effectively captures both spatial and temporal interactions of dynamic objects for behavior classification. Experimental results demonstrate a high accuracy of 98.75% in classifying behaviors into four categories: Going, Coming, Crossing, and Stopped. This paper details the architecture, dataset creation, training methodology, and performance evaluation, highlighting the hybrid model's potential to improve trajectory planning modules in autonomous systems.Item Embargo Multimodal Knowledge Graph for Domain-Specific Intelligence(Institute of Electrical and Electronics Engineers Inc., 2025) Mohan, K; Munasinghe, M; Bandara, L; Wijesinghe, H; Rathnayake, S; Abeywardhana, LIn the era of information abundance, transforming vast amounts of data into meaningful knowledge remains a critical challenge, especially in domains like medicine, engineering, and education, where visual and multimodal elements play a vital role. Traditional Knowledge Graphs (KGs) excel in organizing structured and textual data but struggle to incorporate multimodal information and implicit relationships, limiting their effectiveness. This paper explores the potential of Multimodal Knowledge Graphs (MMKGs) to address these limitations by integrating text, images, videos, and audio into a unified framework. We investigate how MMKGs enhance knowledge retrieval, comprehension, and interactive learning through advanced techniques, including Natural Language Processing and deep learning. Our findings demonstrate that MMKGs significantly improve knowledge retention and application in specialized fields, offering a foundation for more intuitive and effective domain-specific knowledge ecosystems.
