![]()
ViTfuse_GNN: Vision Transformer-based fusion with Graphical Neural Network for detection of Intrusions in Multimedia Healthcare
Yashaswini K1, Sharath M N2
1Yashaswini K, Department of Computer Science and Engineering, Navkis College of Engineering, Thimanahally Hassan, Karnataka, India,
2Dr Sharath M N, Department of Artificial Intelligence and Machine Learning, Rajeev College of Engineering, Hassan, Karnataka, India.
Manuscript received on 03 June 2026 | First Revised Manuscript received on 10 June 2026 | Second Revised Manuscript received on 07 July 2026 | Manuscript Accepted on 15 July 2026 | Manuscript published on 30 July 2026 | PP: 1-9 | Volume-13 Issue-7, July 2026 | Retrieval Number: 100.1/ijies.G115013070726 | DOI: 10.35940/ijies.G1150.13070726
Open Access | Editorial and Publishing Policies | Cite | Zenodo | OJS | Indexing and Abstracting
© The Authors. Blue Eyes Intelligence Engineering and Sciences Publication (BEIESP). This is an open-access article under the CC-BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/)
Abstract: The exponential growth of network traffic and user data has hindered the efficacy and responsiveness of network intrusion detection systems. Intrusion systems are essential in e healthcare since they ensure the security, confidentiality, and accuracy of patients’ medical information. Diagnosis and treatment mistakes may result from any alteration to the patient’s real data. Traditional intrusion detection methods often fail to fully exploit the heterogeneous nature of multimedia healthcare data, which may include medical images, patient records, and network traffic metadata. Hence, this work proposes ViTfuse_GNN, a novel hybrid framework that integrates Vision Transformer (ViT)-based multimodal feature fusion with Graph Neural Networks (GNN) for robust intrusion detection. The proposed model first employs a Vision Transformer to fuse high level semantic features from medical images, videos, and textual metadata. Experimental evaluation on benchmark multimedia healthcare intrusion datasets demonstrates that ViTfuse_GNN outperforms state-of-the-art intrusion detection models with 99% accuracy, 98% precision, 99% recall, and 97% F1-score.
Keywords: Intrusion detection, Multimedia, Neural Network, Vision Transformer, Healthcare, VGG-19.
Scope of the Article: Computer Science and Engineering
