Multimodal Biometric Recognition Based on Face and Video Via Vision Attention Fusion Transformer

Authors

  • S. PREETHI Affiliated to Bharathidasan University, Tiruchirappalli- 620002. https://orcid.org/0009-0007-4184-5463
  • P. JOSEPH CHARLES Affiliated to Bharathidasan University, Tiruchirappalli- 620002.

DOI:

https://doi.org/10.58414/SCIENTIFICTEMPER.2026.17.7.2498

Keywords:

Multimodal Biometric, Multi Task Cascaded Neural, Pre-emphasis Filter, Vision Transformer, Attention Mechanism, Fusion

Abstract

Biometric recognition systems have become an indispensable piece of contemporary security and recognition administration frameworks. Multimodal biometric systems incorporate multiple biometric features for verification provide significant advantages over uni-modal biometric systems, to name a few being, improved security, enhanced accuracy and considerable amount of robustness. Across a sample of individual modalities, face recognition is extensively applied for its reliability and ease of use, voice recognition is eminent for its high accuracy and stability. This study highlights the significance of multimodal biometric authentication method using face images and voice signals called, Multi Task Cascaded Pre-emphasis and Vision Attention Fusion Transformer (MTCP-VAFT) to enhance the security of existing multimodal biometric recognition systems in indoor surveillance videos. It introduces a novel pre-processing model that simultaneously processes face images using Multi Task Cascaded Neural model and voice signals utilizing and Pre-emphasis Filter. This work also proposes a novel Vision Attention Fusion Transformer based multimodal biometric authentication to exploit their complementary advantages and mitigate attacks. Finally the method incorporates the two modalities via an attention fusion mechanism to highlight the multimodal biometric recognition system reliability under varying circumstances. Experimental results on the MSU-AVIS dataset demonstrate the efficiency of our method, showing notable improvements achieving overall multimodal biometric recognition accuracy by 97% and improving the peak signal to noise ratio by 48.25dB. These findings demonstrate that modality-aware fusion using MTCP-VAFT method can delivery secure and flexible biometric authentication suitable for deployment on high-security platforms.

Downloads

Download data is not yet available.

Downloads

Published

28-07-2026

Issue

Section

Research article

How to Cite

Multimodal Biometric Recognition Based on Face and Video Via Vision Attention Fusion Transformer. (2026). The Scientific Temper, 17(07), 6599-6617. https://doi.org/10.58414/SCIENTIFICTEMPER.2026.17.7.2498

Similar Articles

1-10 of 442

You may also start an advanced similarity search for this article.