COMPARATIVE STUDY OF VIT WITH LSTM AND VIT WITH TCN FOR DEEP FAKE VIDEO DETECTION

BUDIONO, MAXIMILLIAN MALEAKHI (2026) COMPARATIVE STUDY OF VIT WITH LSTM AND VIT WITH TCN FOR DEEP FAKE VIDEO DETECTION. Project Report. UNIVERSITAS KATOLIK SOEGIJAPRANATA, Semarang. (Unpublished)

[img]
Preview
Text
22.K3.0001-MAXIMILLIAN MALEAKHI BUDIONO_cover.pdf

Download (1MB) | Preview
[img] Text
22.K3.0001-MAXIMILLIAN MALEAKHI BUDIONO_isi.pdf
Restricted to Registered users only

Download (3MB)
[img]
Preview
Text
22.K3.0001-MAXIMILLIAN MALEAKHI BUDIONO_dapus.pdf

Download (1MB) | Preview
[img] Text
22.K3.0001-MAXIMILLIAN MALEAKHI BUDIONO_lamp.pdf
Restricted to Registered users only

Download (1MB)

Abstract

Deepfake videos are manipulated videos generated by artificial intelligence, deepfakes cause some dangerous problems, they can make people look like they are doing something they have not actually done, spread misinformation, defamation, and digital fraud. Therefore, it is important to develop a deepfake video detector with high stability and accuracy. In this study, we present a comparative study of 1 baseline model and 2 hybrid models, namely, Vision Transformer (ViT) combined with Long Short Term Memory (LSTM), and Vision Transformer (ViT) combined with Temporal Convolutional Networks (TCN). In order to assess its effects on detection performance, these hybrid models use the same ViT for feature extraction. UADFV and FF++ dataset was used, each video extracted into 32 sequence frames, These frames were passed into ViT, resulting in a 768 dimensional feature vector per frame. These extracted feature are then fed into temporal models LSTM or TCN. Both models evaluate using binary classification accuracy, binary cross entropy loss, and visualization metrics especially AUC ROC curve. Results indicate that ViT + TCN has higher AUC value with 0.7228 and more stable training than Vit baseline with 0.6741 and ViT + LSTM with 0.5833 AUC values. This study provides insight of the trade off between temporal modeling in real world scenarios.

Item Type: Monograph (Project Report)
Subjects: 000 Computer Science, Information and General Works > 004 Data processing & computer science
Divisions: Faculty of Computer Science > Department of Informatics Engineering
Depositing User: mr Dwi Purnomo
Date Deposited: 19 Jun 2026 05:02
Last Modified: 19 Jun 2026 05:02
URI: http://repository.unika.ac.id/id/eprint/40041
Keywords: Deepfake Detection, Vision Transformer, LSTM, TCN, Temporal Modeling

Actions (login required)

View Item View Item