IDENTITY-REFERENCED TWO-STAGE DEEPFAKE DETECTION WITH CONSISTENCY-GATED FUSION

Authors

DOI:

https://doi.org/10.35631/JISTM.1144053

Keywords:

Deepfake Detection, Identity Reference, Face Manipulation Detection, Cross-Manipulation Generalization, Cross-Dataset Generalization, Consistency-Gated Fusion

Abstract

Deepfake detection methods that rely on manipulation-specific artifacts tend to overfit the forgery types observed during training and generalize poorly to unseen manipulations and datasets. In this paper, we propose an identity-referenced two-stage framework that detects face forgeries by comparing a test face against a reference face of the claimed identity, rather than by memorizing generator-specific artifacts. Unlike earlier identity-referenced detectors, which assume that the original identity that was replaced is known at test time and fine-tune the identity encoder, our framework relies only on the claimed identity, which is available in deployment (for example from an ID photo), and keeps the identity encoder frozen. In the first stage, the frozen MobileFaceNet encoder, pre-trained with ArcFace for face recognition, produces an identity-consistency score between the reference and the test face. In the second stage, a trained EfficientNet-B0 encoder extracts forgery artifact. The two scores are then combined through a consistency-gated fusion, which combines the two scores equally when they are consistent and progressively suppresses the forgery score when they strongly disagree. Experiments on FaceForensics++ and Celeb-DF show that the identity-consistency signal generalizes across manipulation types and datasets: on the most challenging cross-manipulation case (FaceSwap), the identity reference attains 77.8% video-level AUC while the forgery branch collapses to 43.4%, and the fusion attains 67.7%, outperforming the classical baselines by 7 to 20 percentage points on FaceSwap. In cross-dataset evaluation on Celeb-DF, the fusion reaches 91.3% video-level AUC, competitive with strong recent detectors while using only 6.0M parameters. These results demonstrate that identity consistency is a robust prior for identity-swapping manipulations, and that using it directly as a standalone detection score, combined with an independent artifact branch through a validation-defined gate, is an effective and lightweight route to cross-manipulation and cross-dataset generalization.

Downloads

Download data is not yet available.

References

Chai, L., Bau, D., Lim, S.-N., & Isola, P. (2020). What makes fake images detectable? Understanding properties that generalize. In Computer Vision—ECCV 2020 (pp. 103–120). Springer. https://doi.org/10.1007/978-3-030-58574-7_7

Chen, S., Liu, Y., Gao, X., & Han, Z. (2018). MobileFaceNets: Efficient CNNs for accurate real-time face verification on mobile devices. In Biometric Recognition (pp. 428–438). Springer. https://doi.org/10.1007/978-3-319-97909-0_46

Chollet, F. (2017). Xception: Deep learning with depthwise separable convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 1800–1807). IEEE. https://doi.org/10.1109/CVPR.2017.195

Deng, J., Guo, J., Xue, N., & Zafeiriou, S. (2019). ArcFace: Additive angular margin loss for deep face recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 4685–4694). IEEE. https://doi.org/10.1109/CVPR.2019.00482

Dolhansky, B., Bitton, J., Pflaum, B., Lu, J., Howes, R., Wang, M., & Canton Ferrer, C. (2020). The DeepFake Detection Challenge (DFDC) dataset. arXiv. https://arxiv.org/abs/2006.07397

Haliassos, A., Mira, R., Petridis, S., & Pantic, M. (2022). Leveraging real talking faces via self-supervision for robust forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 14930–14942). IEEE. https://doi.org/10.1109/CVPR52688.2022.01453

Haliassos, A., Vougioukas, K., Petridis, S., & Pantic, M. (2021). Lips don’t lie: A generalisable and robust approach to face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 5037–5047). IEEE. https://doi.org/10.1109/CVPR46437.2021.00500

Li, L., Bao, J., Zhang, T., Yang, H., Chen, D., Wen, F., & Guo, B. (2020). Face X-ray for more general face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 5000–5009). IEEE. https://doi.org/10.1109/CVPR42600.2020.00505

Li, Y., Yang, X., Sun, P., Qi, H., & Lyu, S. (2020). Celeb-DF: A large-scale challenging dataset for deepfake forensics. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 3204–3213). IEEE. https://doi.org/10.1109/CVPR42600.2020.00327

Qian, Y., Yin, G., Sheng, L., Chen, Z., & Shao, J. (2020). Thinking in frequency: Face forgery detection by mining frequency-aware clues. In Computer Vision—ECCV 2020 (pp. 86–103). Springer. https://doi.org/10.1007/978-3-030-58610-2_6

Rössler, A., Cozzolino, D., Verdoliva, L., Riess, C., Thies, J., & Nießner, M. (2019). FaceForensics++: Learning to detect manipulated facial images. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 1–11). IEEE. https://doi.org/10.1109/ICCV.2019.00009

Sabir, E., Cheng, J., Jaiswal, A., AbdAlmageed, W., Masi, I., & Natarajan, P. (2019). Recurrent convolutional strategies for face manipulation detection in videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (pp. 80–87). IEEE. https://doi.org/10.48550/arXiv.1905.00582

Shen, D., Zhao, Y., & Quan, C. (2022). Identity-referenced deepfake detection with contrastive learning. In Proceedings of the 2022 ACM Workshop on Information Hiding and Multimedia Security (pp. 27–32). ACM. https://doi.org/10.1145/3531536.3532964

Shi, L., Zhang, J., Ji, Z., Bai, J., & Shan, S. (2025). Real face foundation representation learning for generalized deepfake detection. Pattern Recognition, 161, 111299. https://doi.org/10.1016/j.patcog.2024.111299

Shiohara, K., & Yamasaki, T. (2022). Detecting deepfakes with self-blended images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 18699–18708). IEEE. https://doi.org/10.1109/CVPR52688.2022.01816

Tan, M., & Le, Q. V. (2019). EfficientNet: Rethinking model scaling for convolutional neural networks. In Proceedings of the 36th International Conference on Machine Learning (pp. 6105–6114). PMLR. https://doi.org/10.48550/arXiv.1905.11946

Wang, S.-Y., Wang, O., Zhang, R., Owens, A., & Efros, A. A. (2020). CNN-generated images are surprisingly easy to spot… for now. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 8692–8701). IEEE. https://doi.org/10.1109/CVPR42600.2020.00872

Zhao, T., Xu, X., Xu, M., Ding, H., Xiong, Y., & Xia, W. (2021). Learning self-consistency for deepfake detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 15003–15013). IEEE. https://doi.org/10.1109/ICCV48922.2021.01475

Zheng, Y., Bao, J., Chen, D., Zeng, M., & Wen, F. (2021). Exploring temporal coherence for more general video face forgery detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 15024–15034). IEEE. https://doi.org/10.1109/ICCV48922.2021.01477

Zhuang, W., Chu, Q., Tan, Z., Liu, Q., Yuan, H., Miao, C., Luo, Z., & Yu, N. (2022). UIA-ViT: Unsupervised inconsistency-aware method based on vision transformer for face forgery detection. In Computer Vision—ECCV 2022 (pp. 391–407). Springer. https://doi.org/10.1007/978-3-031-20065-6_23

Downloads

Published

2026-09-30

How to Cite

Zhen , Y., & Jamaluddin, R. A. (2026). IDENTITY-REFERENCED TWO-STAGE DEEPFAKE DETECTION WITH CONSISTENCY-GATED FUSION. JOURNAL INFORMATION AND TECHNOLOGY MANAGEMENT (JISTM), 11(44), 904–916. https://doi.org/10.35631/JISTM.1144053