IDENTITY-REFERENCED TWO-STAGE DEEPFAKE DETECTION WITH CONSISTENCY-GATED FUSION
DOI:
https://doi.org/10.35631/JISTM.1144053Keywords:
Deepfake Detection, Identity Reference, Face Manipulation Detection, Cross-Manipulation Generalization, Cross-Dataset Generalization, Consistency-Gated FusionAbstract
Deepfake detection methods that rely on manipulation-specific artifacts tend to overfit the forgery types observed during training and generalize poorly to unseen manipulations and datasets. In this paper, we propose an identity-referenced two-stage framework that detects face forgeries by comparing a test face against a reference face of the claimed identity, rather than by memorizing generator-specific artifacts. Unlike earlier identity-referenced detectors, which assume that the original identity that was replaced is known at test time and fine-tune the identity encoder, our framework relies only on the claimed identity, which is available in deployment (for example from an ID photo), and keeps the identity encoder frozen. In the first stage, the frozen MobileFaceNet encoder, pre-trained with ArcFace for face recognition, produces an identity-consistency score between the reference and the test face. In the second stage, a trained EfficientNet-B0 encoder extracts forgery artifact. The two scores are then combined through a consistency-gated fusion, which combines the two scores equally when they are consistent and progressively suppresses the forgery score when they strongly disagree. Experiments on FaceForensics++ and Celeb-DF show that the identity-consistency signal generalizes across manipulation types and datasets: on the most challenging cross-manipulation case (FaceSwap), the identity reference attains 77.8% video-level AUC while the forgery branch collapses to 43.4%, and the fusion attains 67.7%, outperforming the classical baselines by 7 to 20 percentage points on FaceSwap. In cross-dataset evaluation on Celeb-DF, the fusion reaches 91.3% video-level AUC, competitive with strong recent detectors while using only 6.0M parameters. These results demonstrate that identity consistency is a robust prior for identity-swapping manipulations, and that using it directly as a standalone detection score, combined with an independent artifact branch through a validation-defined gate, is an effective and lightweight route to cross-manipulation and cross-dataset generalization.
Downloads
References
Chai, L., Bau, D., Lim, S.-N., & Isola, P. (2020). What makes fake images detectable? Understanding properties that generalize. In Computer Vision—ECCV 2020 (pp. 103–120). Springer. https://doi.org/10.1007/978-3-030-58574-7_7
Chen, S., Liu, Y., Gao, X., & Han, Z. (2018). MobileFaceNets: Efficient CNNs for accurate real-time face verification on mobile devices. In Biometric Recognition (pp. 428–438). Springer. https://doi.org/10.1007/978-3-319-97909-0_46
Chollet, F. (2017). Xception: Deep learning with depthwise separable convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 1800–1807). IEEE. https://doi.org/10.1109/CVPR.2017.195
Deng, J., Guo, J., Xue, N., & Zafeiriou, S. (2019). ArcFace: Additive angular margin loss for deep face recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 4685–4694). IEEE. https://doi.org/10.1109/CVPR.2019.00482
Dolhansky, B., Bitton, J., Pflaum, B., Lu, J., Howes, R., Wang, M., & Canton Ferrer, C. (2020). The DeepFake Detection Challenge (DFDC) dataset. arXiv. https://arxiv.org/abs/2006.07397
Haliassos, A., Mira, R., Petridis, S., & Pantic, M. (2022). Leveraging real talking faces via self-supervision for robust forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 14930–14942). IEEE. https://doi.org/10.1109/CVPR52688.2022.01453
Haliassos, A., Vougioukas, K., Petridis, S., & Pantic, M. (2021). Lips don’t lie: A generalisable and robust approach to face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 5037–5047). IEEE. https://doi.org/10.1109/CVPR46437.2021.00500
Li, L., Bao, J., Zhang, T., Yang, H., Chen, D., Wen, F., & Guo, B. (2020). Face X-ray for more general face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 5000–5009). IEEE. https://doi.org/10.1109/CVPR42600.2020.00505
Li, Y., Yang, X., Sun, P., Qi, H., & Lyu, S. (2020). Celeb-DF: A large-scale challenging dataset for deepfake forensics. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 3204–3213). IEEE. https://doi.org/10.1109/CVPR42600.2020.00327
Qian, Y., Yin, G., Sheng, L., Chen, Z., & Shao, J. (2020). Thinking in frequency: Face forgery detection by mining frequency-aware clues. In Computer Vision—ECCV 2020 (pp. 86–103). Springer. https://doi.org/10.1007/978-3-030-58610-2_6
Rössler, A., Cozzolino, D., Verdoliva, L., Riess, C., Thies, J., & Nießner, M. (2019). FaceForensics++: Learning to detect manipulated facial images. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 1–11). IEEE. https://doi.org/10.1109/ICCV.2019.00009
Sabir, E., Cheng, J., Jaiswal, A., AbdAlmageed, W., Masi, I., & Natarajan, P. (2019). Recurrent convolutional strategies for face manipulation detection in videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (pp. 80–87). IEEE. https://doi.org/10.48550/arXiv.1905.00582
Shen, D., Zhao, Y., & Quan, C. (2022). Identity-referenced deepfake detection with contrastive learning. In Proceedings of the 2022 ACM Workshop on Information Hiding and Multimedia Security (pp. 27–32). ACM. https://doi.org/10.1145/3531536.3532964
Shi, L., Zhang, J., Ji, Z., Bai, J., & Shan, S. (2025). Real face foundation representation learning for generalized deepfake detection. Pattern Recognition, 161, 111299. https://doi.org/10.1016/j.patcog.2024.111299
Shiohara, K., & Yamasaki, T. (2022). Detecting deepfakes with self-blended images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 18699–18708). IEEE. https://doi.org/10.1109/CVPR52688.2022.01816
Tan, M., & Le, Q. V. (2019). EfficientNet: Rethinking model scaling for convolutional neural networks. In Proceedings of the 36th International Conference on Machine Learning (pp. 6105–6114). PMLR. https://doi.org/10.48550/arXiv.1905.11946
Wang, S.-Y., Wang, O., Zhang, R., Owens, A., & Efros, A. A. (2020). CNN-generated images are surprisingly easy to spot… for now. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 8692–8701). IEEE. https://doi.org/10.1109/CVPR42600.2020.00872
Zhao, T., Xu, X., Xu, M., Ding, H., Xiong, Y., & Xia, W. (2021). Learning self-consistency for deepfake detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 15003–15013). IEEE. https://doi.org/10.1109/ICCV48922.2021.01475
Zheng, Y., Bao, J., Chen, D., Zeng, M., & Wen, F. (2021). Exploring temporal coherence for more general video face forgery detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 15024–15034). IEEE. https://doi.org/10.1109/ICCV48922.2021.01477
Zhuang, W., Chu, Q., Tan, Z., Liu, Q., Yuan, H., Miao, C., Luo, Z., & Yu, N. (2022). UIA-ViT: Unsupervised inconsistency-aware method based on vision transformer for face forgery detection. In Computer Vision—ECCV 2022 (pp. 391–407). Springer. https://doi.org/10.1007/978-3-031-20065-6_23
