A Lightweight Hybrid CNN-BiLSTM Model with Adaptive Temporal Attention for Violence Detection in Surveillance Videos

Authors

  • Aicha Khalfaoui
  • Abdelmajid Badri
  • Ilham El Mourabit

DOI:

https://doi.org/10.47839/ijc.25.2.4652

Keywords:

violence detection, edge devices, adaptive temporal attention, MobileNetV3-S, Bi-LSTM

Abstract

Detecting violent actions in surveillance videos is a critical task for ensuring public safety in smart city environments. Although numerous deep learning approaches have been proposed, most of them favor either accuracy or computational efficiency, but rarely achieve both simultaneously. This limitation restricts their deployment on resource-constrained edge devices. In this paper, we propose a lightweight yet effective hybrid architecture that combines MobileNetV3-S for spatial feature extraction with a BiLSTM enhanced by an Adaptive Temporal Attention mechanism for temporal modeling. Despite its compact design (1.00 GFLOPs and 3.66 million parameters), the proposed model achieves competitive performance on four public benchmark datasets. The experimental results demonstrate that the proposed approach provides an excellent trade-off between accuracy and efficiency, making it suitable for real-time smart surveillance applications.

References

H. M. B. Jahlan and L. A. Elrefaei, “Detecting violence in video based on deep features fusion technique,” arXiv preprint arXiv:2204.07443, 2022.

A. Howard, M. Sandler, G. Chu, L.-C. Chen, B. Chen, M. Tan, W. Wang, Y. Zhu, R. Pang, V. Vasudevan et al., “Searching for mobilenetv3,” in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 1314–1324.

A. Traoré and M. A. Akhloufi, “Violence detection in videos using deep recurrent and convolutional neural networks,” in 2020 IEEE international conference on systems, man, and cybernetics (SMC). IEEE, 2020, pp. 154–159.

A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems, vol. 30, 2017.

N. Dalal and B. Triggs, “Histograms of oriented gradients for human detection,” in 2005 IEEE computer society conference on computer vision and pattern recognition (CVPR’05), vol. 1. Ieee, 2005, pp. 886–893.

D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, “Learning spatiotemporal features with 3d convolutional networks,” in Proceedings of the IEEE international conference on computer vision, 2015, pp. 4489– 4497.

K. Simonyan and A. Zisserman, “Two-stream convolutional networks for action recognition in videos,” Advances in neural information processing systems, vol. 27, 2014.

J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell, “Long-term recurrent convolutional networks for visual recognition and description,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 2625–2634.

S. Sharma, B. Sudharsan, S. Naraharisetti, V. Trehan, and K. Jayavel, “A fully integrated violence detection system using cnn and lstm.” International Journal of Electrical & Computer Engineering (2088-8708), vol. 11, no. 4, 2021.

X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 7794–7803.

M.-S. Kang, R.-H. Park, and H.-M. Park, “Efficient spatio-temporal modeling methods for real-time violence recognition,” IEEE Access, vol. 9, pp. 76 270–76 285, 2021.

H. Chen, X. Mei, Z. Ma, X. Wu, and Y. Wei, “Spatial–temporal graph attention network for video anomaly detection,” Image and Vision Computing, vol. 131, p. 104629, 2023.

S. Hochreiter, “The vanishing gradient problem during learning recurrent neural nets and problem solutions,” International Journal of Uncertainty, Fuzziness and Knowledge-Based Systems, vol. 6, no. 02, pp. 107–116, 1998.

Z. Liu, L. Wang, W. Wu, C. Qian, and T. L. Tam, “Temporal adaptive module for video recognition. in 2021 ieee,” in CVF International Conference on Computer Vision (ICCV), 2021, pp. 13 688–13 698.

¸S. Aktı, G. A. Tataroglu, and H. K. Ekenel, “Vision-based fight detection ˘ from surveillance cameras,” in 2019 Ninth International Conference on Image Processing Theory, Tools and Applications (IPTA). IEEE, 2019, pp. 1–6.

E. Bermejo Nievas, O. Deniz Suarez, G. Bueno García, and R. Sukthankar, “Violence detection in video using computer vision techniques,” in Computer Analysis of Images and Patterns: 14th International Conference, CAIP 2011, Seville, Spain, August 29-31, 2011, Proceedings, Part II 14. Springer, 2011, pp. 332–339.

M. Cheng, K. Cai, and M. Li, “Rwf-2000: An open large scale video database for violence detection,” in 2020 25th International Conference on Pattern Recognition (ICPR). IEEE, 2021, pp. 4183–4190.

S. Sudhakaran and O. Lanz, “Learning to detect violent videos using convolutional long short-term memory,” in 2017 14th IEEE international conference on advanced video and signal based surveillance (AVSS). IEEE, 2017, pp. 1–6.

R. Halder and R. Chatterjee, “Cnn-bilstm model for violence detection in smart surveillance,” SN Computer science, vol. 1, no. 4, p. 201, 2020.

L. Ciampi, C. Santiago, J. P. Costeira, F. Falchi, C. Gennaro, and G. Amato, “Unsupervised domain adaptation for video violence detection in the wild.” in IMPROVE, 2023, pp. 37–46.

Q. Liang, Y. Li, B. Chen, and K. Yang, “Violence behavior recognition of two-cascade temporal shift module with attention mechanism,” Journal of Electronic Imaging, vol. 30, no. 4, pp. 043 009–043 009, 2021.

D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri, “A closer look at spatiotemporal convolutions for action recognition,” in Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, 2018, pp. 6450–6459.

K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778.

M. Khan, A. El Saddik, W. Gueaieb, G. De Masi, and F. Karray, “Vd-net: An edge vision-based surveillance system for violence detection,” IEEE Access, vol. 12, pp. 43 796–43 808, 2024.

C. Feichtenhofer, H. Fan, J. Malik, and K. He, “Slowfast networks for video recognition,” in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 6202–6211.

J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,” in proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 6299–6308.

Downloads

Published

2026-06-30

How to Cite

Khalfaoui, A., Badri, A., & El Mourabit, I. (2026). A Lightweight Hybrid CNN-BiLSTM Model with Adaptive Temporal Attention for Violence Detection in Surveillance Videos. International Journal of Computing, 25(2), 262-272. https://doi.org/10.47839/ijc.25.2.4652

Issue

Section

Articles