Violence Detection in Audio: Evaluating the Effectiveness of Deep Learning Models and Data Augmentation

Durães, Dalila; Veloso, Bruno; Novais, Paulo

dc.creator	Durães, Dalila
dc.creator	Veloso, Bruno
dc.creator	Novais, Paulo
dc.date.accessioned	2023-09-06T08:26:38Z
dc.date.accessioned	2023-09-07T15:21:41Z
dc.date.available	2023-09-06T08:26:38Z
dc.date.available	2023-09-07T15:21:41Z
dc.date.created	2023-09-06T08:26:38Z
dc.identifier	1989-1660
dc.identifier	https://reunir.unir.net/handle/123456789/15217
dc.identifier	https://doi.org/10.9781/ijimai.2023.08.007
dc.identifier.uri	https://repositorioslatinoamericanos.uchile.cl/handle/2250/8732532
dc.description.abstract	Human nature is inherently intertwined with violence, impacting the lives of numerous individuals. Various forms of violence pervade our society, with physical violence being the most prevalent in our daily lives. The study of human actions has gained significant attention in recent years, with audio (captured by microphones) and video (captured by cameras) being the primary means to record instances of violence. While video requires substantial processing capacity and hardware-software performance, audio presents itself as a viable alternative, offering several advantages beyond these technical considerations. Therefore, it is crucial to represent audio data in a manner conducive to accurate classification. In the context of violence in a car, specific datasets dedicated to this domain are not readily available. As a result, we had to create a custom dataset tailored to this particular scenario. The purpose of curating this dataset was to assess whether it could enhance the detection of violence in car-related situations. Due to the imbalanced nature of the dataset, data augmentation techniques were implemented. Existing literature reveals that Deep Learning (DL) algorithms can effectively classify audio, with a commonly used approach involving the conversion of audio into a mel spectrogram image. Based on the results obtained for that dataset, the EfficientNetB1 neural network demonstrated the highest accuracy (95.06%) in detecting violence in audios, closely followed by EfficientNetB0 (94.19%). Conversely, MobileNetV2 proved to be less capable in classifying instances of violence.
dc.language	eng
dc.publisher	International Journal of Interactive Multimedia and Artificial Intelligence
dc.relation	;vol. 8, nº 3
dc.relation	https://www.ijimai.org/journal/bibcite/reference/3373
dc.rights	openAccess
dc.subject	audio
dc.subject	deep learning
dc.subject	human detection activity
dc.subject	human activity
dc.subject	machine learning
dc.subject	transfer learning
dc.subject	violence detection
dc.subject	IJIMAI
dc.title	Violence Detection in Audio: Evaluating the Effectiveness of Deep Learning Models and Data Augmentation
dc.type	article

Este ítem pertenece a la siguiente institución

Universidad internacional de la Rioja (Colombia)