Detecting hate crimes through machine learning and natural language processing.
Misidentification and misreporting of hate crimes by victims and law enforcement are significant barriers to accurate data collection of hate crimes, and their consequent study and prevention. The use of machine learning in crime detection can improve the accuracy and speed at which reported inciden...
| Publicado en: | Police Practice & Research Vol. 26; no. 6; pp. 746 - 769 |
|---|---|
| Autor principal: | |
| Formato: | Artículo |
| Publicado: |
Taylor & Francis Ltd
Oct2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ssf&AN=188232690&site=ehost-live header: @attributes: shortDbName: ssf uiTerm: 188232690 longDbName: Social Sciences Full Text (H.W. Wilson) uiTag: AN controlInfo: bkinfo: jinfo: jid: 15614263 J5F jtl: Police Practice & Research issn: 15614263 maglogo: N pubinfo: dt: Oct2025 vid: 26 iid: 6 pid: 377 pub: Taylor & Francis Ltd artinfo: ui: 188232690 10.1080/15614263.2024.2397363 ppf: 746 ppct: 23 formats: tig: atl: Detecting hate crimes through machine learning and natural language processing. aug: au: Ortiz Salazar, Ana affil: Lead Data Scientist and Bias Crime Research Scientist, Seattle Police Department – Performance Analytics and Research, Seattle, WA, USA su: Hate crimes Criminal investigation Machine learning Natural language processing Classification algorithms Police reports Acquisition of data sug: subj: Hate crimes Criminal investigation Machine learning Natural language processing Classification algorithms Police reports Acquisition of data keyword: bias machine learning NLP Seattle bias machine learning NLP Seattle ab: Misidentification and misreporting of hate crimes by victims and law enforcement are significant barriers to accurate data collection of hate crimes, and their consequent study and prevention. The use of machine learning in crime detection can improve the accuracy and speed at which reported incidents with bias elements are identified. This study develops a machine learning classifier that categorizes police reports as either events with bias elements or events with no bias elements. We use incident/offense reports from the Seattle Police Department to train a Natural Language Processing classification algorithm. We collect narratives, location data, and victim and suspect demographics to use as features. We evaluate the performance of logistic regression, random forest, and XGBoost algorithms, as well as several text embedding techniques. Despite substantial class imbalance, our model achieves a macro F1-score of 0.79, demonstrating the benefits of applied machine learning in accurately detecting and reporting hate crimes. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: N holdings: @attributes: islocal: N |
|---|