Prioritizing Recall: Recall‑First Machine Learning for Traffic Accident Severity Detection under Class Imbalance

Authors

  • Geping Cai

DOI:

https://doi.org/10.61173/ym16tb60

Keywords:

Traffic Crash Risk Prediction, Class Imbalance, Recall-First Evaluation, Machine Learning Classification, Cost-Sensitive Decision Thresholds, Public-Safety Operations

Abstract

Traffic safety agencies seek predictive models that identify high-risk crashes before traffic accidents occur, but realworld crash data are highly imbalanced, rendering the overall accuracy of machine learning models a misleading metric for predictive performance. This study investigates a recall-first approach to crash-risk classification in Virginia, arguing that missing a dangerous event carries a much higher cost than issuing an extra alert based on the data provided by Virginia Department of Transportation (VDOT). The modeling framed high-risk identification as a binary classification task and used standard tools, Logistic Regression and Random Forest, augmented with imbalance-aware training. Evaluation emphasized recall and precision, utilizing precision–recall curves and confusion matrices. Across experiments, recall-oriented training and thresholding consistently improved detection of high-risk cases relative to accuracy-optimized baselines, with expected trade-offs resulting in lower precision and overall accuracy. This paper further translates operational priorities into simple threshold policies, showing how agencies can tune models based on recall to align with resource constraints while minimizing costly misses. In conclusion, for safety-critical applications, recall should be the primary metric of success; furthermore, straightforward imbalance treatments—without complex model additions—can realign model performance with public-safety goals.

References

[1] Santos, D., Saias, J., Quaresma, P., & Nogueira, V. B. (2021). Machine Learning Approaches to Traffic Accident Analysis and Hotspot Prediction. Computers, 10(12), 157. https://doi. org/10.3390/computers10120157

[2] Ghanem, M., Ghaith, A. K., El-Hajj, V. G., Bhandarkar, A., de Giorgio, A., Elmi-Terander, A., & Bydon, M. (2023). Limitations in Evaluating Machine Learning Models for Imbalanced Binary Outcome Classification in Spine Surgery: A Systematic Review. Brain sciences, 13(12), 1723. https://doi. org/10.3390/brainsci13121723

[3] Virginia Department of Transportation. Full Crash Dataset. https://www.virginiaroads.org/datasets/VDOT::full-crash-1/expl ore?layer=0&location=0.003769%2C-79.499811%2C0.00&sho wTable=true.

[4] FHWA Safety Program. (n.d.-a). Crash costs for highway safety analysis. U.S. Department of Transportation Federal Highway Administration. https://safety.fhwa.dot.gov/hsip/docs/ fhwasa17071.pdf

[5] U.S. Department of Transportation. (2004, August). Signalized intersections: Informational guide. FHWA. https:// www.fhwa.dot.gov/publications/research/safety/04091/

[6] What is a safe system approach?. U.S. Department of Transportation. (2025, January 14). https://www.transportation. gov/safe-system-approach

[7] EDC-5 Innovations (2019-2020). EDC-5 Innovations | Federal Highway Administration. (2023, May 25). https://www. fhwa.dot.gov/innovation/everydaycounts/edc_5/

[8] Blincoe, L., Miller, T., Wang, J.-S., Swedler, D., Coughlin, T., Lawrence, B., Guo, F., Klauer, S., & Dingus, T. (2023, February). The economic and societal impact of motor vehicle crashes, 2019 (Revised) (Report No. DOT HS 813 403). National Highway Traffic Safety Administration.

[9] Saito, T., & Rehmsmeier, M. (2015). The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PloS one, 10(3), e0118432.

[10] Davis, J., & Goadrich, M. (2006, June). The relationship between Precision-Recall and ROC curves. In Proceedings of the 23rd international conference on Machine learning (pp. 233- 240).

[11] López, V., Fernández, A., García, S., Palade, V., & Herrera, F. (2013). An insight into classification with imbalanced data: Empirical results and current trends on using data intrinsic characteristics. Information sciences, 250, 113-141.

[12] Chen, C., Liaw, A., & Breiman, L. (2004). Using random forest to learn imbalanced data. University of California, Berkeley, 110(1-12), 24.

Downloads

Published

2025-12-19