Visual Reasoning at Urban Intersections: Fine-Tuning GPT-4O for Traffic Conflict Detection
| dc.contributor.author | Masri, Sari | |
| dc.contributor.author | Ashqar, Huthaifa | |
| dc.contributor.author | Elhenawy, Mohammed | |
| dc.date.accessioned | 2025-10-16T15:27:12Z | |
| dc.date.issued | 2025-09-08 | |
| dc.description | 2025 IEEE 4th International Conference on Computing and Machine Intelligence (ICMI), April 5-6, 2025, Mount Pleasant, MI, | |
| dc.description.abstract | Traffic control in unsignalized urban intersections presents significant challenges due to the complexity, frequent conflicts, and blind spots. This study explores the capability of leveraging Multimodal Large Language Models (MLLMs), such as GPT-4o, to provide logical and visual reasoning by directly using birds-eye-view videos of four-legged intersections. In this proposed method, GPT-4o acts as intelligent system to detect conflicts and provide explanations and recommendations for the drivers. The fine-tuned model achieved an accuracy of 77.14 %, while the manual evaluation of the true predicted values of the fine-tuned GPT-4o showed significant achievements of 89.9 % accuracy for model-generated explanations and 92.3 % for the recommended next actions. These results highlight the feasibility of using MLLMs for real-time traffic management using videos as inputs, offering scalable and actionable insights into intersections traffic management and operation. Code used in this study is available at https://github.com/sarimasri3/Traffic-Intersection-Conflict-Detection-using-images.git. | |
| dc.description.uri | https://ieeexplore.ieee.org/abstract/document/11141198 | |
| dc.format.extent | 5 pages | |
| dc.genre | conference papers and proceedings | |
| dc.genre | preprints | |
| dc.identifier | doi:10.13016/m2udaf-n2pf | |
| dc.identifier.citation | Masri, Sari, Huthaifa I. Ashqar, and Mohammed Elhenawy. “Visual Reasoning at Urban Intersections: Fine-Tuning GPT-4O for Traffic Conflict Detection.” 2025 IEEE 4th International Conference on Computing and Machine Intelligence (ICMI), April 2025, 1–5. https://doi.org/10.1109/ICMI65310.2025.11141198. | |
| dc.identifier.uri | https://doi.org/10.1109/ICMI65310.2025.11141198 | |
| dc.identifier.uri | http://hdl.handle.net/11603/40454 | |
| dc.language.iso | en | |
| dc.publisher | IEEE | |
| dc.relation.isAvailableAt | The University of Maryland, Baltimore County (UMBC) | |
| dc.relation.ispartof | UMBC Data Science | |
| dc.rights | © 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works. | |
| dc.subject | Accuracy | |
| dc.subject | Machine intelligence | |
| dc.subject | Conflict Detection | |
| dc.subject | Manuals | |
| dc.subject | Intersection Management | |
| dc.subject | Real-time systems | |
| dc.subject | Predictive models | |
| dc.subject | Videos | |
| dc.subject | Visualization | |
| dc.subject | Traffic control | |
| dc.subject | Cognition | |
| dc.subject | Intersections | |
| dc.subject | Large language models | |
| dc.subject | Multimodal Large Language Models (MLLMs) | |
| dc.title | Visual Reasoning at Urban Intersections: Fine-Tuning GPT-4O for Traffic Conflict Detection | |
| dc.type | Text | |
| dcterms.creator | https://orcid.org/0000-0002-6835-8338 |
Files
Original bundle
1 - 1 of 1
