Please use this identifier to cite or link to this item: http://repository.iiitd.edu.in/xmlui/handle/123456789/2170
Title: Multimodal stream processing for accident detection
Authors: Singh, Vishal
Shah, Rajiv Ratn (Advisor)
Keywords: Accident Anticipation
Monocular Depth
3D Modeling
Autonomous Driving
Issue Date: 26-Nov-2024
Publisher: IIIT-Delhi
Abstract: This report presents a novel architecture for accident anticipation and localization in autonomous driving, integrating monocular depth-enhanced 3D modeling with large language models (LLMs). The system leverages advanced feature extraction techniques, such as MobileNetv2 and Cascade R-CNN, combined with attention mechanisms to process real-time dashcam video inputs. A dual vision attention mechanism enhances feature representation, while a dynamic object attention mechanism prioritizes high-risk objects. The architecture incorporates a GRU-based accident anticipation module and an attention-driven accident localization module to predict accident probabilities and identify hazardous objects. Finally, the model generates real-time verbal alerts using LLMs, offering contextually relevant warnings to passengers, enhancing safety and aware ness. This hybrid approach aims to improve both accident prediction accuracy and anticipation time, making it suitable for deployment in real-world autonomous driving systems. The model’s effectiveness is evaluated using metrics like average precision (AP) and mean Time to Accident (mTTA).
URI: http://repository.iiitd.edu.in/xmlui/handle/123456789/2170
Appears in Collections:Year-2024

Files in This Item:
File Description SizeFormat 
BTP_Report_2021575 - Vishal Singh.pdf
  Restricted Access
197.91 kBAdobe PDFView/Open Request a copy


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.