SearcharxivSearch

arXiv subjects

Ravi Kothari

Publications and source records attributed to Ravi Kothari.

3 recordsLinked to original sources

Self-Supervised Multimodal NeRF for Autonomous Driving

In this paper, we propose a Neural Radiance Fields (NeRF) based framework, referred to as Novel View Synthesis Framework (NVSF). It jointly learns the implicit neural representation of space and time-varying scene for both LiDAR and Camera. We test this on a real-world autonomous driving scenario containing both static and dynamic scenes. Compared to existing multimodal dynamic NeRFs, our framework is self-supervised, thus eliminating the need for 3D labels. For efficient training and faster convergence, we introduce heuristic-based image pixel sampling to focus on pixels with rich information. To preserve the local features of LiDAR points, a Double Gradient based mask is employed. Extensive experiments on the KITTI-360 dataset show that, compared to the baseline models, our framework has reported best performance on both LiDAR and Camera domain. Code of the model is available at https://github.com/gaurav00700/Selfsupervised-NVSF

cs.CV

Raw Radar data based Object Detection and Heading estimation using Cross Attention

Radar is an inevitable part of the perception sensor set for autonomous driving functions. It plays a gap-filling role to complement the shortcomings of other sensors in diverse scenarios and weather conditions. In this paper, we propose a Deep Neural Network (DNN) based end-to-end object detection and heading estimation framework using raw radar data. To this end, we approach the problem in both a Data-centric and model-centric manner. We refine the publicly available CARRADA dataset and introduce Bivariate norm annotations. Besides, the baseline model is improved by a transformer inspired cross-attention fusion and further center-offset maps are added to reduce localisation error. Our proposed model improves the detection mean Average Precision (mAP) by 5%, while reducing the model complexity by almost 23%. For comprehensive scene understanding purposes, we extend our model for heading estimation. The improved ground truth and proposed model is available at Github

cs.IT

A Parameterized Approach to Personalized Variable Length Summarization of Soccer Matches

We present a parameterized approach to produce personalized variable length summaries of soccer matches. Our approach is based on temporally segmenting the soccer video into 'plays', associating a user-specifiable 'utility' for each type of play and using 'bin-packing' to select a subset of the plays that add up to the desired length while maximizing the overall utility (volume in bin-packing terms). Our approach systematically allows a user to override the default weights assigned to each type of play with individual preferences and thus see a highly personalized variable length summarization of soccer matches. We demonstrate our approach based on the output of an end-to-end pipeline that we are building to produce such summaries. Though aspects of the overall end-to-end pipeline are human assisted at present, the results clearly show that the proposed approach is capable of producing semantically meaningful and compelling summaries. Besides the obvious use of producing summaries of superior league matches for news broadcasts, we anticipate our work to promote greater awareness of the local matches and junior leagues by producing consumable summaries of them.

cs.CV