SearcharxivSearch

arXiv subjects

Bo Yu

Publications and source records attributed to Bo Yu.

At least 73 records · Page 4Linked to original sources

Factor Graph Accelerator for LiDAR-Inertial Odometry

Factor graph is a graph representing the factorization of a probability distribution function, and has been utilized in many autonomous machine computing tasks, such as localization, tracking, planning and control etc. We are developing an architecture with the goal of using factor graph as a common abstraction for most, if not, all autonomous machine computing tasks. If successful, the architecture would provide a very simple interface of mapping autonomous machine functions to the underlying compute hardware. As a first step of such an attempt, this paper presents our most recent work of developing a factor graph accelerator for LiDAR-Inertial Odometry (LIO), an essential task in many autonomous machines, such as autonomous vehicles and mobile robots. By modeling LIO as a factor graph, the proposed accelerator not only supports multi-sensor fusion such as LiDAR, inertial measurement unit (IMU), GPS, etc., but solves the global optimization problem of robot navigation in batch or incremental modes. Our evaluation demonstrates that the proposed design significantly improves the real-time performance and energy efficiency of autonomous machine navigation systems. The initial success suggests the potential of generalizing the factor graph architecture as a common abstraction for autonomous machine computing, including tracking, planning, and control etc.

cs.RO

Brief Industry Paper: The Necessity of Adaptive Data Fusion in Infrastructure-Augmented Autonomous Driving System

This paper is the first to provide a thorough system design overview along with the fusion methods selection criteria of a real-world cooperative autonomous driving system, named Infrastructure-Augmented Autonomous Driving or IAAD. We present an in-depth introduction of the IAAD hardware and software on both road-side and vehicle-side computing and communication platforms. We extensively characterize the IAAD system in the context of real-world deployment scenarios and observe that the network condition that fluctuates along the road is currently the main technical roadblock for cooperative autonomous driving. To address this challenge, we propose new fusion methods, dubbed "inter-frame fusion" and "planning fusion" to complement the current state-of-the-art "intra-frame fusion". We demonstrate that each fusion method has its own benefit and constraint.

cs.DC

Dataflow Accelerator Architecture for Autonomous Machine Computing

Commercial autonomous machines is a thriving sector, one that is likely the next ubiquitous computing platform, after Personal Computers (PC), cloud computing, and mobile computing. Nevertheless, a suitable computing substrate for autonomous machines is missing, and many companies are forced to develop ad hoc computing solutions that are neither principled nor extensible. By analyzing the demands of autonomous machine computing, this article proposes Dataflow Accelerator Architecture (DAA), a modern instantiation of the classic dataflow principle, that matches the characteristics of autonomous machine software.

cs.AR

Robotic Computing on FPGAs: Current Progress, Research Challenges, and Opportunities

Robotic computing has reached a tipping point, with a myriad of robots (e.g., drones, self-driving cars, logistic robots) being widely applied in diverse scenarios. The continuous proliferation of robotics, however, critically depends on efficient computing substrates, driven by real-time requirements, robotic size-weight-and-power constraints, cybersecurity considerations, and dynamically changing scenarios. Within all platforms, FPGA is able to deliver both software and hardware solutions with low power, high performance, reconfigurability, reliability, and adaptivity characteristics, serving as the promising computing substrate for robotic applications. This paper highlights the current progress, design techniques, challenges, and open research challenges in the domain of robotic computing on FPGAs.

cs.RO

Studying the potential of QQq at finite temperature in a holographic model

Using the gauge/gravity duality, we investigate the string breaking and dissolution of two heavy quarks coupled to a light quark at finite temperature. It is found that there exist three configurations of QQq with the increase of separate distance for heavy quarks in the confined phase. Besides, the string breaking occurs at the distance $L_{\rm{QQq}} = 1.27 \rm{fm}$($T = 0.1 \rm{GeV}$) for the decay mode $\rm{Q Q q \rightarrow Q q q+Q \bar{q}}$. In the deconfined phase, QQq will melt at a certain distance then becomes free quarks. At last, we compare the potential of QQq with that of $\rm{Q\bar{Q}}$ and find $\rm{Q\bar{Q}}$ is more stable than QQq at high temperature.

hep-ph

An Energy-Efficient and Runtime-Reconfigurable FPGA-Based Accelerator for Robotic Localization Systems

Simultaneous Localization and Mapping (SLAM) estimates agents' trajectories and constructs maps, and localization is a fundamental kernel in autonomous machines at all computing scales, from drones, AR, VR to self-driving cars. In this work, we present an energy-efficient and runtime-reconfigurable FPGA-based accelerator for robotic localization. We exploit SLAM-specific data locality, sparsity, reuse, and parallelism, and achieve >5x performance improvement over the state-of-the-art. Especially, our design is reconfigurable at runtime according to the environment to save power while sustaining accuracy and performance.

cs.AR

Divergent Effects of Factors on Crashes under Autonomous and Conventional Driving Modes Using A Hierarchical Bayesian Approach

Influencing factors on crashes involved with autonomous vehicles (AVs) have been paid increasing attention. However, there is a lack of comparative analyses between influencing factors on crashes of AVs and human-driven vehicles. To fill this research gap, the study aims to explore the divergent effects of factors on crashes under autonomous and conventional driving modes. This study obtained 154 publicly available autonomous vehicle crash data (70 for the autonomous driving mode and 84 for the conventional driving mode), and 36 explanatory variables were extracted from three categories, including environment, roads, and vehicles. Then, a hierarchical Bayesian approach was applied to analyze the impacting factors on crash type and severity under both driving modes. The results showed that some factors affected both driving modes, but their degrees were different. For example, the presence of turning movement had a greater impact on the crash severity under the conventional driving mode, while the presence of turning movement led to a larger decrease in the likelihood of rear-end crashes under the autonomous driving mode. More influencing factors only had a significant impact on one of the driving modes. For example, in the autonomous driving mode, two sidewalks decreased the severity of crashes, and on-street parking was positively associated with rear-end crashes, but they were not significant in the conventional driving mode. This study could contribute to the understanding and development of autonomous driving systems and the better coordination between autonomous driving and conventional driving.

stat.AP

R&D Towards Cryogenic Optical Links

A number of critical active and passive components of optical links have been tested at 77 K or lower temperatures, demonstrating potential development of optical links operating inside the liquid argon time projection chamber (LArTPC) detector cryostat. A ring oscillator, individual MOSFETs, and a high speed 16:1 serializer fabricated in a commercial 0.25-um silicon-on-sapphire CMOS technology continued to function from room temperature to 4.2 K, 15 K, and 77 K respectively. Three types of laser diodes lase from room temperature to 77 K. Optical fibers and optical connectors exhibited minute attenuation changes from room temperature to 77 K.

physics.ins-det

Hint of a truncated primordial spectrum from the CMB large-scale anomalies

Several satellite missions have uncovered a series of potential anomalies in the fluctuation spectrum of the cosmic microwave background temperature, including: (1) an unexpectedly low level of correlation at large angles, manifested via the angular correlation function, C(theta); and (2) missing power in the low multipole moments of the angular power spectrum, C_ell. Their origin is still debated, however, due to a persistent lack of clarity concerning the seeding of quantum fluctuations in the early Universe. A likely explanation for the first of these appears to be a cutoff, k_min=(3.14 +/- 0.36) x 10^{-4} Mpc^{-1}, in the primordial power spectrum, P(k). Our goal in this paper is twofold: (1) we examine whether the same k_min can also self-consistently explain the missing power at large angles, and (2) we confirm that the of this cutoff in P(k) does not adversely affect the remarkable consistency between the prediction of Planck-LCDM and the Planck measurements at ell > 30. We use the publicly available code CAMB to calculate the angular power spectrum, based on a line-of-sight approach. The code is modified slightly to include the additional parameter (i.e., k_min) characterizing the primordial power spectrum. In addition to this cutoff, the code optimizes all of the usual standard-model parameters. In fitting the angular power spectrum, we find an optimized cutoff, k_min = 2.04^{+1.4}_{-0.79} x 10^{-4} Mpc^{-1}, when using the whole range of ell's, and k_min=3.3^{+1.7}_{-1.3} x 10^{-4} Mpc^{-1}, when fitting only the range ell < 30, where the Sachs-Wolfe effect is dominant. These are fully consistent with the value inferred from C(theta), suggesting that both of these large-angle anomalies may be due to the same truncation in P(k).

astro-ph.CO

Investigating The Impacting Factors on The Public's Attitudes Towards Autonomous Vehicles Using Sentiment Analysis from Social Media Data

The public's attitudes play a critical role in the acceptance, purchase, use, and research and development of autonomous vehicles (AVs). To date, the public's attitudes towards AVs were mostly estimated through traditional survey data with high labor costs and a low quantity of samples, which also might be one of the reasons why the influencing factors on the public's attitudes of AVs have not been studied from multiple aspects in a comprehensive way yet. To address the issue, this study aims to propose a method by using large-scale social media data to investigate key factors that affect the public's attitudes and acceptance of AVs. A total of 954,151 Twitter data related to AVs and 53 candidate independent variables from seven categories were extracted using the web scraping method. Then, sentiment analysis was used to measure the public attitudes towards AVs by calculating sentiment scores. Random forests algorithm was employed to preliminarily select candidate independent variables according to their importance, while a linear mixed model was performed to explore the impacting factors considering the unobserved heterogeneities caused by the subjectivity level of tweets. The results showed that the overall attitude of the public on AVs was slightly optimistic. Factors like "drunk", "blind spot", and "mobility" had the largest impacts on public attitudes. In addition, people were more likely to express positive feelings when talking about words such as "lidar" and "Tesla" that relate to high technologies. Conversely, factors such as "COVID-19", "pedestrian", "sleepy", and "highway" were found to have significantly negative effects on the public's attitudes. The findings of this study are beneficial for the development of AV technologies, the guidelines for AV-related policy formulation, and the public's understanding and acceptance of AVs.

cs.SI

Large-Scale Algebraic Riccati Equations with High-Rank Nonlinear Terms and Constant Terms

For large-scale discrete-time algebraic Riccati equations (DAREs) with high-rank nonlinear and constant terms, the stabilizing solutions are no longer numerically low-rank, resulting in the obstacle in the computation and storage. However, in some proper control problems such as power systems, the potential structure of the state matrix -- banded-plus-low-rank, might make the large-scale computation essentially workable. In this paper, a factorized structure-preserving doubling algorithm (FSDA) is developed under the frame of the banded inverse of nonlinear and constant terms. The detailed iterations format, as well as a deflation process of FSDA, are analyzed in detail. A partial truncation and compression technique is introduced to shrink the dimension of columns of low-rank factors as much as possible. The computation of residual, together with the termination condition of the structured version, is also redesigned.

math.NA

Existence of The Solution to The Quadratic Bilinear Equation Arising from A Class of Quadratic Dynamical Systems

A quadratic dynamical system with practical applications is taken into considered. This system is transformed into a new bilinear system with Hadamard products by means of the implicit matrix structure. The corresponding quadratic bilinear equation is subsequently established via the Volterra series. Under proper conditions the existence of the solution to the equation is proved by using a fixed-point iteration.

math.NA

An inexact proximal DC algorithm with sieving strategy for rank constrained least squares semidefinite programming

In this paper, the optimization problem of the supervised distance preserving projection (SDPP) for data dimension reduction (DR) is considered, which is equivalent to a rank constrained least squares semidefinite programming (RCLSSDP). In order to overcome the difficulties caused by rank constraint, the difference-of-convex (DC) regularization strategy was employed, then the RCLSSDP is transferred into a series of least squares semidefinite programming with DC regularization (DCLSSDP). An inexact proximal DC algorithm with sieving strategy (s-iPDCA) is proposed for solving the DCLSSDP, whose subproblems are solved by the accelerated block coordinate descent (ABCD) method. Convergence analysis shows that the generated sequence of s-iPDCA globally converges to stationary points of the corresponding DC problem. To show the efficiency of our proposed algorithm for solving the RCLSSDP, the s-iPDCA is compared with classical proximal DC algorithm (PDCA) and the PDCA with extrapolation (PDCAe) by performing DR experiment on the COIL-20 database, the results show that the s-iPDCA outperforms the PDCA and the PDCAe in solving efficiency. Moreover, DR experiments for face recognition on the ORL database and the YaleB database demonstrate that the rank constrained kernel SDPP (RCKSDPP) is effective and competitive by comparing the recognition accuracy with kernel semidefinite SDPP (KSSDPP) and kernal principal component analysis (KPCA).

math.OC

Eudoxus: Characterizing and Accelerating Localization in Autonomous Machines

We develop and commercialize autonomous machines, such as logistic robots and self-driving cars, around the globe. A critical challenge to our -- and any -- autonomous machine is accurate and efficient localization under resource constraints, which has fueled specialized localization accelerators recently. Prior acceleration efforts are point solutions in that they each specialize for a specific localization algorithm. In real-world commercial deployments, however, autonomous machines routinely operate under different environments and no single localization algorithm fits all the environments. Simply stacking together point solutions not only leads to cost and power budget overrun, but also results in an overly complicated software stack. This paper demonstrates our new software-hardware co-designed framework for autonomous machine localization, which adapts to different operating scenarios by fusing fundamental algorithmic primitives. Through characterizing the software framework, we identify ideal acceleration candidates that contribute significantly to the end-to-end latency and/or latency variation. We show how to co-design a hardware accelerator to systematically exploit the parallelisms, locality, and common building blocks inherent in the localization framework. We build, deploy, and evaluate an FPGA prototype on our next-generation self-driving cars. To demonstrate the flexibility of our framework, we also instantiate another FPGA prototype targeting drones, which represent mobile autonomous machines. We achieve about 2x speedup and 4x energy reduction compared to widely-deployed, optimized implementations on general-purpose platforms.

cs.AR

iELAS: An ELAS-Based Energy-Efficient Accelerator for Real-Time Stereo Matching on FPGA Platform

Stereo matching is a critical task for robot navigation and autonomous vehicles, providing the depth estimation of surroundings. Among all stereo matching algorithms, Efficient Large-scale Stereo (ELAS) offers one of the best tradeoffs between efficiency and accuracy. However, due to the inherent iterative process and unpredictable memory access pattern, ELAS can only run at 1.5-3 fps on high-end CPUs and difficult to achieve real-time performance on low-power platforms. In this paper, we propose an energy-efficient architecture for real-time ELAS-based stereo matching on FPGA platform. Moreover, the original computational-intensive and irregular triangulation module is reformed in a regular manner with points interpolation, which is much more hardware-friendly. Optimizations, including memory management, parallelism, and pipelining, are further utilized to reduce memory footprint and improve throughput. Compared with Intel i7 CPU and the state-of-the-art CPU+FPGA implementation, our FPGA realization achieves up to 38.4x and 3.32x frame rate improvement, and up to 27.1x and 1.13x energy efficiency improvement, respectively.

cs.AR

An Energy-Efficient Quad-Camera Visual System for Autonomous Machines on FPGA Platform

In our past few years' of commercial deployment experiences, we identify localization as a critical task in autonomous machine applications, and a great acceleration target. In this paper, based on the observation that the visual frontend is a major performance and energy consumption bottleneck, we present our design and implementation of an energy-efficient hardware architecture for ORB (Oriented-Fast and Rotated- BRIEF) based localization system on FPGAs. To support our multi-sensor autonomous machine localization system, we present hardware synchronization, frame-multiplexing, and parallelization techniques, which are integrated in our design. Compared to Nvidia TX1 and Intel i7, our FPGA-based implementation achieves 5.6x and 3.4x speedup, as well as 3.0x and 34.6x power reduction, respectively.

cs.AR

The Matter of Time -- A General and Efficient System for Precise Sensor Synchronization in Robotic Computing

Time synchronization is a critical task in robotic computing such as autonomous driving. In the past few years, as we developed advanced robotic applications, our synchronization system has evolved as well. In this paper, we first introduce the time synchronization problem and explain the challenges of time synchronization, especially in robotic workloads. Summarizing these challenges, we then present a general hardware synchronization system for robotic computing, which delivers high synchronization accuracy while maintaining low energy and resource consumption. The proposed hardware synchronization system is a key building block in our future robotic products.

cs.RO

A Survey of FPGA-Based Robotic Computing

Recent researches on robotics have shown significant improvement, spanning from algorithms, mechanics to hardware architectures. Robotics, including manipulators, legged robots, drones, and autonomous vehicles, are now widely applied in diverse scenarios. However, the high computation and data complexity of robotic algorithms pose great challenges to its applications. On the one hand, CPU platform is flexible to handle multiple robotic tasks. GPU platform has higher computational capacities and easy-touse development frameworks, so they have been widely adopted in several applications. On the other hand, FPGA-based robotic accelerators are becoming increasingly competitive alternatives, especially in latency-critical and power-limited scenarios. With specialized designed hardware logic and algorithm kernels, FPGA-based accelerators can surpass CPU and GPU in performance and energy efficiency. In this paper, we give an overview of previous work on FPGA-based robotic accelerators covering different stages of the robotic system pipeline. An analysis of software and hardware optimization techniques and main technical issues is presented, along with some commercial and space applications, to serve as a guide for future work.

cs.RO