SearcharxivSearch

arXiv subjects

Bo Yu

Publications and source records attributed to Bo Yu.

At least 91 records · Page 5Linked to original sources

Towards Fully Intelligent Transportation through Infrastructure-Vehicle Cooperative Autonomous Driving: Challenges and Opportunities

The infrastructure-vehicle cooperative autonomous driving approach depends on the cooperation between intelligent roads and intelligent vehicles. This approach is not only safer but also more economical compared to the traditional on-vehicle-only autonomous driving approach. In this paper, we introduce our real-world deployment experiences of cooperative autonomous driving, and delve into the details of new challenges and opportunities. Specifically, based on our progress towards commercial deployment, we follow a three-stage development roadmap of the cooperative autonomous driving approach:infrastructure-augmented autonomous driving (IAAD), infrastructure-guided autonomous driving (IGAD), and infrastructure-planned autonomous driving (IPAD).

cs.RO

A new analytical approximation of luminosity distance by optimal HPM-Padé technique

By the use of homotopy perturbation method-Padé (HPM-Padé) technique, a new analytical approximation of luminosity distance in the flat universe is proposed, which has the advantage of significant improvement for accuracy in approximating luminosity distance over cosmological redshift range within $0\leq z\leq 2.5$. Then we confront the analytical expression of luminosity distance that is obtained by our new approach with the observational data, for the purpose of checking whether it works well. In order to probe the robustness of the proposed method, we also confront it to supernova type Ia and recent data on the Hubble expansion rate $H(z)$. Markov Chain Monte Carlo (MCMC) code emcee is used in the data fitting. The result indicates that it works fairly well.

astro-ph.CO

Computing the luminosity distance via optimal homotopy perturbation method

We propose a new algorithm for computing the luminosity distance in the flat universe with a cosmological constant based on Shchigolev's homotopy perturbation method, where the optimization idea is applied to prevent the arbitrariness of initial value choice in Shchigolev's homotopy. Compared with the some existing numerical methods, the result of numerical simulation shows that our algorithm is a very promising and powerful technique for computing the luminosity distance, which has obvious advantages in computational accuracy,computing efficiency and robustness for a given {Ω_m}.

astro-ph.CO

On Designing Computing Systems for Autonomous Vehicles: a PerceptIn Case Study

PerceptIn develops and commercializes autonomous vehicles for micromobility around the globe. This paper makes a holistic summary of PerceptIn's development and operating experiences. This paper provides the business tale behind our product, and presents the development of the computing system for our vehicles. We illustrate the design decision made for the computing system, and show the advantage of offloading localization workloads onto an FPGA platform.

cs.RO

An ADMM-LAP method for total variation myopic deconvolution of adaptive optics retinal images

Adaptive optics (AO) corrected ood imaging of the retina is a popular technique for studying the retinal structure and function in the living eye. However, the raw retinal images are usually of poor contrast and the interpretation of such images requires image deconvolution. Different from standard deconvolution problems where the point spread function (PSF) is completely known, the PSF in these retinal imaging problems is only partially known which leads to the more complicated myopic (mildly blind) deconvolution problem. In this paper, we propose an efficient numerical scheme for solving this myopic deconvolution problem with total variational (TV) regularization. First, we apply the alternating direction method of multipliers (ADMM) to tackle the TV regularizer. Specifically, we reformulate the TV problem as an equivalent equality constrained problem where the objective function is separable, and then minimize the augmented Lagrangian function by alternating between two (separated) blocks of unknowns to obtain the solution. Due to the structure of the retinal images, the subproblems with respect to the fidelity term appearing within each ADMM iteration are tightly coupled and a variation of the Linearize And Project (LAP) method is designed to solve these subproblems efficiently. The proposed method is called the ADMM-LAP method. Theoretically, we establish the subsequence convergence of the ADMM-LAP method to a stationary point. Both the theoretical complexity analysis and numerical results are provided to demonstrate the efficiency of the ADMM-LAP method.

math.OC

Realization of ppm level pressure stability for primary thermometry using a primary piston gauge

To achieve an uncertainty of 0.25 mK in single-pressure refractive-index gas thermometry (SPRIGT), the relative pressure variation of He-4 gas in the range 30 kPa to 90 kPa, should not exceed 4 ppm (k=1). To this end, a novel pressure control system has been developed. It consists of two main parts: a piston gauge to control the pressure, and a home-made gas compensation system to supplement the micro-leak of the piston gauge. In addition, to maintain the piston at constant height, a servo loop is used that automatically determines in real time the amount of extra gas required. At room temperature, the standard deviations of the stabilized pressure are 3.0 mPa at 30 kPa, 4.5 mPa at 60 kPa and 2 mPa at 90 kPa. For the temperature region 5 K-25 K used for SPRIGT in the present work, the relative pressure stability is better than 0.16 ppm i.e. 25 times better than required. Moreover, the same pressure stabilization system is readily transposable to other primary gas thermometers.

physics.ins-det

PAL-Hom method for QP and an application to LP

In this paper, a proximal augmented Lagrangian homotopy (PAL-Hom) method for solving convex quadratic programming problems is proposed. This method takes the proximal augmented Lagrangian method as the outer iteration. To solve the proximal augmented Lagrangian subproblems, a homotopy method is presented as the inner iteration. The homotopy method tracks the piecewise-linear solution path of a parametric quadratic programming problem whose start problem takes an approximate solution as its solution and the target problem is the subproblem to be solved. To improve the performance of the homotopy method, the accelerated proximal gradient method is used to obtain a fairly good approximate solution that implies a good prediction of the optimal active set. Moreover, a sorting technique for the Cholesky factor update as well as an $\varepsilon$-relaxation technique for checking primal-dual feasibility and correcting the active sets are presented to improve the efficiency and robustness of the homotopy method. Simultaneously, a proximal-point-based AL-Hom method which is shown to converge in finite number of steps, is applied to linear programming. Numerical experiments on randomly generated problems and the problems from the CUTEr and Netlib test collections, support vector machines (SVMs) and contact problems of elasticity demonstrate that PAL-Hom is faster than the active-set methods and the parametric active set methods and is competitive to the interior-point methods and the specialized algorithms designed for specific models (e.g., sequential minimal optimization (SMO) method for SVMs).

math.OC

APP-Hom Method for Box Constrained Quadratic Programming

In this paper, based on a $Q$-linear convergence analysis and an estimate of the linear convergence factor of the proximal point (PP) algorithm for solving box constrained quadratic programming (BQP) problems, an accelerated proximal point (APP) algorithm for solving BQP problems is presented. To solve the strictly convex BQP problems in each step of the APP algorithm, an efficient homotopy method, which tracks the solution path of a parametric quadratic program, is given. The algorithm with APP algorithm as outer iteration and the homotopy method as inner iteration is named by APP-Hom. The inner homotopy method is efficient by implementing, a warm-start technique based on the accelerated proximal gradient (APG) method, an $\varepsilon$-relaxation technique for checking prime and dual feasibility and determining/correcting the active set. Numerical tests for randomly generated dense and sparse BQPs, BQPs arising from image deblurring, BQPs in SVM, as well as discretized obstacle problem, elastic-plastic torsion problem, and the journal bearing problem show that the APP algorithm takes much less steps than the PP algorithm, the homotopy method is very efficient for strictly-convex BQP, and in consequence, that the APP-Hom is very efficient for non-convex BQP.

math.OC

Mesh Independence of an Accelerated Block Coordinate Descent Method for Sparse Optimal Control Problems

An accelerated block coordinate descent (ABCD) method in Hilbert space is analyzed to solve the sparse optimal control problem via its dual. The finite element approximation of this method is investigated and convergence results are presents. Based on the second order growth condition of the dual objective function, we show that iteration sequence of dual variables has the iteration complexity of $O(1/k)$. Moreover, we also prove iteration complexity for the primal problem. Two types of mesh-independence for ABCD method are proved, which asserts that asymptotically the infinite dimensional ABCD method and finite dimensional discretizations have the same convergence property, and the iterations of ABCD method remain nearly constant as the discretization is refined.

math.OC

A full multigrid multilevel Monte Carlo method for the single phase subsurface flow with random coefficients

The subsurface flow is usually subject to uncertain porous media structures. In most cases, however, we only have partial knowledge about the porous media properties. A common approach is to model the uncertain parameters of porous media as random fields, then the statistical moments (e.g. expectation) of the Quantity of Interest(QoI) can be evaluated by the Monte Carlo method. In this study, we develop a full multigrid-multilevel Monte Carlo (FMG-MLMC) method to speed up the evaluation of random parameters effects on single-phase porous flows. In general, MLMC method applies a series of discretization with increasing resolution and computes the QoI on each of them. The effective variance reduction is the success of the method. We exploit the similar hierarchies of MLMC and multigrid methods and obtain the solution on coarse mesh $Q^c_l$ as a byproduct of the full multigrid solution on fine mesh $Q^f_l$ on each level $l$. In the cases considered in this work, the computational saving due to the coarse mesh samples saving is $20\%$ asymptotically. Besides, a comparison of Monte Carlo and Quasi-Monte Carlo (QMC) methods reveals a smaller estimator variance and a faster convergence rate of the latter approach in this study.

math.NA

A multi-level ADMM algorithm for elliptic PDE-constrained optimization problems

In this paper, the elliptic PDE-constrained optimization problem with box constraints on the control is studied. To numerically solve the problem, we apply the 'optimize-discretize-optimize' strategy. Specifically, the alternating direction method of multipliers (ADMM) algorithm is applied in function space first, then the standard piecewise linear finite element approach is employed to discretize the subproblems in each iteration. Finally, some efficient numerical methods are applied to solve the discretized subproblems based on their structures. Motivated by the idea of the multi-level strategy, instead of fixing the mesh size before the computation process, we propose the strategy of gradually refining the grid. Moreover, the subproblems in each iteration are solved inexactly. Based on the strategies above, an efficient convergent multi-level ADMM (mADMM) algorithm is proposed. We present the convergence analysis and the iteration complexity results o(1/k) of the proposed algorithm for the PDE-constrained optimization problems. Numerical results show the high efficiency of the mADMM algorithm.

math.OC

PI-BA Bundle Adjustment Acceleration on Embedded FPGAs with Co-observation Optimization

Bundle adjustment (BA) is a fundamental optimization technique used in many crucial applications, including 3D scene reconstruction, robotic localization, camera calibration, autonomous driving, space exploration, street view map generation etc. Essentially, BA is a joint non-linear optimization problem, and one which can consume a significant amount of time and power, especially for large optimization problems. Previous approaches of optimizing BA performance heavily rely on parallel processing or distributed computing, which trade higher power consumption for higher performance. In this paper we propose π-BA, the first hardware-software co-designed BA engine on an embedded FPGA-SoC that exploits custom hardware for higher performance and power efficiency. Specifically, based on our key observation that not all points appear on all images in a BA problem, we designed and implemented a Co-Observation Optimization technique to accelerate BA operations with optimized usage of memory and computation resources. Experimental results confirm that π-BA outperforms the existing software implementations in terms of performance and power consumption.

eess.IV

Structured FISTA for Image Restoration

In this paper, we propose an efficient numerical scheme for solving some large scale ill-posed linear inverse problems arising from image restoration. In order to accelerate the computation, two different hidden structures are exploited. First, the coefficient matrix is approximated as the sum of a small number of Kronecker products. This procedure not only introduces one more level of parallelism into the computation but also enables the usage of computationally intensive matrix-matrix multiplications in the subsequent optimization procedure. We then derive the corresponding Tikhonov regularized minimization model and extend the fast iterative shrinkage-thresholding algorithm (FISTA) to solve the resulting optimization problem. Since the matrices appearing in the Kronecker product approximation are all structured matrices (Toeplitz, Hankel, etc.), we can further exploit their fast matrix-vector multiplication algorithms at each iteration. The proposed algorithm is thus called structured fast iterative shrinkage-thresholding algorithm (sFISTA). In particular, we show that the approximation error introduced by sFISTA is well under control and sFISTA can reach the same image restoration accuracy level as FISTA. Finally, both the theoretical complexity analysis and some numerical results are provided to demonstrate the efficiency of sFISTA.

math.NA

PI-Edge: A Low-Power Edge Computing System for Real-Time Autonomous Driving Services

To simultaneously enable multiple autonomous driving services on affordable embedded systems, we designed and implemented π-Edge, a complete edge computing framework for autonomous robots and vehicles. The contributions of this paper are three-folds: first, we developed a runtime layer to fully utilize the heterogeneous computing resources of low-power edge computing systems; second, we developed an extremely lightweight operating system to manage multiple autonomous driving services and their communications; third, we developed an edge-cloud coordinator to dynamically offload tasks to the cloud to optimize client system energy consumption. To the best of our knowledge, this is the first complete edge computing system of a production autonomous vehicle. In addition, we successfully implemented π-Edge on a Nvidia Jetson and demonstrated that we could successfully support multiple autonomous driving services with only 11 W of power consumption, and hence proving the effectiveness of the proposed π-Edge system.

cs.DC

An FE-dABCD algorithm for elliptic optimal control problems with constraints on the gradient of the state and control

In this paper, elliptic control problems with integral constraint on the gradient of the state and box constraints on the control are considered. The optimal conditions of the problem are proved. To numerically solve the problem, we use the 'First discretize, then optimize' approach. Specifically, we discretize both the state and the control by piecewise linear functions. To solve the discretized problem efficiently, we first transform it into a multi-block unconstrained convex optimization problem via its dual, then we extend the inexact majorized accelerating block coordinate descent (imABCD) algorithm to solve it. The entire algorithm framework is called finite element duality-based inexact majorized accelerating block coordinate descent (FE-dABCD) algorithm. Thanks to the inexactness of the FE-dABCD algorithm, each subproblems are allowed to be solved inexactly. For the smooth subproblem, we use the generalized minimal residual (GMRES) method with preconditioner to slove it. For the nonsmooth subproblems, one of them has a closed form solution through introducing appropriate proximal term, another is solved combining semi-smooth Newton (SSN) method. Based on these efficient strategies, we prove that our proposed FE-dABCD algorithm enjoys $O(\frac{1}{k^2})$ iteration complexity. Some numerical experiments are done and the numerical results show the efficiency of the FE-dABCD algorithm.

math.OC

Learning to Generate Structured Queries from Natural Language with Indirect Supervision

Generating structured query language (SQL) from natural language is an emerging research topic. This paper presents a new learning paradigm from indirect supervision of the answers to natural language questions, instead of SQL queries. This paradigm facilitates the acquisition of training data due to the abundant resources of question-answer pairs for various domains in the Internet, and expels the difficult SQL annotation job. An end-to-end neural model integrating with reinforcement learning is proposed to learn SQL generation policy within the answer-driven learning paradigm. The model is evaluated on datasets of different domains, including movie and academic publication. Experimental results show that our model outperforms the baseline models.

cs.CL

Simulation Study on Collaborative Content Distribution in Delay Tolerant Vehicular Networks

Modern vehicles are equipped with more and more sophisticated computer modules, which need to periodically download files from the cloud, such as security certificates, digital maps, system firmwares, etc. Collaborative content distribution utilizes V2V communication to distribute large files across the vehicular networks. It has the potential to significantly reduce the cost of cellular-based communication such as 4G LTE. In this report, we have conducted a simulation study to verify the feasibility of a hybrid cellular and V2V collaborative content distribution network. In our simulation, a small portion of the simulated vehicles download the file directly from the cloud via cellular communication, while other vehicles receive the file via collaborative V2V communications. Our simulation results show that, with only 1\% of vehicles enabled with cellular communication, it takes less than 24 hours to distribute a file to 90\% of the vehicles in a metropolitan area, and around 48 to 72 hours to distribute to 99\%. The results are very promising for many delay-tolerant content distribution applications in vehicular networks.

cs.NI

PIRT: A Runtime Framework to Enable Energy-Efficient Real-Time Robotic Applications on Heterogeneous Architectures

Enabling full robotic workloads with diverse behaviors on mobile systems with stringent resource and energy constraints remains a challenge. In recent years, attempts have been made to deploy single-accelerator-based computing platforms (such as GPU, DSP, or FPGA) to address this challenge, but with little success. The core problem is two-fold: firstly, different robotic tasks require different accelerators, and secondly, managing multiple accelerators simultaneously is overwhelming for developers. In this paper, we propose PIRT, the first robotic runtime framework to efficiently manage dynamic task executions on mobile systems with multiple accelerators as well as on the cloud to achieve better performance and energy savings. With PIRT, we enable a robot to simultaneously perform autonomous navigation with 25 FPS of localization, obstacle detection with 3 FPS, route planning, large map generation, and scene understanding, traveling at a max speed of 5 miles per hour, all within an 11W computing power envelope.

cs.RO