SearcharxivSearch

arXiv subjects

Bo Yu

Publications and source records attributed to Bo Yu.

At least 55 records · Page 3Linked to original sources

SCTBEM: A scaled coordinate transformation boundary element method with 99-line MATLAB code

This paper introduces the Scaled Coordinate Transformation Boundary Element Method (SCTBEM), a novel boundary-type method for solving 3D potential problems. To address the challenges of applying the Boundary Element Method (BEM) to complex problems, it is common practice to use the fundamental solution corresponding to the partial governing equation operator to establish the integral equation. However, this approach introduces domain integral, which may jeopardize the dimensionality reduction advantages of BEM. To preserve the benefits of dimensionality reduction, this paper proposes a novel domain integral transformation method known as the Scaled Coordinate Transformation (SCT). The SCT is purely a mathematical operation that does not rely on particular solution of operators, which requires only discretization on the structure's surface while remaining analytical in the radial direction. An even better novelty is that the lower-order singularity can be eliminated by coordinate translation technique. To facilitate the wider adoption of BEM, the authors present 99-line MATLAB code. Numerical results confirm that the SCTBEM exhibits high numerical accuracy even when dealing with complex model.

math.NA

DaDu-Corki: Algorithm-Architecture Co-Design for Embodied AI-powered Robotic Manipulation

Embodied AI robots have the potential to fundamentally improve the way human beings live and manufacture. Continued progress in the burgeoning field of using large language models to control robots depends critically on an efficient computing substrate, and this trend is strongly evident in manipulation tasks. In particular, today's computing systems for embodied AI robots for manipulation tasks are designed purely based on the interest of algorithm developers, where robot actions are divided into a discrete frame basis. Such an execution pipeline creates high latency and energy consumption. This paper proposes \textsc{Corki}\xspace, an algorithm-architecture co-design framework for real-time embodied AI-powered robotic manipulation applications. We aim to decouple LLM inference, robotic control, and data communication in the embodied AI robots' compute pipeline. Instead of predicting action for one single frame, \textsc{Corki}\xspace predicts the trajectory for the near future to reduce the frequency of LLM inference. The algorithm is coupled with a hardware that accelerates transforming trajectory into actual torque signals used to control robots and an execution pipeline that parallels data communication with computation. \textsc{Corki}\xspace largely reduces LLM inference frequency by up to $5.1\times$, resulting in up to $5.9\times$ speed up. The success rate improvement can be up to 13.9\%.

cs.AR

Zero product and zero Jordan product determined Munn algebras

Let $\mathfrak{M}(\mathbb{D}, m, n, P)$ be the ring of all $m \times n$ matrices over a division ring $\mathbb{D}$, with the product given by $A \bullet B=A P B$, where $P$ is a fixed $n \times m$ matrix over $\mathbb{D}$. When $2\leq m, n <\infty$ and $\operatorname{rank} P \geq 2$, we demonstrate that every element in $\mathcal{A}=\mathfrak{M}(\mathbb{D}, m, n, P)$ is a sum of finite products of pairs of commutators. We also estimate the minimal number $N$ such that $\mathcal{A}= \sum^N [\mathcal{A}, \mathcal{A}][\mathcal{A}, \mathcal{A}]$. Furthermore, if $\operatorname{char}\mathbb{D}\neq 2$, we prove that $\mathfrak{M}(\mathbb{D}, m, n, P)$ is additively spanned by Jordan products of idempotents. For a field $\mathbb{F}$ with $\operatorname{char}\mathbb{F}\neq 2, 3$, we show that the Munn algebra $\mathfrak{M}(\mathbb{F}, m, n, P)$ is zero product determined and zero Jordan product determined.

math.RA

Constellation Dataset: Benchmarking High-Altitude Object Detection for an Urban Intersection

We introduce Constellation, a dataset of 13K images suitable for research on detection of objects in dense urban streetscapes observed from high-elevation cameras, collected for a variety of temporal conditions. The dataset addresses the need for curated data to explore problems in small object detection exemplified by the limited pixel footprint of pedestrians observed tens of meters from above. It enables the testing of object detection models for variations in lighting, building shadows, weather, and scene dynamics. We evaluate contemporary object detection architectures on the dataset, observing that state-of-the-art methods have lower performance in detecting small pedestrians compared to vehicles, corresponding to a 10% difference in average precision (AP). Using structurally similar datasets for pretraining the models results in an increase of 1.8% mean AP (mAP). We further find that incorporating domain-specific data augmentations helps improve model performance. Using pseudo-labeled data, obtained from inference outcomes of the best-performing models, improves the performance of the models. Finally, comparing the models trained using the data collected in two different time intervals, we find a performance drift in models due to the changes in intersection conditions over time. The best-performing model achieves a pedestrian AP of 92.0% with 11.5 ms inference time on NVIDIA A100 GPUs, and an mAP of 95.4%.

cs.CV

An efficient asymptotic DC method for sparse and low-rank matrix recovery

The optimization problem of sparse and low-rank matrix recovery is considered, which involves a least squares problem with a rank constraint and a cardinality constraint. To overcome the challenges posed by these constraints, an asymptotic difference-of-convex (ADC) method that employs a Moreau smoothing approach and an exact penalty approach is proposed to transform this problem into a DC programming format gradually. To solve the gained DC programming, by making full use of its DC structure, an efficient inexact DC algorithm with sieving strategy (siDCA) is introduced. The subproblem of siDCA is solved by an efficient dual-based semismooth Newton method. The convergence of the solution sequence generated by siDCA is proved. To illustrate the effectiveness of ADC-siDCA, matrix recovery experiments on nonnegative and positive semidefinite matrices. The numerical results are compared with those obtained using a successive DC approximation minimization method and a penalty proximal alternating linearized minimization approach. The outcome of the comparison indicates that ADC-siDCA surpasses the other two methods in terms of efficiency and recovery error. Additionally, numerical experiments on sparse phase retrieval demonstrate that ADC-siDCA is a valuable tool for recovering sparse and low-rank Hermitian matrices.

math.OC

Two infinite families of facets of the holographic entropy cone

We verify that the recently proven infinite families of holographic entropy inequalities are maximally tight, i.e. they are facets of the holographic entropy cone. The proof is technical but it offers some heuristic insight. On star graphs, both families of inequalities quantify how concentrated / spread information is with respect to a dihedral symmetry acting on subsystems. In addition, toric inequalities viewed in the K-basis show an interesting interplay between four-party and six-party perfect tensors.

hep-th

A Non-parametric Reconstruction of the Hubble Parameter $H(z)$ Based on Radial Basis Function Neural Networks

Accurately measuring the Hubble parameter is vital for understanding the expansion history and properties of the universe. In this paper, we propose a new method that supplements the covariance between redshift pairs to improve the reconstruction of the Hubble parameter using the OHD dataset. Our approach utilizes a cosmological model-independent radial basis function neural network (RBFNN) to describe the Hubble parameter as a function of redshift effectively. Our experiments show that this method results in a reconstructed Hubble parameter of $H_0 = 67.1\pm9.7~\mathrm{km~s^{-1}~Mpc^{-1}}$ , which is more noise-resistant and fits better with the $Λ$CDM model at high redshifts. Providing the covariance between redshift pairs in subsequent observations will significantly improve the reliability and accuracy of Hubble parametric data reconstruction. Future applications of this method could help overcome the limitations of previous methods and lead to new advances in our understanding of the universe.

astro-ph.CO

Thales: Formulating and Estimating Architectural Vulnerability Factors for DNN Accelerators

As Deep Neural Networks (DNNs) are increasingly deployed in safety critical and privacy sensitive applications such as autonomous driving and biometric authentication, it is critical to understand the fault-tolerance nature of DNNs. Prior work primarily focuses on metrics such as Failures In Time (FIT) rate and the Silent Data Corruption (SDC) rate, which quantify how often a device fails. Instead, this paper focuses on quantifying the DNN accuracy given that a transient error has occurred, which tells us how well a network behaves when a transient error occurs. We call this metric Resiliency Accuracy (RA). We show that existing RA formulation is fundamentally inaccurate, because it incorrectly assumes that software variables (model weights/activations) have equal faulty probability under hardware transient faults. We present an algorithm that captures the faulty probabilities of DNN variables under transient faults and, thus, provides correct RA estimations validated by hardware. To accelerate RA estimation, we reformulate RA calculation as a Monte Carlo integration problem, and solve it using importance sampling driven by DNN specific heuristics. Using our lightweight RA estimation method, we show that transient faults lead to far greater accuracy degradation than what todays DNN resiliency tools estimate. We show how our RA estimation tool can help design more resilient DNNs by integrating it with a Network Architecture Search framework.

cs.AR

Potential and string breaking of doubly heavy baryon at finite temperature and chemical potential

Using gauge/gravity duality, we first study the string breaking and melting of doubly heavy baryon at a finite chemical potential and temperature in this paper. The decay mode $\rm{Q Q q \rightarrow Q q q+Q \bar{q}}$ is investigated with the presence of temperature and chemical potential in this paper. With the increase of temperature and chemical potential, string breaking takes place at a smaller potential energy and the string-breaking distance will increase slightly. It is also found that the $\rm{QQq}$ melts at small separate distance with the increase of temperature and chemical potential. Then, we compare the screening distance of $\rm{QQq}$ with $\rm{Q \bar{Q}}$ under the same conditions. Finally, we draw the melting diagram of $\rm{QQq}$ and $\rm{Q \bar{Q}}$ in the $T-μ$ plane.

hep-ph

FGLQR: Factor Graph Accelerator of LQR Control for Autonomous Machines

Factor graph represents the factorization of a probability distribution function and serves as an effective abstraction in various autonomous machine computing tasks. Control is one of the core applications in autonomous machine computing stacks. Among all control algorithms, Linear Quadratic Regulator (LQR) offers one of the best trade-offs between efficiency and accuracy. However, due to the inherent iterative process and extensive computation, it is a challenging task for the autonomous systems with real-time limits and energy constrained. In this paper, we present FGLQR, an accelerator of LQR control for autonomous machines using the abstraction of a factor graph. By transforming the dynamic equation constraints into least squares constraints, the factor graph solving process is more hardware friendly and accelerated with almost no loss in accuracy. With a domain specific parallel solving pattern, FGLQR achieves 10.2x speed up and 32.9x energy reduction compared to the software implementation on an advanced Intel CPU.

cs.AR

Terahertz reconfigurable multi-functional metamaterials based on 3D printed mortise-tenon structures

The emergence of metamaterial has provided an unprecedented ability to manipulate electromagnetic waves, especially in the terahertz band where there is a lack of natural response materials. However, most metamaterials are fixed single function due to the fixed structure at the beginning of design. The paper reports a reconfigurable multi-functional terahertz metamaterial with variable structures based on mortise and tenon mechanism. And a hybrid 3D printing method based on FDM and E-jet is proposed to fabricate the metamaterials, which simplifies the processing process, improves the speed, and reduces the cost compared to traditional semiconductor processing methods. Through flexible mortise and tenon connections, the metamaterial can achieve: (1) narrowband transmission and broadband absorption; (2) perfect reflection; (3) narrowband reflection and broadband absorption. Relying on ingenious design and processing, the multi-functional metamaterials are expected to be widely used in fields such as electromagnetic shielding, radar stealth, communication and so on.

physics.optics

Autonomy 2.0: The Quest for Economies of Scale

With the advancement of robotics and AI technologies in the past decade, we have now entered the age of autonomous machines. In this new age of information technology, autonomous machines, such as service robots, autonomous drones, delivery robots, and autonomous vehicles, rather than humans, will provide services. In this article, through examining the technical challenges and economic impact of the digital economy, we argue that scalability is both highly necessary from a technical perspective and significantly advantageous from an economic perspective, thus is the key for the autonomy industry to achieve its full potential. Nonetheless, the current development paradigm, dubbed Autonomy 1.0, scales with the number of engineers, instead of with the amount of data or compute resources, hence preventing the autonomy industry to fully benefit from the economies of scale, especially the exponentially cheapening compute cost and the explosion of available data. We further analyze the key scalability blockers and explain how a new development paradigm, dubbed Autonomy 2.0, can address these problems to greatly boost the autonomy industry.

cs.RO

The effect of gluon condensate on the entanglement entropy in a holographic model

In this study, we examine the impact of the gluon condensate on holographic entanglement entropy within an Einstein-Dilaton model at both zero and finite temperatures. A critical length exists for the difference in entanglement entropy between connected and disconnected surfaces in this model, which is typically interpreted as an indicator of phase transition. As the gluon condensate increases, the critical length decreases, suggesting that confinement strengthens at zero temperature. Additionally, the entropic C-function abruptly drops to zero at the critical length, indicating the absence of entangled states. At finite temperatures, the results show that the effect of the gluon condensate on the critical length is qualitatively similar to that at zero temperature. We observe that the entropic C-function increases as a function of $L$ at finite temperature, though it exhibits competitive behaviors when the gluon condensate is large.

hep-ph

Multi-Modality Multi-Scale Cardiovascular Disease Subtypes Classification Using Raman Image and Medical History

Raman spectroscopy (RS) has been widely used for disease diagnosis, e.g., cardiovascular disease (CVD), owing to its efficiency and component-specific testing capabilities. A series of popular deep learning methods have recently been introduced to learn nuance features from RS for binary classifications and achieved outstanding performance than conventional machine learning methods. However, these existing deep learning methods still confront some challenges in classifying subtypes of CVD. For example, the nuance between subtypes is quite hard to capture and represent by intelligent models due to the chillingly similar shape of RS sequences. Moreover, medical history information is an essential resource for distinguishing subtypes, but they are underutilized. In light of this, we propose a multi-modality multi-scale model called M3S, which is a novel deep learning method with two core modules to address these issues. First, we convert RS data to various resolution images by the Gramian angular field (GAF) to enlarge nuance, and a two-branch structure is leveraged to get embeddings for distinction in the multi-scale feature extraction module. Second, a probability matrix and a weight matrix are used to enhance the classification capacity by combining the RS and medical history data in the multi-modality data fusion module. We perform extensive evaluations of M3S and found its outstanding performance on our in-house dataset, with accuracy, precision, recall, specificity, and F1 score of 0.9330, 0.9379, 0.9291, 0.9752, and 0.9334, respectively. These results demonstrate that the M3S has high performance and robustness compared with popular methods in diagnosing CVD subtypes.

eess.IV

Data and Knowledge Co-driving for Cancer Subtype Classification on Multi-Scale Histopathological Slides

Artificial intelligence-enabled histopathological data analysis has become a valuable assistant to the pathologist. However, existing models lack representation and inference abilities compared with those of pathologists, especially in cancer subtype diagnosis, which is unconvincing in clinical practice. For instance, pathologists typically observe the lesions of a slide from global to local, and then can give a diagnosis based on their knowledge and experience. In this paper, we propose a Data and Knowledge Co-driving (D&K) model to replicate the process of cancer subtype classification on a histopathological slide like a pathologist. Specifically, in the data-driven module, the bagging mechanism in ensemble learning is leveraged to integrate the histological features from various bags extracted by the embedding representation unit. Furthermore, a knowledge-driven module is established based on the Gestalt principle in psychology to build the three-dimensional (3D) expert knowledge space and map histological features into this space for metric. Then, the diagnosis can be made according to the Euclidean distance between them. Extensive experimental results on both public and in-house datasets demonstrate that the D&K model has a high performance and credible results compared with the state-of-the-art methods for diagnosing histopathological subtypes. Code: https://github.com/Dennis-YB/Data-and-Knowledge-Co-driving-for-Cancer-Subtypes-Classification

cs.CV

Benchmark modeling and 3D applications of solidification and macro-segregation based on an operator-splitting and fully decoupled scheme with term-wise matrix assembly

The solidification and macro-segregation problem involving unsteady multi-physics and multi-phase fields is typically a complex process with mass, momentum, heat, and species transfers among solid, mushy, and liquid phase regions. The quantitative prediction of phase change, chemical heterogeneities, and multi-phase and multi-component flows plays critical roles in many natural scenarios and industrial applications that involve many disciplines, like material, energy, and even planet science. In view of this, some scholars and research institutions have called for more contributors to join the benchmark analysis of solidification and segregation problems. Our work proposes an operator-splitting and matrix-based method to avoid non-linear systems. Also, the combination of vectorization and forward equation-based matrix assembly techniques enhances the implementability of extensions of 3D applications. Lastly, the novel scheme is well validated through a bunch of 2D and 3D benchmark cases. The numerical results also illustrate that this method can ensure accurate prediction and adequately capture the physical details of phenomena caused by the solutally and thermally driven flow, which include channel segregation, the formation of freckles, edge effect, aspect ratio effect, and 3D effect.

physics.flu-dyn

INTERNEURON: A Middleware with Multi-Network Communication Reliability for Infrastructure Vehicle Cooperative Autonomous Driving

Infrastructure-Vehicle Cooperative Autonomous Driving (IVCAD) is a new paradigm of autonomous driving, which relies on the cooperation between intelligent roads and autonomous vehicles. This paradigm has been shown to be safer and more efficient compared to the on-vehicle-only autonomous driving paradigm. Our real-world deployment data indicates that the effectiveness of IVCAD is constrained by reliability and performance of commercial communication networks. This paper targets this exact problem, and proposes INTERNEURON, a middleware to achieve high communication reliability between intelligent roads and autonomous vehicles, in the context of IVCAD. Specifically, INTERNEURON dynamically matches IVCAD applications and the underlying communication technologies based on varying communication performance and quality needs. Evaluation results confirm that INTERNEURON reduces deadline violations by more than 95\%, significantly improving the reliability of IVCAD systems.

cs.RO

A Reliable Calibration of HII Galaxies Hubble Diagram with Cosmic Chronometers and Artificial Neural Network

The $L-σ$ relation of HII galaxies (HIIGx) calibrated by a distance indicator is a reliable standard candle for measuring the Hubble constant $H_0$. The most straightforward calibration technique anchors them with the first tier of distance ladders from the same galaxies. Recently another promising method that uses the cosmological model-independent Cosmic Chronometers (CC) as a calibrator has been proposed. We promote this technique by removing the assumptions about the cosmic flatness and using a non-parametric Artificial Neural Network for the data reconstruction process. We observe a correlation between the cosmic curvature density parameter and the slope of the $L-σ$ relation, thereby improving the reliability of the calibration. Using the calibrated HIIGx Hubble diagram, we obtain a Type Ia Supernovae Hubble diagram free of the conventional assumption about $H_0$. Finally we get a value of $H_0=65.9_{-2.9}^{+3.0} \mathrm{km s^{-1} Mpc^{-1}}$, which is compatible with latest Planck18 measurement.

astro-ph.CO