Searcharxiv⌕ Search

arXiv subjects

Jun Yan

Publications and source records attributed to Jun Yan.

At least 163 records · Page 9Linked to original sources

Non-LTE ionization potential depression model for warm and hot dense plasma

For warm and hot dense plasma (WHDP), the ionization potential depression (IPD) is a key physical parameter in determining its ionization balance, therefore a reliable and universal IPD model is highly required to understand its microscopic material properties and resolve those existing discrepancies between the theoretical and experimental results. However, the weak temperature dependence of the nowadays IPD models prohibits their application through much of the WHDP regime, especially for the non-LTE dense plasma produced by short-pulse laser. In this work, we propose a universal non-LTE IPD model with the contribution of the inelastic atomic processes, and found that three-body recombination and collision ionization processes become important in determining the electron distribution and further affect the IPD in warm and dense plasma. The proposed IPD model is applied to treat the IPD experiments available in warm and hot dense plasmas and excellent agreements are obtained in comparison with those latest experiments of the IPD for Al plasmas with wide-range conditions of 70-700 eV temperature and 0.2-3 times of solid density, as well as a typical non-LTE system of hollow Al ions. We demonstrate that the present IPD model has a significant temperature dependence due to the consideration of the inelastic atomic processes. With the low computational cost and wide range applicability of WHDP, the proposed model is expected to provide a promising tool to study the ionization balance and the atomic processes as well as the related radiation and particle transports properties of a wide range of WHDP.

physics.plasm-ph↗

Survival Modeling of Suicide Risk with Rare and Uncertain Diagnoses

Motivated by the pressing need for suicide prevention through improving behavioral healthcare, we use medical claims data to study the risk of subsequent suicide attempts for patients who were hospitalized due to suicide attempts and later discharged. Understanding the risk behaviors of such patients at elevated suicide risk is an important step toward the goal of "Zero Suicide." An immediate and unconventional challenge is that the identification of suicide attempts from medical claims contains substantial uncertainty: almost 20% of "suspected" suicide attempts are identified from diagnosis codes indicating external causes of injury and poisoning with undermined intent. It is thus of great interest to learn which of these undetermined events are more likely actual suicide attempts and how to properly utilize them in survival analysis with severe censoring. To tackle these interrelated problems, we develop an integrative Cox cure model with regularization to perform survival regression with uncertain events and a latent cure fraction. We apply the proposed approach to study the risk of subsequent suicide attempts after suicide-related hospitalization for the adolescent and young adult population, using medical claims data from Connecticut. The identified risk factors are highly interpretable; more intriguingly, our method distinguishes the risk factors that are most helpful in assessing either susceptibility or timing of subsequent attempts. The predicted statuses of the uncertain attempts are further investigated, leading to several new insights on suicide event identification.

stat.AP↗

Engineering Blockchain Based Software Systems: Foundations, Survey, and Future Directions

Many scientific and practical areas have shown increasing interest in reaping the benefits of blockchain technology to empower software systems. However, the unique characteristics and requirements associated with Blockchain Based Software (BBS) systems raise new challenges across the development lifecycle that entail an extensive improvement of conventional software engineering. This article presents a systematic literature review of the state-of-the-art in BBS engineering research from a software engineering perspective. We characterize BBS engineering from the theoretical foundations, processes, models, and roles and discuss a rich repertoire of key development activities, principles, challenges, and techniques. The focus and depth of this survey not only gives software engineering practitioners and researchers a consolidated body of knowledge about current BBS development but also underpins a starting point for further research in this field.

cs.SE↗

The generalization error of max-margin linear classifiers: Benign overfitting and high dimensional asymptotics in the overparametrized regime

Modern machine learning classifiers often exhibit vanishing classification error on the training set. They achieve this by learning nonlinear representations of the inputs that maps the data into linearly separable classes. Motivated by these phenomena, we revisit high-dimensional maximum margin classification for linearly separable data. We consider a stylized setting in which data $(y_i,{\boldsymbol x}_i)$, $i\le n$ are i.i.d. with ${\boldsymbol x}_i\sim\mathsf{N}({\boldsymbol 0},{\boldsymbol Σ})$ a $p$-dimensional Gaussian feature vector, and $y_i \in\{+1,-1\}$ a label whose distribution depends on a linear combination of the covariates $\langle {\boldsymbol θ}_*,{\boldsymbol x}_i \rangle$. While the Gaussian model might appear extremely simplistic, universality arguments can be used to show that the results derived in this setting also apply to the output of certain nonlinear featurization maps. We consider the proportional asymptotics $n,p\to\infty$ with $p/n\to ψ$, and derive exact expressions for the limiting generalization error. We use this theory to derive two results of independent interest: $(i)$ Sufficient conditions on $({\boldsymbol Σ},{\boldsymbol θ}_*)$ for `benign overfitting' that parallel previously derived conditions in the case of linear regression; $(ii)$ An asymptotically exact expression for the generalization error when max-margin classification is used in conjunction with feature vectors produced by random one-layer neural networks.

math.ST↗

A representation formula of the viscosity solution of the contact Hamilton-Jacobi equation and its applications

Assume $M$ is a closed, connected and smooth Riemannian manifold. We consider the evolutionary Hamilton-Jacobi equation \begin{equation*} \left\{ \begin{aligned} &\partial_t u(x,t)+H(x,u(x,t),\partial_xu(x,t))=0,\quad (x,t)\in M\times(0,+\infty), \\ &u(x,0)=φ(x), \end{aligned} \right. \end{equation*} where $φ\in C(M)$ and the stationary one \begin{equation*} H(x,u(x),\partial_x u(x))=0, \end{equation*} where $H(x,u,p)$ is continuous, convex and coercive in $p$, uniformly Lipschitz in $u$. By introducing a solution semigroup, we provide a representation formula of the viscosity solution of the evolutionary equation. As its applications, we obtain a necessary and sufficient condition for the existence of the viscosity solutions of the stationary equations. Moreover, we prove a new comparison theorem depending on the neighborhood of the projected Aubry set essentially, which is different from the one for the Hamilton-Jacobi equation independent of $u$.

math.AP↗

Robust Implementation of Foreground Extraction and Vessel Segmentation for X-ray Coronary Angiography Image Sequence

The extraction of contrast-filled vessels from X-ray coronary angiography (XCA) image sequence has important clinical significance for intuitively diagnosis and therapy. In this study, the XCA image sequence is regarded as a 3D tensor input, the vessel layer is regarded as a sparse tensor, and the background layer is regarded as a low-rank tensor. Using tensor nuclear norm (TNN) minimization, a novel method for vessel layer extraction based on tensor robust principal component analysis (TRPCA) is proposed. Furthermore, considering the irregular movement of vessels and the low-frequency dynamic disturbance of surrounding irrelevant tissues, the total variation (TV) regularized spatial-temporal constraint is introduced to smooth the foreground layer. Subsequently, for vessel layer images with uneven contrast distribution, a two-stage region growing (TSRG) method is utilized for vessel enhancement and segmentation. A global threshold method is used as the preprocessing to obtain main branches, and the Radon-Like features (RLF) filter is used to enhance and connect broken minor segments, the final binary vessel mask is constructed by combining the two intermediate results. The visibility of TV-TRPCA algorithm for foreground extraction is evaluated on clinical XCA image sequences and third-party dataset, which can effectively improve the performance of commonly used vessel segmentation algorithms. Based on TV-TRPCA, the accuracy of TSRG algorithm for vessel segmentation is further evaluated. Both qualitative and quantitative results validate the superiority of the proposed method over existing state-of-the-art approaches.

cs.CV↗

Locating the Sources of Sub-synchronous Oscillations Induced by the Control of Voltage Source Converters Based on Energy Structure and Nonlinearity Detection

The oscillation phenomena associated with the control of voltage source converters (VSCs) are widely concerning, and locating the source of these oscillations is crucial to suppressing them; therefore, this paper presents a locating scheme, based on the energy structure and nonlinearity detection. On the one hand, the energy structure, which conforms with the principle of the energy-based method and dissipativity theory, is constructed to describe the transient energy flow for VSCs, and on this basis, a defined characteristic quantity is implemented to narrow the scope of oscillation source location; on the other hand, according to the self-sustained oscillation characteristics of VSCs, an index for nonlinearity detection is applied to locate the VSCs which produce the oscillation energy. The combination of the energy structure and nonlinearity detection could distinguish the contributions of different VSCs to the oscillation. The results of a case study implemented by the PSCAD/EMTDC simulation validate the proposed scheme.

eess.SY↗

Pricing Time-to-Event Contingent Cash Flows: A Discrete-Time Survival Analysis Approach

Prudent management of insurance investment portfolios requires competent asset pricing of fixed-income assets with time-to-event contingent cash flows, such as consumer asset-backed securities (ABS). Current market pricing techniques for these assets either rely on a non-random time-to-event model or may not utilize detailed asset-level data that is now available with most public transactions. We first establish a framework capable of yielding estimates of the time-to-event random variable from securitization data, which is discrete and often subject to left-truncation and right-censoring. We then show that the vector of discrete-time hazard rate estimators is asymptotically multivariate normal with independent components, which has not yet been done in the statistical literature in the case of both left-truncation and right-censoring. The time-to-event distribution estimates are then fed into our cash flow model, which is capable of calculating a formulaic price of a pool of time-to-event contingent cash flows vis-á-vis calculating an expected present value with respect to the estimated time-to-event distribution. In an application to a subset of 29,845 36-month leases from the Mercedes-Benz Auto Lease Trust 2017-A (MBALT 2017-A) bond, our pricing model yields estimates closer to the actual realized future cash flows than the non-random time-to-event model, especially as the fitting window increases. Finally, in certain settings, the asymptotic properties of the hazard rate estimators allow investors to assess the potential uncertainty of the price point estimates, which we illustrate for a subset of 493 24-month leases from MBALT 2017-A.

q-fin.RM↗

Direct and inverse problems for a third order self-adjoint differential operator with periodic boundary conditions and nonlocal potential

A third order self-adjoint differential operator with periodic boundary conditions and an one-dimensional perturbation has been considered. For this operator, we first show that the spectrum consists of simple eigenvalues and finitely many eigenvalues of multiplicity two. Then the expressions of eigenfunctions and resolvent are described. Finally, the inverse problems for recovering all the components of the one-dimensional perturbation are solved. In particular, we prove the Ambarzumyan-type theorem and show that the even or odd potential can be reconstructed by three spectra.

math.CA↗

Distributed Interaction Graph Construction for Dynamic DCOPs in Cooperative Multi-agent Systems

DCOP algorithms usually rely on interaction graphs to operate. In open and dynamic environments, such methods need to address how this interaction graph is generated and maintained among agents. Existing methods require reconstructing the entire graph upon detecting changes in the environment or assuming that new agents know potential neighbors to facilitate connection. We propose a novel distributed interaction graph construction algorithm to address this problem. The proposed method does not assume a predefined constraint graph and stabilizes after disruptive changes in the environment. We evaluate our approach by pairing it with existing DCOP algorithms to solve several generated dynamic problems. The experiment results show that the proposed algorithm effectively constructs and maintains a stable multi-agent interaction graph for open and dynamic environments.

cs.AI↗

Estimating a distribution function for discrete data subject to random truncation with an application to structured finance

Proper econometric analysis should be informed by data structure. Many forms of financial data are recorded in discrete-time and relate to products of a finite term. If the data comes from a financial trust, it will often be further subject to random left-truncation. While the literature for estimating a distribution function from left-truncated data is extensive, a thorough literature search reveals that the case of discrete data over a finite number of possible values has received little attention. A precise discrete framework and suitable sampling procedure for the Woodroofe-type estimator for discrete data over a finite number of possible values is therefore established. Subsequently, the resulting vector of hazard rate estimators is proved to be asymptotically normal with independent components. Asymptotic normality of the survival function estimator is then established. Sister results for the left-truncating random variable are also proved. Taken together, the resulting joint vector of hazard rate estimates for the lifetime and left-truncation random variables is proved to be the maximum likelihood estimate of the parameters of the conditional joint lifetime and left-truncation distribution given the lifetime has not been left-truncated. A hypothesis test for the shape of the distribution function based on our asymptotic results is derived. Such a test is useful to formally assess the plausibility of the stationarity assumption in length-biased sampling. The finite sample performance of the estimators is investigated in a simulation study. Applicability of the theoretical results in an econometric setting is demonstrated with a subset of data from the Mercedes-Benz 2017-A securitized bond.

math.ST↗

Causal Information Bottleneck Boosts Adversarial Robustness of Deep Neural Network

The information bottleneck (IB) method is a feasible defense solution against adversarial attacks in deep learning. However, this method suffers from the spurious correlation, which leads to the limitation of its further improvement of adversarial robustness. In this paper, we incorporate the causal inference into the IB framework to alleviate such a problem. Specifically, we divide the features obtained by the IB method into robust features (content information) and non-robust features (style information) via the instrumental variables to estimate the causal effects. With the utilization of such a framework, the influence of non-robust features could be mitigated to strengthen the adversarial robustness. We make an analysis of the effectiveness of our proposed method. The extensive experiments in MNIST, FashionMNIST, and CIFAR-10 show that our method exhibits the considerable robustness against multiple adversarial attacks. Our code would be released.

cs.LG↗

$p_{T}$ dispersion of inclusive jets in high-energy nuclear collisions

In this paper, we investigate the medium modifications of $p_{T}$ dispersion($p_{T}D$) of inclusive jets with small radius ($R=0.2$) in Pb+Pb collisions at $\sqrt{s}=2.76$~TeV. The partonic spectrum in the initial hard scattering of elementary collisions are obtained by an event generator POWHEG+PYTHIA, which matches the next-to-leading (NLO) matrix elements with parton showering, and energy loss of fast parton traversing in hot/dense QCD medium is calculated by Monte Carlo simulation within Higher-Twist formalism of jet quenching in heavy-ion collisions. We present the model calculations of normalized $p_{T}D$ distributions for inclusive jets in p+p and Pb+Pb collisions at $\sqrt{s}=2.76$~TeV, which give nice descriptions of ALICE measurements. It is shown that the $p_{T}D$ distributions of inclusive jets in Pb+Pb are shifted to higher $p_{T}D$ region relative to that in p+p. Thus the nuclear modification factor of $p_{T}D$ distributions for inclusive jets is smaller than unity at small $p_{T}D$ region, while larger than one at large $p_{T}D$ region. This behaviour results from more uneven $p_T$ of jet constituents as well as the fraction alteration of gluon/quark initiated jets in heavy-ion collisions. The difference of $p_{T}D$ distributions between groomed and ungroomed jets in Pb+Pb collisions are also discussed.

hep-ph↗

Ergodic problems for contact Hamilton-Jacobi equations

This paper deals with the generalized ergodic problem \[ H(x,u(x),Du(x))=c, \quad x\in M, \] where the unknown is a pair $(c,u)$ of a constant $c \in \mathbb{R}$ and a function $u$ on $M$ for which $u$ is a viscosity solution. We assume $H=H(x,u,p)$ satisfies Tonelli conditions in the argument $p\in T^*_xM$ and the Lipschitz condition in the argument $u\in\R$. For a given $c\in \R$, we first discuss necessary and sufficient conditions for the existence of viscosity solutions. Let $\mathfrak{C}$ denote the set of all real numbers $c$'s for which the above equation admits viscosity solutions. Then we show $\mathfrak{C}$ is an interval, whose endpoints $\x$, $\y$ with $\x\leqslant\y$ can be characterized by a min-max formula and a max-min formula, respectively. The most significant finding is that we figure out the structure of $\mathfrak{C}$ without monotonicity assumptions on $u$.

math.AP↗

A Comprehensive Evaluation of Android ICC Resolution Techniques

Inter-component communication (ICC) is a widely used mechanism in mobile apps, which enables message-based control flow transferring and data passing between Android components. Effective ICC resolution requires precisely identifying entry points, analyzing data values of ICC fields, modeling related framework APIs, etc. Due to various control-flow- and data-flow-related characteristics involved and the lack of oracles for real-world apps, the comprehensive evaluation of ICC resolution techniques is challenging. To fill this gap, we collect multiple-type benchmark suites with 4,104 apps, covering hand-made apps, open-source, and commercial ones. Considering their differences, various evaluation metrics, e.g., number count, graph structure, and reliable oracle based metrics, are adopted on-demand. As the oracle for real-world apps is unavailable, we design a dynamic analysis approach to extract the real ICC links triggered during GUI exploration. By auditing the code implementations, we carefully check the extracted ICCs and confirm 1,680 ones to form a reliable oracle set, in which each ICC is labeled with 25 code characteristic tags. The evaluation performed on six state-of-the-art ICC resolution tools shows that 1) the completeness of static ICC resolution results on real-world apps is not satisfactory, as up to 38%-85% ICCs are missed by tools; 2) many wrongly reported ICCs are sent from or received by only a few components and the graph structure information can help the identification; 3) the efficiency of fundamental tools, like ICC resolution ones, should be optimized in both engineering and research aspects. By investigating both the missed and wrongly reported ICCs, we discuss the strengths of different tools for users and summarize eight common FN/FP patterns in ICC resolution for tool developers.

cs.SE↗

AUGER: Automatically Generating Review Comments with Pre-training Models

Code review is one of the best practices as a powerful safeguard for software quality. In practice, senior or highly skilled reviewers inspect source code and provide constructive comments, considering what authors may ignore, for example, some special cases. The collaborative validation between contributors results in code being highly qualified and less chance of bugs. However, since personal knowledge is limited and varies, the efficiency and effectiveness of code review practice are worthy of further improvement. In fact, it still takes a colossal and time-consuming effort to deliver useful review comments. This paper explores a synergy of multiple practical review comments to enhance code review and proposes AUGER (AUtomatically GEnerating Review comments): a review comments generator with pre-training models. We first collect empirical review data from 11 notable Java projects and construct a dataset of 10,882 code changes. By leveraging Text-to-Text Transfer Transformer (T5) models, the framework synthesizes valuable knowledge in the training stage and effectively outperforms baselines by 37.38% in ROUGE-L. 29% of our automatic review comments are considered useful according to prior studies. The inference generates just in 20 seconds and is also open to training further. Moreover, the performance also gets improved when thoroughly analyzed in case study.

cs.SE↗

Variational construction of connecting orbits between Legendrian graphs

Motivated by the problem of global stability of thermodynamical equilibria in non-equilibrium thermodynamics formulated in a recent paper [12], we introduce some mechanisms for constructing semi-infinite orbits of contact Hamiltonian systems connecting two Legendrian graphs from the viewpoint of Aubry-Mather theory and weak KAM theory.

math.DS↗

FedSSO: A Federated Server-Side Second-Order Optimization Algorithm

In this work, we propose FedSSO, a server-side second-order optimization method for federated learning (FL). In contrast to previous works in this direction, we employ a server-side approximation for the Quasi-Newton method without requiring any training data from the clients. In this way, we not only shift the computation burden from clients to server, but also eliminate the additional communication for second-order updates between clients and server entirely. We provide theoretical guarantee for convergence of our novel method, and empirically demonstrate our fast convergence and communication savings in both convex and non-convex settings.

cs.LG↗