SearcharxivSearch

arXiv subjects

Chi-Ken Lu

Publications and source records attributed to Chi-Ken Lu.

At least 19 recordsLinked to original sources

Analytical Charge Density Profile of Vortex Core in Weak-Coupling Superconductor

Self-consistent Bogoliubov-de Gennes calculations have long shown that solving the Poisson equation inside a superconducting vortex turns a one-signed charge depletion into a modulation that alternates in sign with period $\pi/k_F$. We give an elementary account of that result. Taking the Caroli-de Gennes-Matricon bound states in a step-like gap, we show that the normalization of the bound-state spinor is nearly independent of angular momentum, which collapses the mode sum into closed form. Inside the core the vortex winding removes one Bessel channel from a completeness sum, so the density vanishes on the vortex line and carries Friedel-like oscillations of wavevector $2k_F$; outside it the sum gives a $1/r$ envelope decaying over a coherence length, with a residual ripple. The bound-state charge does not integrate to zero, so neutrality obliges the extended states to compensate it exactly. That compensation is complete at long wavelength but fails at the diameter of the Fermi circle, and what survives is a sign-alternating $2k_F$ modulation reduced only by $4k_F^2/(4k_F^2+k_{TF}^2)$, a factor lying between one-half and three-quarters for any metal. The oscillation is therefore not a delicate effect but a consequence of neutrality and the inefficiency of screening at large momentum transfer: in the screened total the smooth terms cancel and only the ripple is left.

cond-mat.supr-con

A Log-Linear Analytics Approach to Cost Model Regularization for Inpatient Stays through Diagnostic Code Merging

Cost models in healthcare research must balance interpretability, accuracy, and parameter consistency. However, interpretable models often struggle to achieve both accuracy and consistency. Ordinary least squares (OLS) models for high-dimensional regression can be accurate but fail to produce stable regression coefficients over time when using highly granular ICD-10 diagnostic codes as predictors. This instability arises because many ICD-10 codes are infrequent in healthcare datasets. While regularization methods such as Ridge can address this issue, they risk discarding important predictors. Here, we demonstrate that reducing the granularity of ICD-10 codes is an effective regularization strategy within OLS while preserving the representation of all diagnostic code categories. By truncating ICD-10 codes from seven characters to six or fewer, we reduce the dimensionality of the regression problem while maintaining model interpretability and consistency. Mathematically, the merging of predictors in OLS leads to increased trace of the Hessian matrix, which reduces the variance of coefficient estimation. Our findings explain why broader diagnostic groupings like DRGs and HCC codes are favored over highly granular ICD-10 codes in real-world risk adjustment and cost models.

cs.LG

Bayesian inference with finitely wide neural networks

The analytic inference, e.g. predictive distribution being in closed form, may be an appealing benefit for machine learning practitioners when they treat wide neural networks as Gaussian process in Bayesian setting. The realistic widths, however, are finite and cause weak deviation from the Gaussianity under which partial marginalization of random variables in a model is straightforward. On the basis of multivariate Edgeworth expansion, we propose a non-Gaussian distribution in differential form to model a finite set of outputs from a random neural network, and derive the corresponding marginal and conditional properties. Thus, we are able to derive the non-Gaussian posterior distribution in Bayesian regression task. In addition, in the bottlenecked deep neural networks, a weight space representation of deep Gaussian process, the non-Gaussianity is investigated through the marginal kernel.

cond-mat.dis-nn

On Connecting Deep Trigonometric Networks with Deep Gaussian Processes: Covariance, Expressivity, and Neural Tangent Kernel

Deep Gaussian Process (DGP) as a model prior in Bayesian learning intuitively exploits the expressive power in function composition. DGPs also offer diverse modeling capabilities, but inference is challenging because marginalization in latent function space is not tractable. With Bochner's theorem, DGP with squared exponential kernel can be viewed as a deep trigonometric network consisting of the random feature layers, sine and cosine activation units, and random weight layers. In the wide limit with a bottleneck, we show that the weight space view yields the same effective covariance functions which were obtained previously in function space. Also, varying the prior distributions over network parameters is equivalent to employing different kernels. As such, DGPs can be translated into the deep bottlenecked trig networks, with which the exact maximum a posteriori estimation can be obtained. Interestingly, the network representation enables the study of DGP's neural tangent kernel, which may also reveal the mean of the intractable predictive distribution. Statistically, unlike the shallow networks, deep networks of finite width have covariance deviating from the limiting kernel, and the inner and outer widths may play different roles in feature learning. Numerical simulations are present to support our findings.

cs.LG

Conditional Deep Gaussian Processes: empirical Bayes hyperdata learning

It is desirable to combine the expressive power of deep learning with Gaussian Process (GP) in one expressive Bayesian learning model. Deep kernel learning showed success in adopting a deep network for feature extraction followed by a GP used as function model. Recently,it was suggested that, albeit training with marginal likelihood, the deterministic nature of feature extractor might lead to overfitting while the replacement with a Bayesian network seemed to cure it. Here, we propose the conditional Deep Gaussian Process (DGP) in which the intermediate GPs in hierarchical composition are supported by the hyperdata and the exposed GP remains zero mean. Motivated by the inducing points in sparse GP, the hyperdata also play the role of function supports, but are hyperparameters rather than random variables. We follow our previous moment matching approach to approximate the marginal prior for conditional DGP with a GP carrying an effective kernel. Thus, as in empirical Bayes, the hyperdata are learned by optimizing the approximate marginal likelihood which implicitly depends on the hyperdata via the kernel. We shall show the equivalence with the deep kernel learning in the limit of dense hyperdata in latent space. However, the conditional DGP and the corresponding approximate inference enjoy the benefit of being more Bayesian than deep kernel learning. Preliminary extrapolation results demonstrate expressive power from the depth of hierarchy by exploiting the exact covariance and hyperdata learning, in comparison with GP kernel composition, DGP variational inference and deep kernel learning. We also address the non-Gaussian aspect of our model as well as way of upgrading to a full Bayes inference.

cs.LG

Conditional Deep Gaussian Processes: multi-fidelity kernel learning

Deep Gaussian Processes (DGPs) were proposed as an expressive Bayesian model capable of a mathematically grounded estimation of uncertainty. The expressivity of DPGs results from not only the compositional character but the distribution propagation within the hierarchy. Recently, [1] pointed out that the hierarchical structure of DGP well suited modeling the multi-fidelity regression, in which one is provided sparse observations with high precision and plenty of low fidelity observations. We propose the conditional DGP model in which the latent GPs are directly supported by the fixed lower fidelity data. Then the moment matching method in [2] is applied to approximate the marginal prior of conditional DGP with a GP. The obtained effective kernels are implicit functions of the lower-fidelity data, manifesting the expressivity contributed by distribution propagation within the hierarchy. The hyperparameters are learned via optimizing the approximate marginal likelihood. Experiments with synthetic and high dimensional data show comparable performance against other multi-fidelity regression methods, variational inference, and multi-output GP. We conclude that, with the low fidelity data and the hierarchical DGP structure, the effective kernel encodes the inductive bias for true function allowing the compositional freedom discussed in [3,4].

cs.LG

Interpretable deep Gaussian processes with moments

Deep Gaussian Processes (DGPs) combine the expressiveness of Deep Neural Networks (DNNs) with quantified uncertainty of Gaussian Processes (GPs). Expressive power and intractable inference both result from the non-Gaussian distribution over composition functions. We propose interpretable DGP based on approximating DGP as a GP by calculating the exact moments, which additionally identify the heavy-tailed nature of some DGP distributions. Consequently, our approach admits interpretation as both NNs with specified activation functions and as a variational approximation to DGP. We identify the expressivity parameter of DGP and find non-local and non-stationary correlation from DGP composition. We provide general recipes for deriving the effective kernels for DGP of two, three, or infinitely many layers, composed of homogeneous or heterogeneous kernels. Results illustrate the expressiveness of our effective kernels through samples from the prior and inference on simulated and real data and demonstrate advantages of interpretability by analysis of analytic forms, and draw relations and equivalences across kernels.

cs.LG

Standing Wave Decomposition Gaussian Process

We propose a Standing Wave Decomposition (SWD) approximation to Gaussian Process regression (GP). GP involves a costly matrix inversion operation, which limits applicability to large data analysis. For an input space that can be approximated by a grid and when correlations among data are short-ranged, the kernel matrix inversion can be replaced by analytic diagonalization using the SWD. We show that this approach applies to uni- and multi-dimensional input data, extends to include longer-range correlations, and the grid can be in a latent space and used as inducing points. Through simulations, we show that our approximate method applied to the squared exponential kernel outperforms existing methods in predictive accuracy per unit time in the regime where data are plentiful. Our SWD-GP is recommended for regression analyses where there is a relatively large amount of data and/or there are constraints on computation time.

stat.ML

Robust pinning of magnetic moments in pyrochlore iridates

Pyrochlore iridates A2Ir2O7 (A = rare earth elements, Y or Bi) hold great promise for realizing novel electronic and magnetic states owing to the interplay of spin-orbit coupling, electron correlation and geometrical frustration. A prominent example is the formation of all-in/all-out (AIAO)antiferromagnetic order in the Ir4+ sublattice that comprises of corner-sharing tetrahedra. Here we report on an unusual magnetic phenomenon, namely a cooling-field induced shift of magnetic hysteresis loop along magnetization axis, and its possible origin in pyrochlore iridates with non-magnetic Ir defects (e.g. Ir3+). In a simple model, we attribute the magnetic hysteresis loop to the formation of ferromagnetic droplets in the AIAO antiferromagnetic background. The weak ferromagnetism originates from canted antiferromagnetic order of the Ir4+ moments surrounding each non-magnetic Ir defect. The shift of hysteresis loop can be understood quantitatively based on an exchange-bias like effect in which the moments at the shell of the FM droplets are pinned by the AIAO AFM background via mainly the Heisenberg (J) and Dzyaloshinsky-Moriya (D) interactions. The magnetic pinning is stable and robust against the sweeping cycle and sweeping field up to 35 T, which is possibly related to the magnetic octupolar nature of the AIAO order.

cond-mat.str-el

New class of 3D topological insulator in double perovskite

We predict a new class of three-dimensional topological insulators (TIs) in which the spin-orbit coupling (SOC) can more effectively generate a large band gap at $Γ$ point. The band gap of conventional TI such as Bi$_2$Se$_3$ is mainly limited by two factors, the strength of SOC and, from electronic structure perspective, the band gap when SOC is absent. While the former is an atomic property, we find that the latter can be minimized in a generic rock-salt lattice model in which a stable crossing of bands {\it at} the Fermi level along with band character inversion occurs for a range of parameters in the absence of SOC. Thus, large-gap TI's or TI's comprised of lighter elements can be expected. In fact, we find by performing first-principle calculations that the model applies to a class of double perovskites A$_2$BiXO$_6$ (A = Ca, Sr, Ba; X = Br, I) and the band gap is predicted up to 0.55 eV. Besides, more detailed calculations considering realistic surface structure indicate that the Dirac cones are robust against the presence of dangling bond at the boundary with a specific termination.

cond-mat.mes-hall

Magnetic order and spin-orbit coupled Mott state in double perovskite (La$_{1-x}$Sr$_x$)$_2$CuIrO$_6$

Double-perovskite oxides that contain both 3d and 5d transition metal elements have attracted growing interest as they provide a model system to study the interplay of strong electron interaction and large spin-orbit coupling (SOC). Here, we report on experimental and theoretical studies of the magnetic and electronic properties of double-perovskites (La$_{1-x}$Sr$_x$)$_2$CuIrO$_6$ ($x$ = 0.0, 0.1, 0.2, and 0.3). The undoped La$_2$CuIrO$_6$ undergoes a magnetic phase transition from paramagnetism to antiferromagnetism at T$_N$ $\sim$ 74 K and exhibits a weak ferromagnetic behavior below $T_C$ $\sim$ 52 K. Two-dimensional magnetism that was observed in many other Cu-based double-perovskites is absent in our samples, which may be due to the existence of weak Cu-Ir exchange interaction. First-principle density-functional theory (DFT) calculations show canted antiferromagnetic (AFM) order in both Cu$^{2+}$ and Ir$^{4+}$ sublattices, which gives rise to weak ferromagnetism. Electronic structure calculations suggest that La$_2$CuIrO$_6$ is an SOC-driven Mott insulator with an energy gap of $\sim$ 0.3 eV. Sr-doping decreases the magnetic ordering temperatures ($T_N$ and $T_C$) and suppresses the electrical resistivity. The high temperatures resistivity can be fitted using a variable-range-hopping model, consistent with the existence of disorders in these double-pervoskite compounds.

cond-mat.str-el

Friedel oscillation near a van Hove singularity in two-dimensional Dirac materials

We consider Friedel oscillation in the two-dimensional Dirac materials when Fermi level is near the van Hove singularity. Twisted graphene bilayer and the surface state of topological crystalline insulator are the representative materials which show low-energy saddle points that are feasible to probe by gating. We approximate the Fermi surface near saddle point with a hyperbola and calculate the static Lindhard response function. Employing a theorem of Lighthill, the induced charge density $δn$ due to an impurity is obtained and the algebraic decay of $δn$ is determined by the singularity of the static response function. Although a hyperbolic Fermi surface is rather different from a circular one, the static Lindhard response function in the present case shows a singularity similar with the response function associated with circular Fermi surface, which leads to the $δn\propto R^{-2}$ at large distance $R$. The dependences of charge density on the Fermi energy are different. Consequently, it is possible to observe in twisted graphene bilayer the evolution that $δn\propto R^{-3}$ near Dirac point changes to $δn\propto R^{-2}$ above the saddle point. Measurements using scanning tunnelling microscopy around the impurity sites could verify the prediction.

cond-mat.mes-hall

Manifestations of topological band crossings in bulk entanglement spectrum: An analytical study for integer quantum Hall states

We consider integer quantum Hall states and calculate bulk entanglement spectrum by formulating the correlation matrix in guiding center representation. Our analytical approach is based on the projection operator with redefining the inner product of states in Hilbert space to take care of the restriction imposed by the (rectangle-tiled) checkerboard partition. The resultant correlation matrix contains the coupling constants between states of different guiding centers parameterized by magnetic length and the period of partition. We find various band-crossings by tuning the flux $Φ$ threading each chekerborad pixel and by changing filling factor $ν$. When $ν=1$ and $Φ=2π$, or $ν=2$ and $Φ=π$, one Dirac band crossing is found. For $ν=1$ and $Φ=π$, the band crossings are in the form of nodal line, enclosing the Brillouin zone. As for $ν=2$ and $Φ=2π$, the doubled Dirac point, or the quadratic point, is seen. Besides, we infer that the quadratic point is protected by C$_4$ symmetry of the checkerboard partition since it evolves into two separate Dirac points when the symmetry is lowered to C$_2$. In addition, we also identify the emerging symmetries responsible for the symmetric bulk entanglement spectra, which are absent in the underlying quantum Hall states.

cond-mat.str-el

Probing Layer Localization in Twisted Graphene Bilayers via Cyclotron Resonance

Electron wavefunctions in twisted bilayer graphene may have a strong single layer character or be intrinsically delocalized between layers, with their nature often determined by how energetically close they are to the Dirac point. In this paper, we demonstrate that in magnetic fields, optical absorption (cyclotron resonance) spectra contain signatures which may be used to distinguish the nature of these wavefunctions at low energies, as well as to locate low energy critical points in the zero-field energy spectrum. Optical absorption for two different configurations -- electric field parallel and perpendicular to the bilayer -- are calculated, which are shown to have different selection rules with respect to which states are connected by the perturbation. Interlayer bias further distinguishes transitions involving states of a single layer nature from those with support in both layers. For doped systems, a sharp increase in intra-Landau level absorption occurs with increasing field as the level passes through the zero-field saddle point energy, where the states change character from single layer to bilayer.

cond-mat.mes-hall

Zero modes of the generalized fermion-vortex system in magnetic field

We show that Dirac fermions moving in two spatial dimensions with a generalized dispersion $E\sim p^N$, subject to an external magnetic field and coupled to a complex scalar field carrying a vortex defect with winding number $Q$ acquire $NQ$ zero modes. This is the same as in the absence of the magnetic field. Our proof is based on selection rules in the Landau level basis that dictate the existence and the number of the zero modes. We show that the result is insensitive to the choice of geometry and is naturally extended to general field profiles, where we also derive a generalization of the Aharonov-Casher theorem. Experimental consequences of our results are briefly discussed.

cond-mat.mes-hall

Magnetic Breakdown in Twisted Bilayer Graphene

We consider magnetic breakdown in twisted bilayer graphene where electrons may hop between semiclassical $k$-space trajectories in different layers. These trajectories within a doubled Brillouin zone constitute a network in which an $S$-matrix at each saddle point is used to model tunneling between different layers. Matching of the semiclassical wavefunctions throughout the network determines the energy spectrum. Semiclassical orbits with energies well below that of the saddle points are Landau levels of the Dirac points in each layer. These continuously evolve into {\it both} electron-like and hole-like levels above the saddle point energy. Possible experimental signatures are discussed.

cond-mat.mes-hall

Conserved charges of order-parameter textures in Dirac systems

A simple expression for the induced fermion current in the presence of a texture in mass-order-parameters in two-dimensional condensed-matter Dirac systems is derived using the representation theory of Clifford algebras. In particular, it is shown that every texture in three mutually anticommuting order parameters, in graphene for example, implies an induced density of a properly defined conserved charge. The sufficient condition for the general charge to be the familiar electrical charge is that the remaining two anticommuting order parameters allowed by the particle-hole symmetry are the two phase components of some superconducting order. This allows eight different types of electrically charged textures in graphene or in the $π$-flux Hamiltonian on the square lattice. Generalized charge of mass-textures on the surfaces of thin films of topological insulators, or in spinless Dirac fermions hopping on the honeycomb lattice is also discussed.

cond-mat.str-el

Zero modes and charged Skyrmions in graphene bilayer

We show that the electric charge of the Skyrmion in the vector order parameters that characterize the quantum anomalous spin Hall state and the layer-antiferromagnet in a graphene bilayer is four and zero, respectively. The result is based on the demonstration that a vortex configuration in two broken symmetry states in bilayer graphene with the quadratic band crossing has the number of zero modes doubled relative to the single layer. The doubling can be understood as a result of Kramers' theorem implied by the "pseudo time reversal" symmetry of the vortex Hamiltonian. Disordering the quantum anomalous spin Hall state by Skyrmion condensation should produce a superconductor of an elementary charge 4e.

cond-mat.str-el