Searcharxiv⌕ Search

arXiv subjects

Cong Ma

Publications and source records attributed to Cong Ma.

90 records · Page 5Linked to original sources

Research of time discrimination circuits for PMT signal readout over large dynamic range in LHAASO WCDA

In the readout electronics of the Water Cerenkov Detector Array (WCDA) in the Large High Altitude Air Shower Observatory (LHAASO), both high-resolution charge and time measurement are required over a dynamic range from 1 photoelectron (P.E.) to 4000 P.E. for the PMT signal readout. In this paper, we present our work on the design of time discrimination circuits in LHAASO WCDA, especially on improvement to reduce the circuit dead time. Several approaches were studied through analysis and simulations, and actual circuits were designed and tested in the laboratory to evaluate the performance. Test results indicate that a time resolution better than 500 ps RMS is achieved in the whole large dynamic range, and the circuit dead time is successfully reduced to less than 200 ns.

physics.ins-det↗

Data and clock transmission interface for the WCDA in LHAASO

The Water Cherenkov Detector Array (WCDA) is one of the major components of the Large High Altitude Air Shower Observatory (LHAASO). In the WCDA, 3600 Photomultiplier Tubes (PMTs) and the Front End Electronics (FEEs) are scattered over a 90000 m2 area, while high precision time measurements (0.5 ns RMS) are required in the readout electronics. To meet this requirement, the clock has to be distributed to the FEEs with high precision. Due to the "triggerless" architecture, high speed data transfer is required based on the TCP/IP protocol. To simplify the readout electronics architecture and be consistent with the whole LHAASO readout electronics, the White Rabbit (WR) switches are used to transfer clock, data, and commands via a single fiber of about 400 meters. In this paper, a prototype of data and clock transmission interface for LHAASO WCDA is developed. The performance tests are conducted and the results indicate that the clock synchronization precision of the data and clock transmission is better than 50 ps. The data transmission throughput can reach 400 Mbps for one FEE board and 180 Mbps for 4 FEE boards sharing one up link port in WR switch, which is better than the requirement of the LHAASO WCDA.

physics.ins-det↗

A Selective Overview of Deep Learning

Deep learning has arguably achieved tremendous success in recent years. In simple words, deep learning uses the composition of many nonlinear functions to model the complex dependency between input features and labels. While neural networks have a long history, recent advances have greatly improved their performance in computer vision, natural language processing, etc. From the statistical and scientific perspective, it is natural to ask: What is deep learning? What are the new characteristics of deep learning, compared with classical methods? What are the theoretical foundations of deep learning? To answer these questions, we introduce common neural network models (e.g., convolutional neural nets, recurrent neural nets, generative adversarial nets) and training techniques (e.g., stochastic gradient descent, dropout, batch normalization) from a statistical point of view. Along the way, we highlight new characteristics of deep learning (including depth and over-parametrization) and explain their practical and theoretical benefits. We also sample recent results on theories of deep learning, many of which are only suggestive. While a complete understanding of deep learning remains elusive, we hope that our perspectives and discussions serve as a stimulus for new statistical research.

stat.ML↗

Gradient Descent with Random Initialization: Fast Global Convergence for Nonconvex Phase Retrieval

This paper considers the problem of solving systems of quadratic equations, namely, recovering an object of interest $\mathbf{x}^{\natural}\in\mathbb{R}^{n}$ from $m$ quadratic equations/samples $y_{i}=(\mathbf{a}_{i}^{\top}\mathbf{x}^{\natural})^{2}$, $1\leq i\leq m$. This problem, also dubbed as phase retrieval, spans multiple domains including physical sciences and machine learning. We investigate the efficiency of gradient descent (or Wirtinger flow) designed for the nonconvex least squares problem. We prove that under Gaussian designs, gradient descent --- when randomly initialized --- yields an $ε$-accurate solution in $O\big(\log n+\log(1/ε)\big)$ iterations given nearly minimal samples, thus achieving near-optimal computational and sample complexities at once. This provides the first global convergence guarantee concerning vanilla gradient descent for phase retrieval, without the need of (i) carefully-designed initialization, (ii) sample splitting, or (iii) sophisticated saddle-point escaping schemes. All of these are achieved by exploiting the statistical models in analyzing optimization algorithms, via a leave-one-out approach that enables the decoupling of certain statistical dependency between the gradient descent iterates and the data.

stat.ML↗

Nonconvex Matrix Factorization from Rank-One Measurements

We consider the problem of recovering low-rank matrices from random rank-one measurements, which spans numerous applications including covariance sketching, phase retrieval, quantum state tomography, and learning shallow polynomial neural networks, among others. Our approach is to directly estimate the low-rank factor by minimizing a nonconvex quadratic loss function via vanilla gradient descent, following a tailored spectral initialization. When the true rank is small, this algorithm is guaranteed to converge to the ground truth (up to global ambiguity) with near-optimal sample complexity and computational complexity. To the best of our knowledge, this is the first guarantee that achieves near-optimality in both metrics. In particular, the key enabler of near-optimal computational guarantees is an implicit regularization phenomenon: without explicit regularization, both spectral initialization and the gradient descent iterates automatically stay within a region incoherent with the measurement vectors. This feature allows one to employ much more aggressive step sizes compared with the ones suggested in prior literature, without the need of sample splitting.

cs.IT↗

Statistical Test of Distance--Duality Relation with Type Ia Supernovae and Baryon Acoustic Oscillations

We test the distance--duality relation $η\equiv d_L / [ (1 + z)^2 d_A ] = 1$ between cosmological luminosity distance ($d_L$) from the JLA SNe Ia compilation (arXiv:1401.4064) and angular-diameter distance ($d_A$) based on Baryon Oscillation Spectroscopic Survey (BOSS; arXiv:1607.03155) and WiggleZ baryon acoustic oscillation measurements (arXiv:1105.2862, arXiv:1204.3674). The $d_L$ measurements are matched to $d_A$ redshift by a statistically consistent compression procedure. With Monte Carlo methods, nontrivial and correlated distributions of $η$ can be explored in a straightforward manner without resorting to a particular evolution template $η(z)$. Assuming independent constraints on cosmological parameters that are necessary to obtain $d_L$ and $d_A$ values, we find 9% constraints consistent with $η= 1$ from the analysis of SNIa + BOSS and an 18% bound results from SNIa + WiggleZ. These results are contrary to previous claims that $η< 1$ has been found close to or above the $1 σ$ level. We discuss the effect of different cosmological parameter inputs and the use of the apparent deviation from distance--duality as a proxy of systematic effects on cosmic distance measurements. The results suggest possible systematic overestimation of SNIa luminosity distances compared with $d_A$ data when a Planck ΛCDM cosmological parameter inference (arXiv:1502.01589) is used to enhance the precision. If interpreted as an extinction correction due to a gray dust component, the effect is broadly consistent with independent observational constraints.

astro-ph.CO↗

Spectral Method and Regularized MLE Are Both Optimal for Top-$K$ Ranking

This paper is concerned with the problem of top-$K$ ranking from pairwise comparisons. Given a collection of $n$ items and a few pairwise comparisons across them, one wishes to identify the set of $K$ items that receive the highest ranks. To tackle this problem, we adopt the logistic parametric model --- the Bradley-Terry-Luce model, where each item is assigned a latent preference score, and where the outcome of each pairwise comparison depends solely on the relative scores of the two items involved. Recent works have made significant progress towards characterizing the performance (e.g. the mean square error for estimating the scores) of several classical methods, including the spectral method and the maximum likelihood estimator (MLE). However, where they stand regarding top-$K$ ranking remains unsettled. We demonstrate that under a natural random sampling model, the spectral method alone, or the regularized MLE alone, is minimax optimal in terms of the sample complexity --- the number of paired comparisons needed to ensure exact top-$K$ identification, for the fixed dynamic range regime. This is accomplished via optimal control of the entrywise error of the score estimates. We complement our theoretical studies by numerical experiments, confirming that both methods yield low entrywise errors for estimating the underlying scores. Our theory is established via a novel leave-one-out trick, which proves effective for analyzing both iterative and non-iterative procedures. Along the way, we derive an elementary eigenvector perturbation bound for probability transition matrices, which parallels the Davis-Kahan $\sinΘ$ theorem for symmetric matrices. This also allows us to close the gap between the $\ell_2$ error upper bound for the spectral method and the minimax lower limit.

stat.ML↗

Trajectory Factory: Tracklet Cleaving and Re-connection by Deep Siamese Bi-GRU for Multiple Object Tracking

Multi-Object Tracking (MOT) is a challenging task in the complex scene such as surveillance and autonomous driving. In this paper, we propose a novel tracklet processing method to cleave and re-connect tracklets on crowd or long-term occlusion by Siamese Bi-Gated Recurrent Unit (GRU). The tracklet generation utilizes object features extracted by CNN and RNN to create the high-confidence tracklet candidates in sparse scenario. Due to mis-tracking in the generation process, the tracklets from different objects are split into several sub-tracklets by a bidirectional GRU. After that, a Siamese GRU based tracklet re-connection method is applied to link the sub-tracklets which belong to the same object to form a whole trajectory. In addition, we extract the tracklet images from existing MOT datasets and propose a novel dataset to train our networks. The proposed dataset contains more than 95160 pedestrian images. It has 793 different persons in it. On average, there are 120 images for each person with positions and sizes. Experimental results demonstrate the advantages of our model over the state-of-the-art methods on MOT16.

cs.CV↗

Inter-Subject Analysis: Inferring Sparse Interactions with Dense Intra-Graphs

We develop a new modeling framework for Inter-Subject Analysis (ISA). The goal of ISA is to explore the dependency structure between different subjects with the intra-subject dependency as nuisance. It has important applications in neuroscience to explore the functional connectivity between brain regions under natural stimuli. Our framework is based on the Gaussian graphical models, under which ISA can be converted to the problem of estimation and inference of the inter-subject precision matrix. The main statistical challenge is that we do not impose sparsity constraint on the whole precision matrix and we only assume the inter-subject part is sparse. For estimation, we propose to estimate an alternative parameter to get around the non-sparse issue and it can achieve asymptotic consistency even if the intra-subject dependency is dense. For inference, we propose an "untangle and chord" procedure to de-bias our estimator. It is valid without the sparsity assumption on the inverse Hessian of the log-likelihood function. This inferential method is general and can be applied to many other statistical problems, thus it is of independent theoretical interest. Numerical experiments on both simulated and brain imaging data validate our methods and theory.

stat.ME↗

Application of Bayesian graphs to SN Ia data analysis and compression

Bayesian graphical models are an efficient tool for modelling complex data and derive self-consistent expressions of the posterior distribution of model parameters. We apply Bayesian graphs to perform statistical analyses of Type Ia supernova (SN Ia) luminosity distance measurements from the joint light-curve analysis (JLA) data set. In contrast to the $χ^2$ approach used in previous studies, the Bayesian inference allows us to fully account for the standard-candle parameter dependence of the data covariance matrix. Comparing with $χ^2$ analysis results, we find a systematic offset of the marginal model parameter bounds. We demonstrate that the bias is statistically significant in the case of the SN Ia standardization parameters with a maximal 6 $σ$ shift of the SN light-curve colour correction. In addition, we find that the evidence for a host galaxy correction is now only 2.4 $σ$. Systematic offsets on the cosmological parameters remain small, but may increase by combining constraints from complementary cosmological probes. The bias of the $χ^2$ analysis is due to neglecting the parameter-dependent log-determinant of the data covariance, which gives more statistical weight to larger values of the standardization parameters. We find a similar effect on compressed distance modulus data. To this end, we implement a fully consistent compression method of the JLA data set that uses a Gaussian approximation of the posterior distribution for fast generation of compressed data. Overall, the results of our analysis emphasize the need for a fully consistent Bayesian statistical approach in the analysis of future large SN Ia data sets.

astro-ph.CO↗

The Analog Front-end Prototype Electronics Designed for LHAASO WCDA

In the readout electronics of the Water Cerenkov Detector Array (WCDA) in the Large High Altitude Air Shower Observatory (LHAASO) experiment, both high-resolution charge and time measurement are required over a dynamic range from 1 photoelectron (P.E.) to 4000 P.E. The Analog Front-end (AFE) circuit is one of the crucial parts in the whole readout electronics. We designed and optimized a prototype of the AFE through parameter calculation and circuit simulation, and conducted initial electronics tests on this prototype to evaluate its performance. Test results indicate that the charge resolution is better than 1% @ 4000 P.E. and remains better than 10% @ 1 P.E., and the time resolution is better than 0.5 ns RMS, which is better than application requirement.

physics.ins-det↗

Panther: Fast Top-k Similarity Search in Large Networks

Estimating similarity between vertices is a fundamental issue in network analysis across various domains, such as social networks and biological networks. Methods based on common neighbors and structural contexts have received much attention. However, both categories of methods are difficult to scale up to handle large networks (with billions of nodes). In this paper, we propose a sampling method that provably and accurately estimates the similarity between vertices. The algorithm is based on a novel idea of random path, and an extended method is also presented, to enhance the structural similarity when two vertices are completely disconnected. We provide theoretical proofs for the error-bound and confidence of the proposed algorithm. We perform extensive empirical study and show that our algorithm can obtain top-k similar vertices for any vertex in a network approximately 300x faster than state-of-the-art methods. We also use identity resolution and structural hole spanner finding, two important applications in social networks, to evaluate the accuracy of the estimated similarities. Our experimental results demonstrate that the proposed algorithm achieves clearly better performance than several alternative methods.

cs.SI↗

Testing the Copernican Principle with Hubble Parameter

Using the longitudinal expression of Hubble expansion rate for the general Lemaître-Tolman-Bondi (LTB) metric as a function of cosmic time, we examine the scale on which the Copernican Principle holds in the context of a void model. By way of performing parameter estimation on the CGBH void model, we show that the Hubble parameter data favors a void with characteristic radius of 2 ~ 3 Gpc. This brings the void model closer, but not yet enough, to harmony with observational indications given by the background kinetic Sunyaev-Zel'dovich effect and the normalization of near-infrared galaxy luminosity function. However, the test of such void models may ultimately lie in the future detection of the discrepancy between longitudinal and transverse expansion rates, a touchstone of inhomogeneous models. With the proliferation of observational Hubble parameter data and future large-scale structure observation, a definitive test could be performed on the question of cosmic homogeneity. Particularly, the spherical LTB void models have been ruled out, but more general non-spherical inhomogeneities still need to be tested by observation. In this paper, we utilise a spherical void model to provide guidelines into how observational tests may be done with more general models in the future.

astro-ph.CO↗

Cosmological constraints on holographic dark energy models under the energy conditions

We study the holographic and agegraphic dark energy models without interaction using the latest observational Hubble parameter data (OHD), the Union2.1 compilation of type Ia supernovae (SNIa), and the energy conditions. Scenarios of dark energy are distinguished by the cut-off of cosmic age, conformal time, and event horizon. The best-fit value of matter density for the three scenarios almost steadily located at $Ω_{m0}=0.26$ by the joint constraint. For the agegraphic models, they can be recovered to the standard cosmological model when the constant $c$ which presents the fraction of dark energy approaches to infinity. Absence of upper limit of $c$ by the joint constraint demonstrates the recovery possibility. Using the fitted result, we also reconstruct the current equation of state of dark energy at different scenarios, respectively. Employing the model criteria $χ^2_{\textrm{min}}/dof$, we find that conformal time model is the worst, but they can not be distinguished clearly. Comparing with the observational constraints, we find that SEC is fulfilled at redshift $0.2 \lesssim z \lesssim 0.3$ with $1σ$ confidence level. We also find that NEC gives a meaningful constraint for the event horizon cut-off model, especially compared with OHD only. We note that the energy condition maybe could play an important role in the interacting models because of different degeneracy between $Ω_m$ and constant $c$.

astro-ph.CO↗

Reconstructing the History of Energy Condition Violation from Observational Data

We study the likelihood of energy condition violations in the history of the Universe. Our method is based on a set of functions that characterize energy condition violation. FLRW cosmological models are built around these "indication functions". By computing the Fisher matrix of model parameters using type Ia supernova and Hubble parameter data, we extract the principal modes of these functions' redshift evolution. These modes allow us to obtain general reconstructions of energy condition violation history independent of the dark energy model. We find that the data suggest a history of strong energy condition violation, but the null and dominant energy conditions are likely to be fulfilled. Implications for dark energy models are discussed.

astro-ph.CO↗

Power of Observational Hubble Parameter Data: a Figure of Merit Exploration

We use simulated Hubble parameter data in the redshift range 0 \leq z \leq 2 to explore the role and power of observational H(z) data in constraining cosmological parameters of the ΛCDM model. The error model of the simulated data is empirically constructed from available measurements and scales linearly as z increases. By comparing the median figures of merit calculated from simulated datasets with that of current type Ia supernova data, we find that as many as 64 further independent measurements of H(z) are needed to match the parameter constraining power of SNIa. If the error of H(z) could be lowered to 3%, the same number of future measurements would be needed, but then the redshift coverage would only be required to reach z = 1. We also show that accurate measurements of the Hubble constant H_0 can be used as priors to increase the H(z) data's figure of merit.

astro-ph.CO↗

Constraints on the Dark Side of the Universe and Observational Hubble Parameter Data

This paper is a review on the observational Hubble parameter data that have gained increasing attention in recent years for their illuminating power on the dark side of the universe --- the dark matter, dark energy, and the dark age. Currently, there are two major methods of independent observational H(z) measurement, which we summarize as the "differential age method" and the "radial BAO size method". Starting with fundamental cosmological notions such as the spacetime coordinates in an expanding universe, we present the basic principles behind the two methods. We further review the two methods in greater detail, including the source of errors. We show how the observational H(z) data presents itself as a useful tool in the study of cosmological models and parameter constraint, and we also discuss several issues associated with their applications. Finally, we point the reader to a future prospect of upcoming observation programs that will lead to some major improvements in the quality of observational H(z) data.

astro-ph.CO↗

Numerical Strategies of Computing the Luminosity Distance

We propose two efficient numerical methods of evaluating the luminosity distance in the spatially flat ΛCDM universe. The first method is based on the Carlson symmetric form of elliptic integrals, which is highly accurate and can replace numerical quadratures. The second method, using a modified version of Hermite interpolation, is less accurate but involves only basic numerical operations and can be easily implemented. We compare our methods with other numerical approximation schemes and explore their respective features and limitations. Possible extensions of these methods to other cosmological models are also discussed.

astro-ph.IM↗