Searcharxiv⌕ Search

arXiv subjects

Rui Luo

Publications and source records attributed to Rui Luo.

At least 109 records · Page 6Linked to original sources

On the FRB luminosity function -- II. Event rate density

The luminosity function of Fast Radio Bursts (FRBs), defined as the event rate per unit cosmic co-moving volume per unit luminosity, may help to reveal the possible origins of FRBs and design the optimal searching strategy. With the Bayesian modelling, we measure the FRB luminosity function using 46 known FRBs. Our Bayesian framework self-consistently models the selection effects, including the survey sensitivity, the telescope beam response, and the electron distributions from Milky Way / the host galaxy / local environment of FRBs. Different from the previous companion paper, we pay attention to the FRB event rate density and model the event counts of FRB surveys based on the Poisson statistics. Assuming a Schechter luminosity function form, we infer (at the 95% confidence level) that the characteristic FRB event rate density at the upper cut-off luminosity $L^*=2.9_{-1.7}^{+11.9}\times10^{44}\,\rm erg\, s^{-1}$ is $ϕ^*=339_{-313}^{+1074}\,\rm Gpc^{-3}\, yr^{-1}$, the power-law index is $α=-1.79_{-0.35}^{+0.31}$, and the lower cut-off luminosity is $L_0\le9.1\times10^{41}\,\rm erg\, s^{-1}$. The event rate density of FRBs is found to be $3.5_{-2.4}^{+5.7}\times10^4\,\rm Gpc^{-3}\, yr^{-1}$ above $10^{42}\,\rm erg\, s^{-1}$, $5.0_{-2.3}^{+3.2}\times10^3\,\rm Gpc^{-3}\, yr^{-1}$ above $10^{43}\,\rm erg\, s^{-1}$, and $3.7_{-2.0}^{+3.5}\times10^2\,\rm Gpc^{-3}\, yr^{-1}$ above $10^{44}\,\rm erg\, s^{-1}$. As a result, we find that, for searches conducted at 1.4 GHz, the optimal diameter of single-dish radio telescopes to detect FRBs is 30-40 m. The possible astrophysical implications of the measured event rate density are also discussed in the current paper.

astro-ph.HE↗

Modelling Bounded Rationality in Multi-Agent Interactions by Generalized Recursive Reasoning

Though limited in real-world decision making, most multi-agent reinforcement learning (MARL) models assume perfectly rational agents -- a property hardly met due to individual's cognitive limitation and/or the tractability of the decision problem. In this paper, we introduce generalized recursive reasoning (GR2) as a novel framework to model agents with different \emph{hierarchical} levels of rationality; our framework enables agents to exhibit varying levels of "thinking" ability thereby allowing higher-level agents to best respond to various less sophisticated learners. We contribute both theoretically and empirically. On the theory side, we devise the hierarchical framework of GR2 through probabilistic graphical models and prove the existence of a perfect Bayesian equilibrium. Within the GR2, we propose a practical actor-critic solver, and demonstrate its convergent property to a stationary point in two-player games through Lyapunov analysis. On the empirical side, we validate our findings on a variety of MARL benchmarks. Precisely, we first illustrate the hierarchical thinking process on the Keynes Beauty Contest, and then demonstrate significant improvements compared to state-of-the-art opponent modeling baselines on the normal-form games and the cooperative navigation benchmark.

cs.AI↗

FRB 171019: An event of binary neutron star merger?

The fast radio burst, FRB 171019, was relatively bright when discovered first by ASKAP, but was identified as a repeater with three faint bursts detected later by GBT and CHIME. These observations lead to the discussion of whether the first bright burst shares the same mechanism with the following repeating bursts. A model of binary neutron star merger is proposed for FRB 171019, in which the first bright burst occurred during the merger event, while the subsequent repeating bursts are starquake-induced, and generally fainter, as the energy release rate for the starquakes can hardly exceed that of the catastrophic merger event. This scenario is consistent with the observation that no burst detected is as bright as the first one.

astro-ph.HE↗

A weakly supervised adaptive triplet loss for deep metric learning

We address the problem of distance metric learning in visual similarity search, defined as learning an image embedding model which projects images into Euclidean space where semantically and visually similar images are closer and dissimilar images are further from one another. We present a weakly supervised adaptive triplet loss (ATL) capable of capturing fine-grained semantic similarity that encourages the learned image embedding models to generalize well on cross-domain data. The method uses weakly labeled product description data to implicitly determine fine grained semantic classes, avoiding the need to annotate large amounts of training data. We evaluate on the Amazon fashion retrieval benchmark and DeepFashion in-shop retrieval data. The method boosts the performance of triplet loss baseline by 10.6% on cross-domain data and out-performs the state-of-art model on all evaluation metrics.

cs.CV↗

Wasserstein Robust Reinforcement Learning

Reinforcement learning algorithms, though successful, tend to over-fit to training environments hampering their application to the real-world. This paper proposes $\text{W}\text{R}^{2}\text{L}$ -- a robust reinforcement learning algorithm with significant robust performance on low and high-dimensional control tasks. Our method formalises robust reinforcement learning as a novel min-max game with a Wasserstein constraint for a correct and convergent solver. Apart from the formulation, we also propose an efficient and scalable solver following a novel zero-order optimisation method that we believe can be useful to numerical optimisation in general. We empirically demonstrate significant gains compared to standard and robust state-of-the-art algorithms on high-dimensional MuJuCo environments.

cs.LG↗

Non-detection of fast radio bursts from six gamma-ray burst remnants with possible magnetar engines

The analogy of the host galaxy of the repeating fast radio burst (FRB) source FRB 121102 and those of long gamma-ray bursts (GRBs) and super-luminous supernovae (SLSNe) has led to the suggestion that young magnetars born in GRBs and SLSNe could be the central engine of repeating FRBs. We test such a hypothesis by performing dedicated observations of the remnants of six GRBs with evidence of having a magnetar central engine using the Arecibo telescope and the Robert C. Byrd Green Bank Telescope (GBT). A total of $\sim 20$ hrs of observations of these sources did not detect any FRB from these remnants. Under the assumptions that all these GRBs left behind a long-lived magnetar and that the bursting rate of FRB 121102 is typical for a magnetar FRB engine, we estimate a non-detection probability of $8.9\times10^{-6}$. Even though these non-detections cannot exclude the young magnetar model of FRBs, we place constraints on the burst rate and luminosity function of FRBs from these GRB targets.

astro-ph.HE↗

Automatic Radiology Report Generation based on Multi-view Image Fusion and Medical Concept Enrichment

Generating radiology reports is time-consuming and requires extensive expertise in practice. Therefore, reliable automatic radiology report generation is highly desired to alleviate the workload. Although deep learning techniques have been successfully applied to image classification and image captioning tasks, radiology report generation remains challenging in regards to understanding and linking complicated medical visual contents with accurate natural language descriptions. In addition, the data scales of open-access datasets that contain paired medical images and reports remain very limited. To cope with these practical challenges, we propose a generative encoder-decoder model and focus on chest x-ray images and reports with the following improvements. First, we pretrain the encoder with a large number of chest x-ray images to accurately recognize 14 common radiographic observations, while taking advantage of the multi-view images by enforcing the cross-view consistency. Second, we synthesize multi-view visual features based on a sentence-level attention mechanism in a late fusion fashion. In addition, in order to enrich the decoder with descriptive semantics and enforce the correctness of the deterministic medical-related contents such as mentions of organs or diagnoses, we extract medical concepts based on the radiology reports in the training data and fine-tune the encoder to extract the most frequent medical concepts from the x-ray images. Such concepts are fused with each decoding step by a word-level attention model. The experimental results conducted on the Indiana University Chest X-Ray dataset demonstrate that the proposed model achieves the state-of-the-art performance compared with other baseline approaches.

eess.IV↗

Photon-level tuning of photonic nanocavities

Energy-efficient optical control of photonic device properties is crucial for diverse photonic signal processing. Here we demonstrate extremely efficient optical tuning of photonic nanocavities, with only photon-level optical energy. With a lithium niobate photonic crystal nanocavity with an optical Q up to 1.41 million and an effective mode volume down to 0.78$(λ/n)^3$, we are able to achieve a resonance tuning rate of about 88.4~MHz/photon (0.67~GHz/aJ), which allows us to tune across the whole cavity resonance with only about 2.4 photons on average inside the cavity. Such a photon-level resonance tuning is of great potential for energy-efficient optical switching, wavelength routing, and reconfiguration of photonic devices/circuits that are indispensable for future photonic interconnect.

physics.optics↗

Probabilistic Recursive Reasoning for Multi-Agent Reinforcement Learning

Humans are capable of attributing latent mental contents such as beliefs or intentions to others. The social skill is critical in daily life for reasoning about the potential consequences of others' behaviors so as to plan ahead. It is known that humans use such reasoning ability recursively by considering what others believe about their own beliefs. In this paper, we start from level-$1$ recursion and introduce a probabilistic recursive reasoning (PR2) framework for multi-agent reinforcement learning. Our hypothesis is that it is beneficial for each agent to account for how the opponents would react to its future behaviors. Under the PR2 framework, we adopt variational Bayes methods to approximate the opponents' conditional policies, to which each agent finds the best response and then improve their own policies. We develop decentralized-training-decentralized-execution algorithms, namely PR2-Q and PR2-Actor-Critic, that are proved to converge in the self-play scenarios when there exists one Nash equilibrium. Our methods are tested on both the matrix game and the differential game, which have a non-trivial equilibrium where common gradient-based methods fail to converge. Our experiments show that it is critical to reason about how the opponents believe about what the agent believes. We expect our work to contribute a new idea of modeling the opponents to the multi-agent reinforcement learning community.

cs.LG↗

Adversarial Variational Bayes Methods for Tweedie Compound Poisson Mixed Models

The Tweedie Compound Poisson-Gamma model is routinely used for modeling non-negative continuous data with a discrete probability mass at zero. Mixed models with random effects account for the covariance structure related to the grouping hierarchy in the data. An important application of Tweedie mixed models is pricing the insurance policies, e.g. car insurance. However, the intractable likelihood function, the unknown variance function, and the hierarchical structure of mixed effects have presented considerable challenges for drawing inferences on Tweedie. In this study, we tackle the Bayesian Tweedie mixed-effects models via variational inference approaches. In particular, we empower the posterior approximation by implicit models trained in an adversarial setting. To reduce the variance of gradients, we reparameterize random effects, and integrate out one local latent variable of Tweedie. We also employ a flexible hyper prior to ensure the richness of the approximation. Our method is evaluated on both simulated and real-world data. Results show that the proposed method has smaller estimation bias on the random effects compared to traditional inference methods including MCMC; it also achieves a state-of-the-art predictive performance, meanwhile offering a richer estimation of the variance function.

stat.ML↗

Thermostat-assisted continuously-tempered Hamiltonian Monte Carlo for Bayesian learning

We propose a new sampling method, the thermostat-assisted continuously-tempered Hamiltonian Monte Carlo, for Bayesian learning on large datasets and multimodal distributions. It simulates the Nosé-Hoover dynamics of a continuously-tempered Hamiltonian system built on the distribution of interest. A significant advantage of this method is that it is not only able to efficiently draw representative i.i.d. samples when the distribution contains multiple isolated modes, but capable of adaptively neutralising the noise arising from mini-batches and maintaining accurate sampling. While the properties of this method have been studied using synthetic distributions, experiments on three real datasets also demonstrated the gain of performance over several strong baselines with various types of neural networks plunged in.

stat.ML↗

A self-starting bi-chromatic LiNbO3 soliton microcomb

For its many useful properties, including second and third-order optical nonlinearity as well as electro-optic control, lithium niobate is considered an important potential microcomb material. Here, a soliton microcomb is demonstrated in a monolithic high-Q lithium niobate resonator. Besides the demonstration of soliton mode locking, the photorefractive effect enables mode locking to self-start and soliton switching to occur bi-directionally. Second-harmonic generation of the soliton spectrum is also observed, an essential step for comb self-referencing. The Raman shock time constant of lithium niobate is also determined by measurement of soliton self-frequency-shift. Besides the considerable technical simplification provided by a self-starting soliton system, these demonstrations, together with the electro-optic and piezoelectric properties of lithium niobate, open the door to a multi-functional microcomb providing f-2f generation and fast electrical control of optical frequency and repetition rate, all of which are critical in applications including time keeping, frequency synthesis/division, spectroscopy and signal generation.

physics.optics↗

Parallel-tempered Stochastic Gradient Hamiltonian Monte Carlo for Approximate Multimodal Posterior Sampling

We propose a new sampler that integrates the protocol of parallel tempering with the Nosé-Hoover (NH) dynamics. The proposed method can efficiently draw representative samples from complex posterior distributions with multiple isolated modes in the presence of noise arising from stochastic gradient. It potentially facilitates deep Bayesian learning on large datasets where complex multimodal posteriors and mini-batch gradient are encountered.

stat.ML↗

A Neural Stochastic Volatility Model

In this paper, we show that the recent integration of statistical models with deep recurrent neural networks provides a new way of formulating volatility (the degree of variation of time series) models that have been widely used in time series analysis and prediction in finance. The model comprises a pair of complementary stochastic recurrent neural networks: the generative network models the joint distribution of the stochastic volatility process; the inference network approximates the conditional distribution of the latent variables given the observables. Our focus here is on the formulation of temporal dynamics of volatility over time under a stochastic recurrent neural network framework. Experiments on real-world stock price datasets demonstrate that the proposed model generates a better volatility estimation and prediction that outperforms mainstream methods, e.g., deterministic models such as GARCH and its variants, and stochastic models namely the MCMC-based model \emph{stochvol} as well as the Gaussian process volatility model \emph{GPVol}, on average negative log-likelihood.

cs.LG↗

Pulsar giant pulse: coherent instability near light cylinder

Giant pulses (GPs) are extremely bright individual pulses of radio pulsar. In microbursts of Crab pulsar, which is an active GP emitter, zebra-pattern-like spectral structures are observed, which are reminiscent of the `zebra bands' that are observed in type IV solar radio flares. However, band spacing linearly increases with the band center frequency of $\sim5-30$\,GHz. In this study, we propose that the Crab pulsar GP can originate from the coherent instability of plasma near a light cylinder. Further, the growth of coherent instability can be attributed to the resonance observed between the cyclotron-resonant-excited wave and the background plasma oscillation. The particles can be injected into the closed-field line regions owing to magnetic reconnection near a light cylinder. These particles introduce a large amount of free energy that further causes cyclotron-resonant instability, which grows and amplifies radiative waves at frequencies close to the electron cyclotron harmonics that exhibit zebra-pattern-like spectral band structures. Further, these structures can be modulated by the resonance between the cyclotron-resonant-excited wave and the background plasma oscillation. In this scenario, the band structures of the Crab pulsar can be well fitted by a coherent instability model, where the plasma density of a light cylinder should be $\sim10^{13-15}\,\rm{cm^{-3}}$, with an estimated gradient of $>5.5\times10^5\,\rm{cm^{-4}}$. This process may be accompanied by high-energy emissions. Similar phenomena are expected to be detected in other types of GP sources that have magnetic fields of $\simeq10^6$\,G in a light cylinder.

astro-ph.HE↗

Clumpy jets from black hole-massive star binaries as engines of Fast Radio Bursts

We propose a new model of Fast Radio Bursts (FRBs) based on stellar mass black hole-massive star binaries. We argue that the inhomogeneity of the circumstellar materials or/and the time varying wind activities of the stellar companion will cause the black hole to accrete at a transient super-Eddington rate. The collision among the clumpy ejecta in the resulted jet could trigger plasma instability. As a result, the plasma in the jet will emit coherent curvature radiation. When the jet cone aims toward the observer, the apparent luminosity can be $10^{41}-10^{42}$\,erg/s. The duration of the resulted flare is $\sim$ millisecond. The high event rate of the observed non-repeating FRBs can be explained. A similar scenario in the vicinity of a supermassive black hole can be used to explain the interval distribution of the repeating source FRB121102 qualitatively.

astro-ph.HE↗

Benchmarking Deep Sequential Models on Volatility Predictions for Financial Time Series

Volatility is a quantity of measurement for the price movements of stocks or options which indicates the uncertainty within financial markets. As an indicator of the level of risk or the degree of variation, volatility is important to analyse the financial market, and it is taken into consideration in various decision-making processes in financial activities. On the other hand, recent advancement in deep learning techniques has shown strong capabilities in modelling sequential data, such as speech and natural language. In this paper, we empirically study the applicability of the latest deep structures with respect to the volatility modelling problem, through which we aim to provide an empirical guidance for the theoretical analysis of the marriage between deep learning techniques and financial applications in the future. We examine both the traditional approaches and the deep sequential models on the task of volatility prediction, including the most recent variants of convolutional and recurrent networks, such as the dilated architecture. Accordingly, experiments with real-world stock price datasets are performed on a set of 1314 daily stock series for 2018 days of transaction. The evaluation and comparison are based on the negative log likelihood (NLL) of real-world stock price time series. The result shows that the dilated neural models, including dilated CNN and Dilated RNN, produce most accurate estimation and prediction, outperforming various widely-used deterministic models in the GARCH family and several recently proposed stochastic models. In addition, the high flexibility and rich expressive power are validated in this study.

cs.LG↗

Optical parametric generation in a lithium niobate microring with modal phase matching

The lithium niobate integrated photonic platform has recently shown great promise in nonlinear optics on a chip scale. Here, we report second-harmonic generation in a high-Q lithium niobate microring resonator through modal phase matching, with a conversion efficiency of 1,500% W$^{-1}$. Our device also allows us to observe difference-frequency generation in the telecom band. Our work demonstrates the great potential of the lithium niobate integrated platform for nonlinear wavelength conversion with high efficiencies.

physics.optics↗