Searcharxiv⌕ Search

arXiv subjects

Yu Lu

Publications and source records attributed to Yu Lu.

At least 109 records · Page 6Linked to original sources

The assignments of the $B_s$ mesons within the screened potential model and $^3P_0$ model

We investigate the mass spectrum and the decay properties of the $B_s$ mesons within the screened nonrelativistic quark model and the $^3P_0$ model. Our results suggest that the $B_{sJ}(6064)$ and $B_{sJ}(6114)$ states, as the first solution of the recently LHCb measurements, could be explained as the $B_s(1^3D_3)$ and $B_s(1^3D_1)$, respectively. In addition, the $B_{sJ}(6109)$ and $B_{sJ}(6158)$ states, as the second solution of the LHCb measurements, could be explained as the $B'_{s2}(1D)$ and $B_{s1}(2P)$, respectively. Meanwhile, the $B_{s1}(5830)$ could be interpreted as the candidate of the $B_{s1}(1P)$. We also calculated the decay properties of the other excited $B_s$ mesons with the predicted masses, which should be helpful for the experimental searching in future.

hep-ph↗

AUGUST: an Automatic Generation Understudy for Synthesizing Conversational Recommendation Datasets

High-quality data is essential for conversational recommendation systems and serves as the cornerstone of the network architecture development and training strategy design. Existing works contribute heavy human efforts to manually labeling or designing and extending recommender dialogue templates. However, they suffer from (i) the limited number of human annotators results in that datasets can hardly capture rich and large-scale cases in the real world, (ii) the limited experience and knowledge of annotators account for the uninformative corpus and inappropriate recommendations. In this paper, we propose a novel automatic dataset synthesis approach that can generate both large-scale and high-quality recommendation dialogues through a data2text generation process, where unstructured recommendation conversations are generated from structured graphs based on user-item information from the real world. In doing so, we comprehensively exploit: (i) rich personalized user profiles from traditional recommendation datasets, (ii) rich external knowledge from knowledge graphs, and (iii) the conversation ability contained in human-to-human conversational recommendation datasets. Extensive experiments validate the benefit brought by the automatically synthesized data under low-resource scenarios and demonstrate the promising potential to facilitate the development of a more effective conversational recommendation system.

cs.CL↗

Pretrained Language Model based Web Search Ranking: From Relevance to Satisfaction

Search engine plays a crucial role in satisfying users' diverse information needs. Recently, Pretrained Language Models (PLMs) based text ranking models have achieved huge success in web search. However, many state-of-the-art text ranking approaches only focus on core relevance while ignoring other dimensions that contribute to user satisfaction, e.g., document quality, recency, authority, etc. In this work, we focus on ranking user satisfaction rather than relevance in web search, and propose a PLM-based framework, namely SAT-Ranker, which comprehensively models different dimensions of user satisfaction in a unified manner. In particular, we leverage the capacities of PLMs on both textual and numerical inputs, and apply a multi-field input that modularizes each dimension of user satisfaction as an input field. Overall, SAT-Ranker is an effective, extensible, and data-centric framework that has huge potential for industrial applications. On rigorous offline and online experiments, SAT-Ranker obtains remarkable gains on various evaluation sets targeting different dimensions of user satisfaction. It is now fully deployed online to improve the usability of our search engine.

cs.IR↗

Hierarchical Beam Training for Extremely Large-Scale MIMO: From Far-Field to Near-Field

Extremely large-scale MIMO (XL-MIMO) is a promising technique for future 6G communications. The sharp increase in the number of antennas causes electromagnetic propagation to change from far-field to near-field. Due to the near-field effect, the exhaustive near-field beam training at all angles and distances requires very high overhead. The improved fast near-field beam training scheme based on time-delay structure can reduce the overhead, but it suffers from very high hardware costs and energy consumption caused by time-delay circuits. In this paper, we propose a near-field two dimension (2D) hierarchical beam training scheme to reduce the overhead without the need for extra hardware circuits. Specifically, we first formulate the multi-resolution near-field codewords design problem covering different angle and distance coverages. Next, inspired by phase retrieval problems in digital holography imaging technology, we propose a Gerchberg-Saxton (GS)-based algorithm to acquire the theoretical codeword by considering the ideal fully digital architecture. Based on the theoretical codeword, an alternating optimization algorithm is then proposed to acquire the practical codeword by considering the hybrid digital-analog architecture. Finally, with the help of multi-resolution codebooks, we propose a near-field 2D hierarchical beam training scheme to significantly reduce the training overhead, which is verified by extensive simulation results.

cs.IT↗

Near-Field Channel Estimation in Mixed LoS/NLoS Environments for Extremely Large-Scale MIMO Systems

Accurate channel model and channel estimation are essential to empower extremely large-scale MIMO (XL-MIMO) in 6G networks with ultra-high spectral efficiency. With the sharp increase in the antenna array aperture of the XL-MIMO scenario, the electromagnetic propagation field will change from far-field to near-field. Unfortunately, due to the near-field effect, most of the existing XL-MIMO channel models fail to describe mixed line-of-sight (LoS) and non-line-of-sight (NLoS) path components simultaneously. In this paper, a mixed LoS/NLoS near-field XL-MIMO channel model is proposed to match the practical near-field XL-MIMO scenario, where the LoS path component is modeled by the geometric free space propagation assumption while NLoS path components are modeled by the near-field array response vectors. Then, to define the range of near-field for XL-MIMO, the MIMO Rayleigh distance (MIMO-RD) and MIMO advanced RD (MIMO-ARD) is derived. Next, a two stage channel estimation algorithm is proposed, where the LoS path component and NLoS path components are estimated separately. Moreover, the Cramer-Rao lower bound (CRLB) of the proposed algorithm is derived in this paper. Numerical simulation results demonstrate that, the proposed two stage scheme is able to outperform the existing methods in both the theoretical channel model and the QuaDRiGa channel emulation platform.

cs.IT↗

A Simple yet Effective Framework for Active Learning to Rank

While China has become the biggest online market in the world with around 1 billion internet users, Baidu runs the world largest Chinese search engine serving more than hundreds of millions of daily active users and responding billions queries per day. To handle the diverse query requests from users at web-scale, Baidu has done tremendous efforts in understanding users' queries, retrieve relevant contents from a pool of trillions of webpages, and rank the most relevant webpages on the top of results. Among these components used in Baidu search, learning to rank (LTR) plays a critical role and we need to timely label an extremely large number of queries together with relevant webpages to train and update the online LTR models. To reduce the costs and time consumption of queries/webpages labeling, we study the problem of Activ Learning to Rank (active LTR) that selects unlabeled queries for annotation and training in this work. Specifically, we first investigate the criterion -- Ranking Entropy (RE) characterizing the entropy of relevant webpages under a query produced by a sequence of online LTR models updated by different checkpoints, using a Query-By-Committee (QBC) method. Then, we explore a new criterion namely Prediction Variances (PV) that measures the variance of prediction results for all relevant webpages under a query. Our empirical studies find that RE may favor low-frequency queries from the pool for labeling while PV prioritizing high-frequency queries more. Finally, we combine these two complementary criteria as the sample selection strategies for active learning. Extensive experiments with comparisons to baseline algorithms show that the proposed approach could train LTR models achieving higher Discounted Cumulative Gain (i.e., the relative improvement ΔDCG4=1.38%) with the same budgeted labeling efforts.

cs.IR↗

Proposal for the search for exotic spin-spin interactions at the micrometer scale using functionalized cantilever force sensors

Spin-dependent exotic interactions can be generated by exchanging hypothetical bosons, which were introduced to solve some puzzles in physics. Many precision experiments have been performed to search for such interactions, but no confirmed observation has been made. Here, we propose new experiments to search for the exotic spin-spin interactions that can be mediated by axions or Z$^\prime$ bosons. A sensitive functionalized cantilever is utilized as a force sensor to measure the interactions between the spin-polarized electrons in a periodic magnetic source structure and a closed-loop magnetic structure integrated on the cantilever. The source is set to oscillate during data acquisition to modulate the exotic force signal to high harmonics of the oscillating frequency. This helps to suppress the spurious signals at the signal frequency. Different magnetic source structures are designed for different interaction detections. A magnetic stripe structure is designed for Z$^\prime$-mediated interaction, which is insensitive to the detection of axion-mediated interaction. This allows us to measure the coupling constant of both if we assume both exist. With the force sensitivity achievable at low temperature, the proposed experiments are expected to search for the parameter spaces with much smaller coupling constant than the current stringent constraints from micrometer to millimeter range. Specifically, the lower bound of the parameter space will be seven orders of magnitude lower than the stringent constraints for Z$^\prime$-mediated interaction, and an order of magnitude lower for axion-mediated interaction, at the interaction range of $10\, μ$m.

hep-ex↗

On the critical exponent $p_c$ of the 3D quasilinear wave equation $-\big(1+(\partial_tϕ)^p\big)\partial_t^2ϕ+Δϕ=0$ with short pulse initial data. II, shock formation

In the previous paper [Ding Bingbing, Lu Yu, Yin Huicheng, On the critical exponent $p_c$ of the 3D quasilinear wave equation $-\big(1+(\partial_tϕ)^p\big)\partial_t^2ϕ+Δϕ=0$ with short pulse initial data. I, global existence, Preprint, 2022], for the 3D quasilinear wave equation $-\big(1+(\partial_tϕ)^p\big)\partial_t^2ϕ+Δϕ=0$ with short pulse initial data $(ϕ,\partial_tϕ)(1,x)=\big(δ^{2-\varepsilon_0}ϕ_0(\frac{r-1}δ,ω),δ^{1-\varepsilon_0}ϕ_1(\frac{r-1}δ,ω)\big)$, where $p\in\mathbb{N}$, $0<\varepsilon_0<1$, under the outgoing constraint condition $(\partial_t+\partial_r)^kϕ(1,x)=O(δ^{2-\varepsilon_0-k\max\{0,1-(1-\varepsilon_0)p\}})$ for $k=1,2$, the authors establish the global existence of smooth large solution $ϕ$ when $p>p_c$ with $p_c=\frac{1}{1-\varepsilon_0}$. In the present paper, under the same outgoing constraint condition, when $1\leq p\leq p_c$, we will show that the smooth solution $ϕ$ blows up and further the outgoing shock is formed in finite time.

math.AP↗

On the critical exponent $p_c$ of the 3D quasilinear wave equation $-\big(1+(\partial_tϕ)^p\big)\partial_t^2ϕ+Δϕ=0$ with short pulse initial data. I, global existence

For the 3D quasilinear wave equation $-\big(1+(\partial_tϕ)^p\big)\partial_t^2ϕ+Δϕ=0$ with the short pulse initial data $(ϕ,\partial_tϕ)(1,x)=\big(δ^{2-\varepsilon_0}ϕ_0(\frac{r-1}δ,ω), δ^{1-\varepsilon_0}ϕ_1(\frac{r-1}δ,ω)\big)$, where $p\in\mathbb N$, $p\geq 2$, $0<\varepsilon_0<1$, $r=|x|$, $ω= \frac{x}{r}\in\mathbb S^2$, and $δ>0$ is sufficiently small, under the outgoing constraint condition $(\partial_t+\partial_r)^kϕ(1,x)=O(δ^{2-\varepsilon_0})$ for $k=1,2$, we will establish the global existence of smooth large data solution $ϕ$ when $p>p_c$ with $p_c=\frac{1}{1-\varepsilon_0}$ being the critical exponent. In the forthcoming paper, when $1\leq p\leq p_c$, we show the formation of the outgoing shock before the time $t=2$ under the suitable assumptions of $(ϕ_0,ϕ_1)$.

math.AP↗

Clustering High-dimensional Data via Feature Selection

High-dimensional clustering analysis is a challenging problem in statistics and machine learning, with broad applications such as the analysis of microarray data and RNA-seq data. In this paper, we propose a new clustering procedure called Spectral Clustering with Feature Selection (SC-FS), where we first obtain an initial estimate of labels via spectral clustering, then select a small fraction of features with the largest R-squared with these labels, i.e., the proportion of variation explained by group labels, and conduct clustering again using selected features. Under mild conditions, we prove that the proposed method identifies all informative features with high probability and achieves minimax optimal clustering error rate for the sparse Gaussian mixture model. Applications of SC-FS to four real world data sets demonstrate its usefulness in clustering high-dimensional data.

stat.ME↗

The reaction $πN \to ωN$ in a dynamical coupled-channel approach

A refined investigation on light flavor meson-baryon scatterings is performed using a dynamical coupled-channel approach, the Jülich-Bonn model, that respects unitartiy and analyticity constraints. The channel space of $πN$, $πΔ$, $σN$, $ρN$, $ηN$, $K Λ$ and $K Σ$ is extended by adding the $ωN$ final state. The spectra of $N^*$ and $Δ$ resonances are extracted in terms of complex poles of the scattering amplitudes, based on the result of a global fit to a worldwide collection of data, in the energy region from the $πN$ threshold to center-of-mass energy $z=2.3$ GeV. A negative value of the $ωN$ elastic spin-averaged scattering length is extracted, questioning the existence of bound states of the $ω$ meson in the nuclear matter.

nucl-th↗

Personalizing or Not: Dynamically Personalized Federated Learning with Incentives

Personalized federated learning (FL) facilitates collaborations between multiple clients to learn personalized models without sharing private data. The mechanism mitigates the statistical heterogeneity commonly encountered in the system, i.e., non-IID data over different clients. Existing personalized algorithms generally assume all clients volunteer for personalization. However, potential participants might still be reluctant to personalize models since they might not work well. In this case, clients choose to use the global model instead. To avoid making unrealistic assumptions, we introduce the personalization rate, measured as the fraction of clients willing to train personalized models, into federated settings and propose DyPFL. This dynamically personalized FL technique incentivizes clients to participate in personalizing local models while allowing the adoption of the global model when it performs better. We show that the algorithmic pipeline in DyPFL guarantees good convergence performance, allowing it to outperform alternative personalized methods in a broad range of conditions, including variation in heterogeneity, number of clients, local epochs, and batch sizes.

cs.LG↗

Near-Field Communications for 6G: Fundamentals, Challenges, Potentials, and Future Directions

Extremely large antenna array (ELAA) is a common feature of several key candidate technologies for sixth-generation mobile networks (6G), such as ultra-massive multiple-input-multiple-output (UM-MIMO), cell-free massive MIMO, reconfigurable intelligent surface (RIS), and terahertz communications. Since the number of antennas is very large for ELAA, the electromagnetic radiation field needs to be modeled by near-field spherical waves, which is opposed to the conventional planar-wave-based radiation model of 5G massive MIMO. As a result, near-field communications will become essential in 6G wireless networks. In this article, we systematically investigate the emerging near-field communication techniques. Firstly, we present the fundamentals of near-field communications and the metric to determine the near-field ranges in typical communication scenarios. Then, we investigate recent studies specific to near-field communications by classifying them into two categories, i.e., techniques addressing the challenges and those exploiting the potentials in near-field regions. Their principles, recent progress, pros and cons are discussed. More importantly, several open problems and future research directions for near-field communications are pointed out. We believe that this article would inspire more innovations for this important research topic of near-field communications for 6G.

cs.IT↗

Coupled Channel Effects for the Charmed-Strange Mesons

We make a systematic calculation of the spectra and hadronic decays of the $D_s$ system in a coupled channel framework, where the unquenched effects are induced by the $^3P_0$ model. In the calculation, the wave functions are obtained by using a nonrelativistic potential model and are handled precisely with Gaussian Expansion Method. Even though the fitting mainly focuses on the spectrum, our model agrees well with the experiments on both the spectra and the hadronic decays, suggesting that the coupled channel effect could result in a reasonable and coherent description of the $D_s$ mesons. Based on the calculation, we give a detailed analysis on various aspects of the excited states, especially $D_s(2317), D_s(2460), D_s(2536), D_s(2860), D_s(3040)$. We also predict that $D_s(2{}^3P_0)$ should be a $D^*K^*$ dominant molecule with mass $2894$~MeV, which is only $5$~MeV below the $D^*K^*$ threshold.

hep-ph↗

Interpretable Fault Diagnosis of Rolling Element Bearings with Temporal Logic Neural Network

Machine learning-based methods have achieved successful applications in machinery fault diagnosis. However, the main limitation that exists for these methods is that they operate as a black box and are generally not interpretable. This paper proposes a novel neural network structure, called temporal logic neural network (TLNN), in which the neurons of the network are logic propositions. More importantly, the network can be described and interpreted as a weighted signal temporal logic. TLNN not only keeps the nice properties of traditional neuron networks but also provides a formal interpretation of itself with formal language. Experiments with real datasets show the proposed neural network can obtain highly accurate fault diagnosis results with good computation efficiency. Additionally, the embedded formal language of the neuron network can provide explanations about the decision process, thus achieve interpretable fault diagnosis.

cs.LG↗

Learning Confidence for Transformer-based Neural Machine Translation

Confidence estimation aims to quantify the confidence of the model prediction, providing an expectation of success. A well-calibrated confidence estimate enables accurate failure prediction and proper risk measurement when given noisy samples and out-of-distribution data in real-world settings. However, this task remains a severe challenge for neural machine translation (NMT), where probabilities from softmax distribution fail to describe when the model is probably mistaken. To address this problem, we propose an unsupervised confidence estimate learning jointly with the training of the NMT model. We explain confidence as how many hints the NMT model needs to make a correct prediction, and more hints indicate low confidence. Specifically, the NMT model is given the option to ask for hints to improve translation accuracy at the cost of some slight penalty. Then, we approximate their level of confidence by counting the number of hints the model uses. We demonstrate that our learned confidence estimate achieves high accuracy on extensive sentence/word-level quality estimation tasks. Analytical results verify that our confidence estimate can correctly assess underlying risk in two real-world scenarios: (1) discovering noisy samples and (2) detecting out-of-domain data. We further propose a novel confidence-based instance-specific label smoothing approach based on our learned confidence estimate, which outperforms standard label smoothing.

cs.CL↗

CRIS: CLIP-Driven Referring Image Segmentation

Referring image segmentation aims to segment a referent via a natural linguistic expression.Due to the distinct data properties between text and image, it is challenging for a network to well align text and pixel-level features. Existing approaches use pretrained models to facilitate learning, yet separately transfer the language/vision knowledge from pretrained models, ignoring the multi-modal corresponding information. Inspired by the recent advance in Contrastive Language-Image Pretraining (CLIP), in this paper, we propose an end-to-end CLIP-Driven Referring Image Segmentation framework (CRIS). To transfer the multi-modal knowledge effectively, CRIS resorts to vision-language decoding and contrastive learning for achieving the text-to-pixel alignment. More specifically, we design a vision-language decoder to propagate fine-grained semantic information from textual representations to each pixel-level activation, which promotes consistency between the two modalities. In addition, we present text-to-pixel contrastive learning to explicitly enforce the text feature similar to the related pixel-level features and dissimilar to the irrelevances. The experimental results on three benchmark datasets demonstrate that our proposed framework significantly outperforms the state-of-the-art performance without any post-processing. The code will be released.

cs.CV↗

A Double-Spring Model for Nanoparticle Diffusion in a Polymer Network

The transport of nanoparticles (NPs) in polymer networks, as a typical simplified model describing various structures in living systems, is profoundly important in biomedical engineering and nanotechnology. Predicting the effective diffusivity of NP confined in an ordered network has been an intriguing focus in this frontier field. In the present study, the diffusion of NPs in an unentangled polymer network for different NP radii and network stiffness is numerically investigated by single particle dissipative particle dynamics (DPD). It is found that, the deformation due to the junction deviation contributes significantly to the the potential barrier $U$ for the NP to overcome during hopping, and it is dominated over the strain energy induced by loop stretching for larger NPs and lower network rigidity. Analyses based on the theory of continuum mechanics reveal that the relation between this deformed energy and the junction deviation can be described by a non-linear spring. Taking into account both effects of the loop stretching and junction deviation, a double-spring model is proposed to characterize the diffusivity of the NPs in the ordered network. The theoretical prediction is in good agreement with our numerical simulations, and qualitatively consistent with the investigations available. This model is helpful to improve our understanding on the dynamic behavior of nanoparticle in complex biological environment, and provide theoretical guidance in designing biomedical applications.

physics.chem-ph↗