SearcharxivSearch

arXiv subjects

Dayong Wang

Publications and source records attributed to Dayong Wang.

At least 19 recordsLinked to original sources

Probing Quantum Entanglement in $\tau^+\tau^-$ Pairs via the $\pi\pi$ Channel at STCF

Quantum entanglement and Bell-inequality violation in $\tau^+\tau^-$ pairs provide a sensitive probe of quantum correlations in high-energy interactions. We present a feasibility study of $e^+e^- \to \tau^+\tau^-$ at the proposed Super Tau-Charm Facility based on full Monte Carlo simulation at $\sqrt{s} = 7$ GeV, focusing on the $\pi\pi$ channel ($\tau^\pm \to \pi^\pm\nu$), which offers the maximal spin-analyzing power $|\kappa| = 1$ and the simplest final-state topology for validating the quantum-tomography framework. We establish a consistency chain from the tree-level QED prediction through truth-level and detector-level reconstruction, yielding a reconstructed concurrence of $0.279 \pm 0.007$ with the good-solution approach. A complementary full-simulation study of the $\rho\rho$ channel is also briefly reported. These results demonstrate that the STCF can provide a competitive platform for precision studies of quantum correlations in $\tau$-lepton pairs.

hep-ex

A hierarchy tree data structure for behavior-based user segment representation

User attributes are essential in multiple stages of modern recommendation systems and are particularly important for mitigating the cold-start problem and improving the experience of new or infrequent users. We propose Behavior-based User Segmentation (BUS), a novel tree-based data structure that hierarchically segments the user universe with various users' categorical attributes based on the users' product-specific engagement behaviors. During the BUS tree construction, we use Normalized Discounted Cumulative Gain (NDCG) as the objective function to maximize the behavioral representativeness of marginal users relative to active users in the same segment. The constructed BUS tree undergoes further processing and aggregation across the leaf nodes and internal nodes, allowing the generation of popular social content and behavioral patterns for each node in the tree. To further mitigate bias and improve fairness, we use the social graph to derive the user's connection-based BUS segments, enabling the combination of behavioral patterns extracted from both the user's own segment and connection-based segments as the connection aware BUS-based recommendation. Our offline analysis shows that the BUS-based retrieval significantly outperforms traditional user cohort-based aggregation on ranking quality. We have successfully deployed our data structure and machine learning algorithm and tested it with various production traffic serving billions of users daily, achieving statistically significant improvements in the online product metrics, including music ranking and email notifications. To the best of our knowledge, our study represents the first list-wise learning-to-rank framework for tree-based recommendation that effectively integrates diverse user categorical attributes while preserving real-world semantic interpretability at a large industrial scale.

cs.LG

Studies of the tracking and identification efficiencies of electrons and positrons at BESIII

The efficiencies for electron and positron tracking and identification in the BESIII experiment are investigated with the radiative Bhabha process $e^+e^-\rightarrow e^+e^-\gamma$ from the data samples collected at the center-of-mass energies of 3.08 GeV and 3.097 GeV. The relative differences between data and MC associated with tracking and identification efficiencies of electrons and positrons, as well as the corresponding correction factors are determined. It turns out the relative differences of tracking efficiency and particle identification efficiency after correction are mostly less than 0.5$\%$ for transverse momenta $p_T>0.4$ GeV and for the entire momentum region, respectively.

hep-ex

Flavor Physics at the CEPC: a General Perspective

We discuss the landscape of flavor physics at the Circular Electron-Positron Collider (CEPC), based on the nominal luminosity outlined in its Technical Design Report. The CEPC is designed to operate in multiple modes to address a variety of tasks. At the $Z$ pole, the expected production of 4 Tera $Z$ bosons will provide unique and highly precise measurements of $Z$ boson couplings, while the substantial number of boosted heavy-flavored quarks and leptons produced in clean $Z$ decays will facilitate investigations into their flavor physics with unprecedented precision. We investigate the prospects of measuring various physics benchmarks and discuss their implications for particle theories and phenomenological models. Our studies indicate that, with its highlighted advantages and anticipated excellent detector performance, the CEPC can explore beauty and $τ$ physics in ways that are superior to or complementary with the Belle II and Large-Hadron-Collider-beauty experiments, potentially enabling the detection of new physics at energy scales of 10 TeV and above. This potential also extends to the observation of yet-to-be-discovered rare and exotic processes, as well as testing fundamental principles such as lepton flavor universality, lepton and baryon number conservation, etc., making the CEPC a vibrant platform for flavor physics research. The $WW$ threshold scan, Higgs-factory operation and top-pair productions of the CEPC further enhance its merits in this regard, especially for measuring the Cabibbo-Kobayashi-Maskawa matrix elements, and Flavor-Changing-Neutral-Current physics of Higgs boson and top quarks. We outline the requirements for detector performance and considerations for future development to achieve the anticipated scientific goals.

hep-ex

Probing darK Matter Using free leptONs: PKMUON

We propose a new method to detect sub-GeV dark matter, through their scatterings from free leptons and the resulting kinematic shifts. Specially, such an experiment can detect dark matter interacting solely with muons. The experiment proposed here is to directly probe muon-philic dark matter, in a model-independent way. Its complementarity with the muon on target proposal, is similar to, e.g. XENON/PandaX and ATLAS/CMS on dark matter searches. Moreover, our proposal can work better for relatively heavy dark matter such as in the sub-GeV region. We start with a small device of a size around 0.1 to 1 meter, using atmospheric muons to set up a prototype. Within only one year of operation, the sensitivity on cross section of dark matter scattering with muons can already reach $σ_D\sim 10^{-19 (-20,\,-18)}\rm{cm}^{2}$ for a dark mater $\rm{M_D}=100\, (10,\,1000)$ MeV. We can then interface the device with a high intensity muon beam of $10^{12}$/bunch. Within one year, the sensitivity can reach $σ_D\sim 10^{-27 (-28,\,-26)}\rm{cm}^{2}$ for $\rm{M_D}=100\, (10,\,1000)$ MeV.

hep-ph

Designing a Deep Learning-Driven Resource-Efficient Diagnostic System for Metastatic Breast Cancer: Reducing Long Delays of Clinical Diagnosis and Improving Patient Survival in Developing Countries

Breast cancer is one of the leading causes of cancer mortality. Breast cancer patients in developing countries, especially sub-Saharan Africa, South Asia, and South America, suffer from the highest mortality rate in the world. One crucial factor contributing to the global disparity in mortality rate is long delay of diagnosis due to a severe shortage of trained pathologists, which consequently has led to a large proportion of late-stage presentation at diagnosis. The delay between the initial development of symptoms and the receipt of a diagnosis could stretch upwards 15 months. To tackle this critical healthcare disparity, this research has developed a deep learning-based diagnosis system for metastatic breast cancer that can achieve high diagnostic accuracy as well as computational efficiency. Based on our evaluation, the MobileNetV2-based diagnostic model outperformed the more complex VGG16, ResNet50 and ResNet101 models in diagnostic accuracy, model generalization, and model training efficiency. The visual comparisons between the model prediction and ground truth have demonstrated that the MobileNetV2 diagnostic models can identify very small cancerous nodes embedded in a large area of normal cells which is challenging for manual image analysis. Equally Important, the light weighted MobleNetV2 models were computationally efficient and ready for mobile devices or devices of low computational power. These advances empower the development of a resource-efficient and high performing AI-based metastatic breast cancer diagnostic system that can adapt to under-resourced healthcare facilities in developing countries. This research provides an innovative technological solution to address the long delays in metastatic breast cancer diagnosis and the consequent disparity in patient survival outcome in developing countries.

eess.IV

New Physics Searches at Kaon and Hyperon Factories

Rare meson decays are among the most sensitive probes of both heavy and light new physics. Among them, new physics searches using kaons benefit from their small total decay widths and the availability of very large datasets. On the other hand, useful complementary information is provided by hyperon decay measurements. We summarize the relevant phenomenological models and the status of the searches in a comprehensive list of kaon and hyperon decay channels. We identify new search strategies for under-explored signatures, and demonstrate that the improved sensitivities from current and next-generation experiments could lead to a qualitative leap in the exploration of light dark sectors.

hep-ph

Predicting breast tumor proliferation from whole-slide images: the TUPAC16 challenge

Tumor proliferation is an important biomarker indicative of the prognosis of breast cancer patients. Assessment of tumor proliferation in a clinical setting is highly subjective and labor-intensive task. Previous efforts to automate tumor proliferation assessment by image analysis only focused on mitosis detection in predefined tumor regions. However, in a real-world scenario, automatic mitosis detection should be performed in whole-slide images (WSIs) and an automatic method should be able to produce a tumor proliferation score given a WSI as input. To address this, we organized the TUmor Proliferation Assessment Challenge 2016 (TUPAC16) on prediction of tumor proliferation scores from WSIs. The challenge dataset consisted of 500 training and 321 testing breast cancer histopathology WSIs. In order to ensure fair and independent evaluation, only the ground truth for the training dataset was provided to the challenge participants. The first task of the challenge was to predict mitotic scores, i.e., to reproduce the manual method of assessing tumor proliferation by a pathologist. The second task was to predict the gene expression based PAM50 proliferation scores from the WSI. The best performing automatic method for the first task achieved a quadratic-weighted Cohen's kappa score of $κ$ = 0.567, 95% CI [0.464, 0.671] between the predicted scores and the ground truth. For the second task, the predictions of the top method had a Spearman's correlation coefficient of r = 0.617, 95% CI [0.581 0.651] with the ground truth. This was the first study that investigated tumor proliferation assessment from WSIs. The achieved results are promising given the difficulty of the tasks and weakly-labelled nature of the ground truth. However, further research is needed to improve the practical utility of image analysis methods for this task.

cs.CV

Precision Higgs Physics at CEPC

The discovery of the Higgs boson with its mass around 125 GeV by the ATLAS and CMS Collaborations marked the beginning of a new era in high energy physics. The Higgs boson will be the subject of extensive studies of the ongoing LHC program. At the same time, lepton collider based Higgs factories have been proposed as a possible next step beyond the LHC, with its main goal to precisely measure the properties of the Higgs boson and probe potential new physics associated with the Higgs boson. The Circular Electron Positron Collider~(CEPC) is one of such proposed Higgs factories. The CEPC is an $e^+e^-$ circular collider proposed by and to be hosted in China. Located in a tunnel of approximately 100~km in circumference, it will operate at a center-of-mass energy of 240~GeV as the Higgs factory. In this paper, we present the first estimates on the precision of the Higgs boson property measurements achievable at the CEPC and discuss implications of these measurements.

hep-ex

Searching for flavor changing neutral currents at BESIII

The Flavor Changing Neutral Current decays(FCNC) is forbidden at tree level in the Standard Model (SM) and could only contribute through loops. Any direct observation beyond SM expectations could be a good probe of physics beyond SM. BESIII is a currently running tau-charm factory with the largest samples of on threshold charm meson pairs, directly produced charmonia and some other unique datasets at BEPCII collider. It has great potential to probe these FCNC decays from multiple channels. Here we review some of latest search results of FCNC decays from BESIII. We present the latest results of searches for the decay of $J/ψ\to D^{0} e^+ e^-$, $ψ(3686) \to D^{0} e^+ e^-$, $ψ(3686) \to Λ_c^+ \overline{p} e^+ e^-$, $D \to h (h') e^+ e^-$, $D^{+} \to h^{+} e^+ e^-$, etc. Other related searches, together with the prospects and challenges with other channels and the future data are also discussed.

hep-ex

Angular analyses of $b \to s μ^+ μ^-$ transitions at CMS

The flavour changing neutral current decays can be interesting probes for searching for new physics. Angular distributions of $b \to s \ell^+ \ell^-$ transition processes of both $\mathrm{B}^0 \to \mathrm{K}^{*0} μ^ +μ^-$ and $\mathrm{B}^+ \to \mathrm{K}^+ μ^+μ^-$ are studied using a sample of proton-proton collisions at $\sqrt{s} = 8~\mathrm{TeV}$ collected with the CMS detector at the LHC, corresponding to an integrated luminosity of $20.5~\mathrm{fb}^{-1}$. Angular analyses are performed to determine $P_1$ and $P_5'$ angular parameters for $\mathrm{B}^0 \to \mathrm{K}^{*0} μ^ +μ^-$ and $A_{FB}$ and $F_{H}$ parameters for $\mathrm{B}^+ \to \mathrm{K}^+ μ^+μ^-$, all as functions of the dimuon invariant mass squared. The $P_5'$ parameter is of particular interest due to recent measurements that indicate a potential discrepancy with the standard model. All the measurements are consistent with the standard model predictions. Efforts with more channels and more coming data will be continued to further test the standard model in higher precision in future.

hep-ex

Cross Section and Higgs Mass Measurement with Higgsstrahlung at the CEPC

The Circular Electron Positron Collider (CEPC) is a future Higgs factory proposed by the Chinese high energy physics community. It will operate at a center-of-mass energy of 240-250 GeV. The CEPC will accumulate an integrated luminosity of 5 ab$^{\rm{-1}}$ in ten years' operation, producing one million Higgs bosons via the Higgsstrahlung and vector boson fusion processes. This sample allows a percent or even sub-percent level determination of the Higgs boson couplings. With GEANT4-based full simulation and dedicated fast simulation tool, we evaluated the statistical precisions of the Higgstrahlung cross section $σ_{ZH}$ and the Higgs mass $m_{H}$ measurement at the CEPC in the $Z\rightarrowμ^+μ^-$ channel. The statistical precision of $σ_{ZH}$ ($m_{H}$) measurement could reach 0.97\% (6.9 MeV) in the model-independent analysis which uses only the information of Z boson decay. For the standard model Higgs boson, the $m_{H}$ precision could be improved to 5.4 MeV by including the information of Higgs decays. Impact of the TPC size to these measurements is investigated. In addition, we studied the prospect of measuring the Higgs boson decaying into invisible final states at the CEPC. With the standard model $ZH$ production rate, the upper limit of ${\cal B}(H\rightarrow \rm{inv.})$ could reach 1.2\% at 95\% confidence level.

hep-ex

Deep Learning Assessment of Tumor Proliferation in Breast Cancer Histological Images

Current analysis of tumor proliferation, the most salient prognostic biomarker for invasive breast cancer, is limited to subjective mitosis counting by pathologists in localized regions of tissue images. This study presents the first data-driven integrative approach to characterize the severity of tumor growth and spread on a categorical and molecular level, utilizing multiple biologically salient deep learning classifiers to develop a comprehensive prognostic model. Our approach achieves pathologist-level performance on three-class categorical tumor severity prediction. It additionally pioneers prediction of molecular expression data from a tissue image, obtaining a Spearman's rank correlation coefficient of 0.60 with ex vivo mean calculated RNA expression. Furthermore, our framework is applied to identify over two hundred unprecedented biomarkers critical to the accurate assessment of tumor proliferation, validating our proposed integrative pipeline as the first to holistically and objectively analyze histopathological images.

cs.CV

Deep Learning for Identifying Metastatic Breast Cancer

The International Symposium on Biomedical Imaging (ISBI) held a grand challenge to evaluate computational systems for the automated detection of metastatic breast cancer in whole slide images of sentinel lymph node biopsies. Our team won both competitions in the grand challenge, obtaining an area under the receiver operating curve (AUC) of 0.925 for the task of whole slide image classification and a score of 0.7051 for the tumor localization task. A pathologist independently reviewed the same images, obtaining a whole slide image classification AUC of 0.966 and a tumor localization score of 0.733. Combining our deep learning system's predictions with the human pathologist's diagnoses increased the pathologist's AUC to 0.995, representing an approximately 85 percent reduction in human error rate. These results demonstrate the power of using deep learning to produce significant improvements in the accuracy of pathological diagnoses.

q-bio.QM

Clustering Millions of Faces by Identity

In this work, we attempt to address the following problem: Given a large number of unlabeled face images, cluster them into the individual identities present in this data. We consider this a relevant problem in different application scenarios ranging from social media to law enforcement. In large-scale scenarios the number of faces in the collection can be of the order of hundreds of million, while the number of clusters can range from a few thousand to millions--leading to difficulties in terms of both run-time complexity and evaluating clustering and per-cluster quality. An efficient and effective Rank-Order clustering algorithm is developed to achieve the desired scalability, and better clustering accuracy than other well-known algorithms such as k-means and spectral clustering. We cluster up to 123 million face images into over 10 million clusters, and analyze the results in terms of both external cluster quality measures (known face labels) and internal cluster quality measures (unknown face labels) and run-time. Our algorithm achieves an F-measure of 0.87 on a benchmark unconstrained face dataset (LFW, consisting of 13K faces), and 0.27 on the largest dataset considered (13K images in LFW, plus 123M distractor images). Additionally, we present preliminary work on video frame clustering (achieving 0.71 F-measure when clustering all frames in the benchmark YouTube Faces dataset). A per-cluster quality measure is developed which can be used to rank individual clusters and to automatically identify a subset of good quality clusters for manual exploration.

cs.CV

Face Search at Scale: 80 Million Gallery

Due to the prevalence of social media websites, one challenge facing computer vision researchers is to devise methods to process and search for persons of interest among the billions of shared photos on these websites. Facebook revealed in a 2013 white paper that its users have uploaded more than 250 billion photos, and are uploading 350 million new photos each day. Due to this humongous amount of data, large-scale face search for mining web images is both important and challenging. Despite significant progress in face recognition, searching a large collection of unconstrained face images has not been adequately addressed. To address this challenge, we propose a face search system which combines a fast search procedure, coupled with a state-of-the-art commercial off the shelf (COTS) matcher, in a cascaded framework. Given a probe face, we first filter the large gallery of photos to find the top-k most similar faces using deep features generated from a convolutional neural network. The k candidates are re-ranked by combining similarities from deep features and the COTS matcher. We evaluate the proposed face search system on a gallery containing 80 million web-downloaded face images. Experimental results demonstrate that the deep features are competitive with state-of-the-art methods on unconstrained face recognition benchmarks (LFW and IJB-A). Further, the proposed face search system offers an excellent trade-off between accuracy and scalability on datasets consisting of millions of images. Additionally, in an experiment involving searching for face images of the Tsarnaev brothers, convicted of the Boston Marathon bombing, the proposed face search system could find the younger brother's (Dzhokhar Tsarnaev) photo at rank 1 in 1 second on a 5M gallery and at rank 8 in 7 seconds on an 80M gallery.

cs.CV

A Framework of Sparse Online Learning and Its Applications

The amount of data in our society has been exploding in the era of big data today. In this paper, we address several open challenges of big data stream classification, including high volume, high velocity, high dimensionality, high sparsity, and high class-imbalance. Many existing studies in data mining literature solve data stream classification tasks in a batch learning setting, which suffers from poor efficiency and scalability when dealing with big data. To overcome the limitations, this paper investigates an online learning framework for big data stream classification tasks. Unlike some existing online data stream classification techniques that are often based on first-order online learning, we propose a framework of Sparse Online Classification (SOC) for data stream classification, which includes some state-of-the-art first-order sparse online learning algorithms as special cases and allows us to derive a new effective second-order online learning algorithm for data stream classification. In addition, we also propose a new cost-sensitive sparse online learning algorithm by extending the framework with application to tackle online anomaly detection tasks where class distribution of data could be very imbalanced. We also analyze the theoretical bounds of the proposed method, and finally conduct an extensive set of experiments, in which encouraging results validate the efficacy of the proposed algorithms in comparison to a family of state-of-the-art techniques on a variety of data stream classification tasks.

cs.LG

Terahertz in-line digital holography of dragonfly hindwing: amplitude and phase reconstruction at enhanced resolution by extrapolation

We report here on terahertz (THz) digital holography on a biological specimen. A continuous-wave (CW) THz in-line holographic setup was built based on a 2.52 THz CO2 pumped THz laser and a pyroelectric array detector. We introduced novel statistical method of obtaining true intensity values for the pyroelectric array detector's pixels. Absorption and phase-shifting images of a dragonfly's hind wing were reconstructed simultaneously from single in-line hologram. Furthermore, we applied phase retrieval routines to eliminate twin image and enhanced the resolution of the reconstructions by hologram extrapolation beyond the detector area. The finest observed features are 35 μm width cross veins.

physics.optics