SearcharxivSearch

arXiv subjects

Lin Yang

Publications and source records attributed to Lin Yang.

At least 199 records · Page 11Linked to original sources

Generalizing Nucleus Recognition Model in Multi-source Images via Pruning

Ki67 is a significant biomarker in the diagnosis and prognosis of cancer, whose index can be evaluated by quantifying its expression in Ki67 immunohistochemistry (IHC) stained images. However, quantitative analysis on multi-source Ki67 images is yet a challenging task in practice due to cross-domain distribution differences, which result from imaging variation, staining styles, and lesion types. Many recent studies have made some efforts on domain generalization (DG), whereas there are still some noteworthy limitations. Specifically in the case of Ki67 images, learning invariant representation is at the mercy of the insufficient number of domains and the cell categories mismatching in different domains. In this paper, we propose a novel method to improve DG by searching the domain-agnostic subnetwork in a domain merging scenario. Partial model parameters are iteratively pruned according to the domain gap, which is caused by the data converting from a single domain into merged domains during training. In addition, the model is optimized by fine-tuning on merged domains to eliminate the interference of class mismatching among various domains. Furthermore, an appropriate implementation is attained by applying the pruning method to different parts of the framework. Compared with known DG methods, our method yields excellent performance in multiclass nucleus recognition of Ki67 IHC images, especially in the lost category cases. Moreover, our competitive results are also evaluated on the public dataset over the state-of-the-art DG methods.

cs.CV

A comparative study of neural network techniques for automatic software vulnerability detection

Software vulnerabilities are usually caused by design flaws or implementation errors, which could be exploited to cause damage to the security of the system. At present, the most commonly used method for detecting software vulnerabilities is static analysis. Most of the related technologies work based on rules or code similarity (source code level) and rely on manually defined vulnerability features. However, these rules and vulnerability features are difficult to be defined and designed accurately, which makes static analysis face many challenges in practical applications. To alleviate this problem, some researchers have proposed to use neural networks that have the ability of automatic feature extraction to improve the intelligence of detection. However, there are many types of neural networks, and different data preprocessing methods will have a significant impact on model performance. It is a great challenge for engineers and researchers to choose a proper neural network and data preprocessing method for a given problem. To solve this problem, we have conducted extensive experiments to test the performance of the two most typical neural networks (i.e., Bi-LSTM and RVFL) with the two most classical data preprocessing methods (i.e., the vector representation and the program symbolization methods) on software vulnerability detection problems and obtained a series of interesting research conclusions, which can provide valuable guidelines for researchers and engineers. Specifically, we found that 1) the training speed of RVFL is always faster than BiLSTM, but the prediction accuracy of Bi-LSTM model is higher than RVFL; 2) using doc2vec for vector representation can make the model have faster training speed and generalization ability than using word2vec; and 3) multi-level symbolization is helpful to improve the precision of neural network models.

cs.SE

TransfoRNN: Capturing the Sequential Information in Self-Attention Representations for Language Modeling

In this paper, we describe the use of recurrent neural networks to capture sequential information from the self-attention representations to improve the Transformers. Although self-attention mechanism provides a means to exploit long context, the sequential information, i.e. the arrangement of tokens, is not explicitly captured. We propose to cascade the recurrent neural networks to the Transformers, which referred to as the TransfoRNN model, to capture the sequential information. We found that the TransfoRNN models which consists of only shallow Transformers stack is suffice to give comparable, if not better, performance than a deeper Transformer model. Evaluated on the Penn Treebank and WikiText-2 corpora, the proposed TransfoRNN model has shown lower model perplexities with fewer number of model parameters. On the Penn Treebank corpus, the model perplexities were reduced up to 5.5% with the model size reduced up to 10.5%. On the WikiText-2 corpus, the model perplexity was reduced up to 2.2% with a 27.7% smaller model. Also, the TransfoRNN model was applied on the LibriSpeech speech recognition task and has shown comparable results with the Transformer models.

cs.CL

Hydrophobic interaction determines docking affinity of SARS CoV 2 variants with antibodies

Preliminary epidemiologic, phylogenetic and clinical findings suggest that several novel severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) variants have increased transmissibility and decreased efficacy of several existing vaccines. Four mutations in the receptor-binding domain (RBD) of the spike protein that are reported to contribute to increased transmission. Understanding physical mechanism responsible for the affinity enhancement between the SARS-CoV-2 variants and ACE2 is the "urgent challenge" for developing blockers, vaccines and therapeutic antibodies against the coronavirus disease 2019 (COVID-19) pandemic. Based on a hydrophobic-interaction-based protein docking mechanism, this study reveals that the mutation N501Y obviously increased the hydrophobic attraction and decrease hydrophilic repulsion between the RBD and ACE2 that most likely caused the transmissibility increment of the variants. By analyzing the mutation-induced hydrophobic surface changes in the attraction and repulsion at the binding site of the complexes of the SARS-CoV-2 variants and antibodies, we found out that all the mutations of N501Y, E484K, K417N and L452R can selectively decrease or increase their binding affinity with some antibodies.

q-bio.BM

The DKU-Duke-Lenovo System Description for the Third DIHARD Speech Diarization Challenge

In this paper, we present the submitted system for the third DIHARD Speech Diarization Challenge from the DKU-Duke-Lenovo team. Our system consists of several modules: voice activity detection (VAD), segmentation, speaker embedding extraction, attentive similarity scoring, agglomerative hierarchical clustering. In addition, the target speaker VAD (TSVAD) is used for the phone call data to further improve the performance. Our final submitted system achieves a DER of 15.43% for the core evaluation set and 13.39% for the full evaluation set on task 1, and we also get a DER of 21.63% for core evaluation set and 18.90% for full evaluation set on task 2.

eess.AS

A Thermodynamics Model for Mechanochemical Synthesis of Gold Nanoparticles: Implications for Solvent-free Nanoparticle Production

Mechanochemistry is becoming an established method for the sustainable, solid-phase synthesis of scores of nano-materials and molecules, ranging from active pharmaceutical ingredients to materials for cleantech. Yet we are still lacking a good model to rationalize experimental observations and develop a mechanistic understanding of the factors at play during mechanically assisted, solid-phase nanoparticle synthesis. We propose herein a structural-phase-field-crystal (XPFC) model with a ballistic driving force to describe such a process, with the specific example of the growth of gold nanoparticles in a two component mixture. The reaction path is described in the context of free energy landscape of the model, and dynamical simulations are performed based on phenomenological model parameters closely corresponding to the experimental conditions, so as to draw conclusions on nanoparticle growth dynamics. It is shown that the ballistic term lowers the activation energy barrier of reaction, enabling the reaction in a temperature regime compatible with experimental observations. The model also explains the mechanism of precipitated grain size reduction that is consistent with experimental observations. Our simulation results afford novel mechanistic insights into mechanosynthesis with implications for nanaparticle production and beyond.

physics.chem-ph

Correction to the photometric magnitudes of the Gaia Early Data Release 3

In this letter, we have carried out an independent validation of the Gaia EDR3 photometry using about 10,000 Landolt standard stars from Clem & Landolt (2013). Using a machine learning technique, the UBVRI magnitudes are converted into the Gaia magnitudes and colors and then compared to those in the EDR3, with the effect of metallicity incorporated. Our result confirms the significant improvements in the calibration process of the Gaia EDR3. Yet modest trends up to 10 mmag with G magnitude are found for all the magnitudes and colors for the 10 < G < 19 mag range, particularly for the bright and faint ends. With the aid of synthetic magnitudes computed on the CALSPEC spectra with the Gaia EDR3 passbands, absolute corrections are further obtained, paving the way for optimal usage of the Gaia EDR3 photometry in high accuracy investigations.

astro-ph.SR

Comparison of ablators for the polar direct drive exploding pusher platform

We examine the performance of pure boron, boron carbide, high density carbon, and boron nitride ablators in the polar direct drive exploding pusher (PDXP) platform. The platform uses the polar direct drive configuration at the National Ignition Facility to drive high ion temperatures in a room temperature capsule and has potential applications for plasma physics studies and as a neutron source. The higher tensile strength of these materials compared to plastic enables a thinner ablator to support higher gas pressures, which could help optimize its performance for plasma physics experiments, while ablators containing boron enable the possiblity of collecting addtional data to constrain models of the platform. Applying recently developed and experimentally validated equation of state models for the boron materials, we examine the performance of these materials as ablators in 2D simulations, with particular focus on changes to the ablator and gas areal density, as well as the predicted symmetry of the inherently 2D implosion.

physics.comp-ph

Unlabeled Data Guided Semi-supervised Histopathology Image Segmentation

Automatic histopathology image segmentation is crucial to disease analysis. Limited available labeled data hinders the generalizability of trained models under the fully supervised setting. Semi-supervised learning (SSL) based on generative methods has been proven to be effective in utilizing diverse image characteristics. However, it has not been well explored what kinds of generated images would be more useful for model training and how to use such images. In this paper, we propose a new data guided generative method for histopathology image segmentation by leveraging the unlabeled data distributions. First, we design an image generation module. Image content and style are disentangled and embedded in a clustering-friendly space to utilize their distributions. New images are synthesized by sampling and cross-combining contents and styles. Second, we devise an effective data selection policy for judiciously sampling the generated images: (1) to make the generated training set better cover the dataset, the clusters that are underrepresented in the original training set are covered more; (2) to make the training process more effective, we identify and oversample the images of "hard cases" in the data for which annotated training data may be scarce. Our method is evaluated on glands and nuclei datasets. We show that under both the inductive and transductive settings, our SSL method consistently boosts the performance of common segmentation models and attains state-of-the-art results.

cs.CV

Lesion Harvester: Iteratively Mining Unlabeled Lesions and Hard-Negative Examples at Scale

Acquiring large-scale medical image data, necessary for training machine learning algorithms, is frequently intractable, due to prohibitive expert-driven annotation costs. Recent datasets extracted from hospital archives, e.g., DeepLesion, have begun to address this problem. However, these are often incompletely or noisily labeled, e.g., DeepLesion leaves over 50% of its lesions unlabeled. Thus, effective methods to harvest missing annotations are critical for continued progress in medical image analysis. This is the goal of our work, where we develop a powerful system to harvest missing lesions from the DeepLesion dataset at high precision. Accepting the need for some degree of expert labor to achieve high fidelity, we exploit a small fully-labeled subset of medical image volumes and use it to intelligently mine annotations from the remainder. To do this, we chain together a highly sensitive lesion proposal generator and a very selective lesion proposal classifier. While our framework is generic, we optimize our performance by proposing a 3D contextual lesion proposal generator and by using a multi-view multi-scale lesion proposal classifier. These produce harvested and hard-negative proposals, which we then re-use to finetune our proposal generator by using a novel hard negative suppression loss, continuing this process until no extra lesions are found. Extensive experimental analysis demonstrates that our method can harvest an additional 9,805 lesions while keeping precision above 90%. To demonstrate the benefits of our approach, we show that lesion detectors trained on our harvested lesions can significantly outperform the same variants only trained on the original annotations, with boost of average precision of 7% to 10%. We open source our annotations at https://github.com/JimmyCai91/DeepLesionAnnotation.

cs.CV

Exploring Voice Conversion based Data Augmentation in Text-Dependent Speaker Verification

In this paper, we focus on improving the performance of the text-dependent speaker verification system in the scenario of limited training data. The speaker verification system deep learning based text-dependent generally needs a large scale text-dependent training data set which could be labor and cost expensive, especially for customized new wake-up words. In recent studies, voice conversion systems that can generate high quality synthesized speech of seen and unseen speakers have been proposed. Inspired by those works, we adopt two different voice conversion methods as well as the very simple re-sampling approach to generate new text-dependent speech samples for data augmentation purposes. Experimental results show that the proposed method significantly improves the Equal Error Rare performance from 6.51% to 4.51% in the scenario of limited training data.

cs.SD

SuperOCR: A Conversion from Optical Character Recognition to Image Captioning

Optical Character Recognition (OCR) has many real world applications. The existing methods normally detect where the characters are, and then recognize the character for each detected location. Thus the accuracy of characters recognition is impacted by the performance of characters detection. In this paper, we propose a method for recognizing characters without detecting the location of each character. This is done by converting the OCR task into an image captioning task. One advantage of the proposed method is that the labeled bounding boxes for the characters are not needed during training. The experimental results show the proposed method outperforms the existing methods on both the license plate recognition and the watermeter character recognition tasks. The proposed method is also deployed into a low-power (300mW) CNN accelerator chip connected to a Raspberry Pi 3 for on-device applications.

cs.CV

Wideband acoustic modulation using periodic poroelastic composite structures

We proposed an effective acoustic abatement solution comprised of periodic resonators and multi-panel structures with porous lining, which incorporates the wideband capability of porous materials and the low-frequency advantage of locally resonant structures together. Theoretical model and numerical implementation are developed and validated. Two-dimensional poroelastic field expressions are used and resonator forces are incorporated. The results agree well with those reported in the literature and the results obtained from the finite element method. It turns out that sound insulation concerning both the amplitude and tuning bandwidth can be achieved effectively using porous additions and locally-resonant designs. This study presents a promising and practical alternative for wideband acoustic modulation.

physics.app-ph

Ferrimagnetic 120$^\circ$ magnetic structure in Cu2OSO4

We report magnetic properties of a 3d$^9$ (Cu$^{2+}$) magnetic insulator Cu2OSO4 measured on both powder and single crystal. The magnetic atoms of this compound form layers, whose geometry can be described either as a system of chains coupled through dimers or as a Kagomé lattice where every 3rd spin is replaced by a dimer. Specific heat and DC-susceptibility show a magnetic transition at 20 K, which is also confirmed by neutron scattering. Magnetic entropy extracted from the specific heat data is consistent with a $S=1/2$ degree of freedom per Cu$^{2+}$, and so is the effective moment extracted from DC-susceptibility. The ground state has been identified by means of neutron diffraction on both powder and single crystal and corresponds to a $\sim120$ degree spin structure in which ferromagnetic intra-dimer alignment results in a net ferrimagnetic moment. No evidence is found for a change in lattice symmetry down to 2 K. Our results suggest that \sample \ represents a new type of model lattice with frustrated interactions where interplay between magnetic order, thermal and quantum fluctuations can be explored.

cond-mat.str-el

Origin of the hump anomalies in the Hall resistance loops of ultrathin SrRuO$_3$/SrIrO$_3$ multilayers

The proposal that very small Néel skyrmions can form in SrRuO$_3$/SrIrO$_3$ epitaxial bilayers and that the electric field-effect can be used to manipulate these skyrmions in gated devices strongly stimulated the recent research of SrRuO$_3$ heterostructures. A strong interfacial Dzyaloshinskii-Moriya interaction, combined with the breaking of inversion symmetry, was considered as the driving force for the formation of skyrmions in SrRuO$_3$/SrIrO$_3$ bilayers. Here, we investigated nominally symmetric heterostructures in which an ultrathin ferromagnetic SrRuO$_3$ layer is sandwiched between large spin-orbit coupling SrIrO$_3$ layers, for which the conditions are not favorable for the emergence of a net interfacial Dzyaloshinskii-Moriya interaction. Previously the formation of skyrmions in the asymmetric SrRuO$_3$/SrIrO$_3$ bilayers was inferred from anomalous Hall resistance loops showing humplike features that resembled topological Hall effect contributions. Symmetric SrIrO$_3$/SrRuO$_3$/SrIrO$_3$ trilayers do not show hump anomalies in the Hall loops. However, the anomalous Hall resistance loops of symmetric multilayers, in which the trilayer is stacked several times, do exhibit the humplike structures, similar to the asymmetric SrRuO$_3$/SrIrO$_3$ bilayers. The origin of the Hall effect loop anomalies likely resides in unavoidable differences in the electronic and magnetic properties of the individual SrRuO$_3$ layers rather than in the formation of skyrmions.

cond-mat.mtrl-sci

The role of hydrophobic interactions in folding of $β$-sheets

Exploring the protein-folding problem has been a long-standing challenge in molecular biology. Protein folding is highly dependent on folding of secondary structures as the way to pave a native folding pathway. Here, we demonstrate that a feature of a large hydrophobic surface area covering most side-chains on one side or the other side of adjacent $β$-strands of a $β$-sheet is prevail in almost all experimentally determined $β$-sheets, indicating that folding of $β$-sheets is most likely triggered by multistage hydrophobic interactions among neighbored side-chains of unfolded polypeptides, enable $β$-sheets fold reproducibly following explicit physical folding codes in aqueous environments. $β$-turns often contain five types of residues characterized with relatively small exposed hydrophobic proportions of their side-chains, that is explained as these residues can block hydrophobic effect among neighbored side-chains in sequence. Temperature dependence of the folding of $β$-sheet is thus attributed to temperature dependence of the strength of the hydrophobicity. The hydrophobic-effect-based mechanism responsible for $β$-sheets folding is verified by bioinformatics analyses of thousands of results available from experiments. The folding codes in amino acid sequence that dictate formation of a $β$-hairpin can be deciphered through evaluating hydrophobic interaction among side-chains of an unfolded polypeptide from a $β$-strand-like thermodynamic metastable state.

q-bio.BM

A hydrophobic-interaction-based mechanism trigger docking between the SARS CoV 2 spike and angiotensin-converting enzyme 2

A recent experimental study found that the binding affinity between the cellular receptor human angiotensin converting enzyme 2 (ACE2) and receptor-binding domain (RBD) in spike (S) protein of novel severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) is more than 10-fold higher than that of the original severe acute respiratory syndrome coronavirus (SARS-CoV). However, main-chain structures of the SARS-CoV-2 RBD are almost the same with that of the SARS-CoV RBD. Understanding physical mechanism responsible for the outstanding affinity between the SARS-CoV-2 S and ACE2 is the "urgent challenge" for developing blockers, vaccines and therapeutic antibodies against the coronavirus disease 2019 (COVID-19) pandemic. Considering the mechanisms of hydrophobic interaction, hydration shell, surface tension, and the shielding effect of water molecules, this study reveals a hydrophobic-interaction-based mechanism by means of which SARS-CoV-2 S and ACE2 bind together in an aqueous environment. The hydrophobic interaction between the SARS-CoV-2 S and ACE2 protein is found to be significantly greater than that between SARS-CoV S and ACE2. At the docking site, the hydrophobic portions of the hydrophilic side chains of SARS-CoV-2 S are found to be involved in the hydrophobic interaction between SARS-CoV-2 S and ACE2. We propose a method to design live attenuated viruses by mutating several key amino acid residues of the spike protein to decrease the hydrophobic surface areas at the docking site. Mutation of a small amount of residues can greatly reduce the hydrophobic binding of the coronavirus to the receptor, which may be significant reduce infectivity and transmissibility of the virus.

q-bio.BM

Mask Detection and Breath Monitoring from Speech: on Data Augmentation, Feature Representation and Modeling

This paper introduces our approaches for the Mask and Breathing Sub-Challenge in the Interspeech COMPARE Challenge 2020. For the mask detection task, we train deep convolutional neural networks with filter-bank energies, gender-aware features, and speaker-aware features. Support Vector Machines follows as the back-end classifiers for binary prediction on the extracted deep embeddings. Several data augmentation schemes are used to increase the quantity of training data and improve our models' robustness, including speed perturbation, SpecAugment, and random erasing. For the speech breath monitoring task, we investigate different bottleneck features based on the Bi-LSTM structure. Experimental results show that our proposed methods outperform the baselines and achieve 0.746 PCC and 78.8% UAR on the Breathing and Mask evaluation set, respectively.

eess.AS