Searcharxiv⌕ Search

arXiv subjects

Lin Zhang

Publications and source records attributed to Lin Zhang.

At least 271 records · Page 15Linked to original sources

$ \text{T}^3 $OMVP: A Transformer-based Time and Team Reinforcement Learning Scheme for Observation-constrained Multi-Vehicle Pursuit in Urban Area

Smart Internet of Vehicles (IoVs) combined with Artificial Intelligence (AI) will contribute to vehicle decision-making in the Intelligent Transportation System (ITS). Multi-Vehicle Pursuit games (MVP), a multi-vehicle cooperative ability to capture mobile targets, is becoming a hot research topic gradually. Although there are some achievements in the field of MVP in the open space environment, the urban area brings complicated road structures and restricted moving spaces as challenges to the resolution of MVP games. We define an Observation-constrained MVP (OMVP) problem in this paper and propose a Transformer-based Time and Team Reinforcement Learning scheme ($ \text{T}^3 $OMVP) to address the problem. First, a new multi-vehicle pursuit model is constructed based on decentralized partially observed Markov decision processes (Dec-POMDP) to instantiate this problem. Second, by introducing and modifying the transformer-based observation sequence, QMIX is redefined to adapt to the complicated road structure, restricted moving spaces and constrained observations, so as to control vehicles to pursue the target combining the vehicle's observations. Third, a multi-intersection urban environment is built to verify the proposed scheme. Extensive experimental results demonstrate that the proposed $ \text{T}^3 $OMVP scheme achieves significant improvements relative to state-of-the-art QMIX approaches by 9.66%~106.25%. Code is available at https://github.com/pipihaiziguai/T3OMVP.

cs.AI↗

Changeable Rate and Novel Quantization for CSI Feedback Based on Deep Learning

Deep learning (DL)-based channel state information (CSI) feedback improves the capacity and energy efficiency of massive multiple-input multiple-output (MIMO) systems in frequency division duplexing mode. However, multiple neural networks with different lengths of feedback overhead are required by time-varying bandwidth resources. The storage space required at the user equipment (UE) and the base station (BS) for these models increases linearly with the number of models. In this paper, we propose a DL-based changeable-rate framework with novel quantization scheme to improve the efficiency and feasibility of CSI feedback systems. This framework can reutilize all the network layers to achieve overhead-changeable CSI feedback to optimize the storage efficiency at the UE and the BS sides. Designed quantizer in this framework can avoid the normalization and gradient problems faced by traditional quantization schemes. Specifically, we propose two DL-based changeable-rate CSI feedback networks CH-CsiNetPro and CH-DualNetSph by introducing a feedback overhead control unit. Then, a pluggable quantization block (PQB) is developed to further improve the encoding efficiency of CSI feedback in an end-to-end way. Compared with existing CSI feedback methods, the proposed framework saves the storage space by about 50% with changeable-rate scheme and improves the encoding efficiency with the quantization module.

cs.IT↗

Sharp decay estimates for massless Dirac fields on a Schwarzschild background

We consider the explicit asymptotic profile of massless Dirac fields on a Schwarzschild background. First, we prove for the spin $s=\pm \frac{1}{2}$ components of the Dirac field a uniform bound of a positive definite energy and an integrated local energy decay estimate from a symmetric hyperbolic wave system. Based on these estimates, we further show that these components have globally pointwise decay $fv^{-3/2-s}τ^{-5/2+s}$ as both an upper and a lower bound outside the black hole, with function $f$ finite and explicitly expressed in terms of the initial data and the coordinates. This establishes the validity of the conjectured Price's law for massless Dirac fields outside a Schwarzschild black hole.

math.AP↗

A characterization of maximally entangled two-qubit states

As already known by Rana's result \href{https://doi.org/10.1103/PhysRevA.87.054301}{[\pra {\bf87} (2013) 054301]}, all eigenvalues of any partial-transposed bipartite state fall within the closed interval $[-\frac12,1]$. In this note, we study a family of bipartite quantum states whose minimal eigenvalues of partial-transposed states being $-\frac12$. For a two-qubit system, we find that the minimal eigenvalue of its partial-transposed state is $-\frac12$ if and only if such two-qubit state must be maximally entangled. However this result does not hold in general for a two-qudit system when the dimensions of the underlying space are larger than two.

quant-ph↗

A discussion of measuring the top-1 percent most-highly cited publications: Quality and impact of Chinese papers

The top 1 percent most highly cited articles are watched closely as the vanguards of the sciences. Using Web of Science data, one can find that China had overtaken the USA in the relative participation in the top 1 percent in 2019, after outcompeting the EU on this indicator in 2015. However, this finding contrasts with repeated reports of Western agencies that the quality of Chinese output in science is lagging other advanced nations, even as it has caught up in numbers of articles. The difference between the results presented here and the previous results depends mainly upon field normalizations, which classify source journals by discipline. Average citation rates of these subsets are commonly used as a baseline so that one can compare among disciplines. However, the expected value of the top 1 percent of a sample of N papers is N 100, ceteris paribus. Using the average citation rates as expected values, errors are introduced by using the mean of highly skewed distributions and a specious precision in the delineations of the subsets. Classifications can be used for the decomposition, but not for the normalization. When the data is thus decomposed, the USA ranks ahead of China in biomedical fields such as virology. Although the number of papers is smaller, China outperforms the US in the field of Business and Finance in the Social Sciences Citation Index when p is less than .05. Using percentile ranks, subsets other than indexing based classifications can be tested for the statistical significance of differences among them.

cs.DL↗

A General Auxiliary Controller for Multi-agent Flocking

We aim to improve the performance of multi-agent flocking behavior by quantifying the structural significance of each agent. We designed a confidence score(ConfScore) to measure the spatial significance of each agent. The score will be used by an auxiliary controller to refine the velocity of agents. The agents will be enforced to follow the motion of the leader agents whose ConfScores are high. We demonstrate the efficacy of the auxiliary controller by applying it to several existing algorithms including learning-based and non-learning-based methods. Furthermore, we examined how the auxiliary controller can help improve the performance under different settings of communication radius, number of agents and maximum initial velocity.

cs.MA↗

Multi-resolution Super Learner for Voxel-wise Classification of Prostate Cancer Using Multi-parametric MRI

While current research has shown the importance of Multi-parametric MRI (mpMRI) in diagnosing prostate cancer (PCa), further investigation is needed for how to incorporate the specific structures of the mpMRI data, such as the regional heterogeneity and between-voxel correlation within a subject. This paper proposes a machine learning-based method for improved voxel-wise PCa classification by taking into account the unique structures of the data. We propose a multi-resolution modeling approach to account for regional heterogeneity, where base learners trained locally at multiple resolutions are combined using the super learner, and account for between-voxel correlation by efficient spatial Gaussian kernel smoothing. The method is flexible in that the super learner framework allows implementation of any classifier as the base learner, and can be easily extended to classifying cancer into more sub-categories. We describe detailed classification algorithm for the binary PCa status, as well as the ordinal clinical significance of PCa for which a weighted likelihood approach is implemented to enhance the detection of the less prevalent cancer categories. We illustrate the advantages of the proposed approach over conventional modeling and machine learning approaches through simulations and application to in vivo data.

stat.ML↗

Uncertainty regions of observables and state-independent uncertainty relations

The optimal state-independent lower bounds for the sum of variances or deviations of observables are of significance for the growing number of experiments that reach the uncertainty limited regime. We present a framework for computing the tight uncertainty relations of variance or deviation via determining the uncertainty regions, which are formed by the tuples of two or more of quantum observables in random quantum states induced from the uniform Haar measure on the purified states. From the analytical formulae of these uncertainty regions, we present state-independent uncertainty inequalities satisfied by the sum of variances or deviations of two, three and arbitrary many observables, from which experimentally friend entanglement detection criteria are derived for bipartite and tripartite systems.

quant-ph↗

1D photonic crystal direct bandgap GeSn-on-insulator laser

GeSn alloys have been regarded as a potential lasing material for a complementary metal-oxide-semiconductor (CMOS)-compatible light source. Despite their remarkable progress, all GeSn lasers reported to date have large device footprints and active areas, which prevent the realization of densely integrated on-chip lasers operating at low power consumption. Here, we present a 1D photonic crystal (PC) nanobeam with a very small device footprint of 7 $μm^2$ and a compact active area of ~1.2 $μm^2$ on a high-quality GeSn-on-insulator (GeSnOI) substrate. We also report that the improved directness in our strain-free nanobeam lasers leads to a lower threshold density and a higher operating temperature compared to the compressive strained counterparts. The threshold density of the strain-free nanobeam laser is ~18.2 kW cm$^{ -2}$ at 4 K, which is significantly lower than that of the unreleased nanobeam laser (~38.4 kW cm$^{ -2}$ at 4 K). Lasing in the strain-free nanobeam device persists up to 90 K, whereas the unreleased nanobeam shows a quenching of the lasing at a temperature of 70 K. Our demonstration offers a new avenue towards developing practical group-IV light sources with high-density integration and low power consumption.

physics.optics↗

CdtGRN: Construction of qualitative time-delayed gene regulatory networks with a deep learning method

Background:Gene regulations often change over time rather than being constant. But many of gene regulatory networks extracted from databases are static. The tumor suppressor gene $P53$ is involved in the pathogenesis of many tumors, and its inhibition effects occur after a certain period. Therefore, it is of great significance to elucidate the regulation mechanism over time points. Result:A qualitative method for representing dynamic gene regulatory network is developed, called CdtGRN. It adopts the combination of convolutional neural networks(CNN) and fully connected networks(DNN) as the core mechanism of prediction. The ionizing radiation Affymetrix dataset (E-MEXP-549) was obtained at ArrayExpress, by microarray gene expression levels predicting relations between regulation. CdtGRN is tested against a time-delayed gene regulatory network with $22,284$ genes related to $P53$. The accuracy of CdtGRN reaches 92.07$\%$ on the classification of conservative verification set, and a kappa coefficient reaches $0.84$ and an average AUC accuracy is 94.25$\%$. This resulted in the construction of. Conclusion:The algorithm and program we developed in our study would be useful for identifying dynamic gene regulatory networks, and objectively analyze the delay of the regulatory relationship by analyzing the gene expression levels at different time points. The time-delayed gene regulatory network of $P53$ is also inferred and represented qualitatively, which is helpful to understand the pathological mechanism of tumors.

q-bio.MN↗

Inter-intra Variant Dual Representations forSelf-supervised Video Recognition

Contrastive learning applied to self-supervised representation learning has seen a resurgence in deep models. In this paper, we find that existing contrastive learning based solutions for self-supervised video recognition focus on inter-variance encoding but ignore the intra-variance existing in clips within the same video. We thus propose to learn dual representations for each clip which (\romannumeral 1) encode intra-variance through a shuffle-rank pretext task; (\romannumeral 2) encode inter-variance through a temporal coherent contrastive loss. Experiment results show that our method plays an essential role in balancing inter and intra variances and brings consistent performance gains on multiple backbones and contrastive learning frameworks. Integrated with SimCLR and pretrained on Kinetics-400, our method achieves $\textbf{82.0\%}$ and $\textbf{51.2\%}$ downstream classification accuracy on UCF101 and HMDB51 test sets respectively and $\textbf{46.1\%}$ video retrieval accuracy on UCF101, outperforming both pretext-task based and contrastive learning based counterparts. Our code is available at \href{https://github.com/lzhangbj/DualVar}{https://github.com/lzhangbj/DualVar}.

cs.CV↗

Content-Preserving Unpaired Translation from Simulated to Realistic Ultrasound Images

Interactive simulation of ultrasound imaging greatly facilitates sonography training. Although ray-tracing based methods have shown promising results, obtaining realistic images requires substantial modeling effort and manual parameter tuning. In addition, current techniques still result in a significant appearance gap between simulated images and real clinical scans. Herein we introduce a novel content-preserving image translation framework (ConPres) to bridge this appearance gap, while maintaining the simulated anatomical layout. We achieve this goal by leveraging both simulated images with semantic segmentations and unpaired in-vivo ultrasound scans. Our framework is based on recent contrastive unpaired translation techniques and we propose a regularization approach by learning an auxiliary segmentation-to-real image translation task, which encourages the disentanglement of content and style. In addition, we extend the generator to be class-conditional, which enables the incorporation of additional losses, in particular a cyclic consistency loss, to further improve the translation quality. Qualitative and quantitative comparisons against state-of-the-art unpaired translation methods demonstrate the superiority of our proposed framework.

eess.IV↗

Estimating coherence with respect to general quantum measurements

The conventional coherence is defined with respect to a fixed orthonormal basis, i.e., to a von Neumann measurement. Recently, generalized quantum coherence with respect to general positive operator-valued measurements (POVMs) has been presented. Several well-defined coherence measures, such as the relative entropy of coherence $C_{r}$, the $l_{1}$ norm of coherence $C_{l_{1}}$ and the coherence $C_{T,α}$ based on Tsallis relative entropy with respect to general POVMs have been obtained. In this work, we investigate the properties of $C_{r}$, $l_{1}$ and $C_{T,α}$. We estimate the upper bounds of $C_{l_{1}}$; we show that the minimal error probability of the least square measurement state discrimination is given by $C_{T,1/2}$; we derive the uncertainty relations given by $C_{r}$, and calculate the average values of $C_{r}$, $C_{T,α}$ and $C_{l_{1}}$ over random pure quantum states. All these results include the corresponding results of the conventional coherence as special cases.

quant-ph↗

CSI Sensing and Feedback: A Semi-Supervised Learning Approach

Deep learning-based (DL-based) channel state information (CSI) feedback for a Massive multiple-input multiple-output (MIMO) system has proved to be a creative and efficient application. However, the existing systems ignored the wireless channel environment variation sensing, e.g., indoor and outdoor scenarios. Moreover, systems training requires excess pre-labeled CSI data, which is often unavailable. In this letter, to address these issues, we first exploit the rationality of introducing semi-supervised learning on CSI feedback, then one semi-supervised CSI sensing and feedback Network ($S^2$CsiNet) with three classifiers comparisons is proposed. Experiment shows that $S^2$CsiNet primarily improves the feasibility of the DL-based CSI feedback system by \textbf{\textit{indoor}} and \textbf{\textit{outdoor}} environment sensing and at most 96.2\% labeled dataset decreasing and secondarily boost the system performance by data distillation and latent information mining.

eess.SP↗

Estimating Mean Speed-of-Sound from Sequence-Dependent Geometric Disparities

In ultrasound beamforming, focusing time delays are typically computed with a spatially constant speed-of-sound (SoS) assumption. A mismatch between beamforming and true medium SoS then leads to aberration artifacts. Other imaging techniques such as spatially-resolved SoS reconstruction using tomographic techniques also rely on a good SoS estimate for initial beamforming. In this work, we exploit spatially-varying geometric disparities in the transmit and receive paths of multiple sequences for estimating a mean medium SoS. We use images from diverging waves beamformed with an assumed SoS, and propose a model fitting method for estimating the SoS offset. We demonstrate the effectiveness of our proposed method for tomographic SoS reconstruction. With corrected beamforming SoS, the reconstruction accuracy on simulated data was improved by 63% and 29%, respectively, for an initial SoS over- and under-estimation of 1.5%. We further demonstrate our proposed method on a breast phantom, indicating substantial improvement in contrast-to-noise ratio for local SoS mapping.

eess.IV↗

Multi-Task Learning in Utterance-Level and Segmental-Level Spoof Detection

In this paper, we provide a series of multi-tasking benchmarks for simultaneously detecting spoofing at the segmental and utterance levels in the PartialSpoof database. First, we propose the SELCNN network, which inserts squeeze-and-excitation (SE) blocks into a light convolutional neural network (LCNN) to enhance the capacity of hidden feature selection. Then, we implement multi-task learning (MTL) frameworks with SELCNN followed by bidirectional long short-term memory (Bi-LSTM) as the basic model. We discuss MTL in PartialSpoof in terms of architecture (uni-branch/multi-branch) and training strategies (from-scratch/warm-up) step-by-step. Experiments show that the multi-task model performs relatively better than single-task models. Also, in MTL, a binary-branch architecture more adequately utilizes information from two levels than a uni-branch model. For the binary-branch architecture, fine-tuning a warm-up model works better than training from scratch. Models can handle both segment-level and utterance-level predictions simultaneously overall under a binary-branch multi-task architecture. Furthermore, the multi-task model trained by fine-tuning a segmental warm-up model performs relatively better at both levels except on the evaluation set for segmental detection. Segmental detection should be explored further.

cs.SD↗

Maximum Likelihood Estimation for Multimodal Learning with Missing Modality

Multimodal learning has achieved great successes in many scenarios. Compared with unimodal learning, it can effectively combine the information from different modalities to improve the performance of learning tasks. In reality, the multimodal data may have missing modalities due to various reasons, such as sensor failure and data transmission error. In previous works, the information of the modality-missing data has not been well exploited. To address this problem, we propose an efficient approach based on maximum likelihood estimation to incorporate the knowledge in the modality-missing data. Specifically, we design a likelihood function to characterize the conditional distribution of the modality-complete data and the modality-missing data, which is theoretically optimal. Moreover, we develop a generalized form of the softmax function to effectively implement maximum likelihood estimation in an end-to-end manner. Such training strategy guarantees the computability of our algorithm capably. Finally, we conduct a series of experiments on real-world multimodal datasets. Our results demonstrate the effectiveness of the proposed approach, even when 95% of the training data has missing modality.

cs.LG↗

MT-ORL: Multi-Task Occlusion Relationship Learning

Retrieving occlusion relation among objects in a single image is challenging due to sparsity of boundaries in image. We observe two key issues in existing works: firstly, lack of an architecture which can exploit the limited amount of coupling in the decoder stage between the two subtasks, namely occlusion boundary extraction and occlusion orientation prediction, and secondly, improper representation of occlusion orientation. In this paper, we propose a novel architecture called Occlusion-shared and Path-separated Network (OPNet), which solves the first issue by exploiting rich occlusion cues in shared high-level features and structured spatial information in task-specific low-level features. We then design a simple but effective orthogonal occlusion representation (OOR) to tackle the second issue. Our method surpasses the state-of-the-art methods by 6.1%/8.3% Boundary-AP and 6.5%/10% Orientation-AP on standard PIOD/BSDS ownership datasets. Code is available at https://github.com/fengpanhe/MT-ORL.

cs.CV↗