SearcharxivSearch

arXiv subjects

Dan Sun

Publications and source records attributed to Dan Sun.

At least 19 recordsLinked to original sources

Ultrabroadband Integrated Photonics Empowering Full-Spectrum Adaptive Wireless Communications

The forthcoming sixth-generation (6G) and beyond (XG) wireless networks are poised to operate across an expansive frequency range from microwave, millimeter-wave to terahertz bands to support ubiquitous connectivity in diverse application scenarios. This necessitates a one-size-fits-all hardware solution that can be adaptively reconfigured within this wide spectrum to support full-band coverage and dynamic spectrum management. However, existing electrical or photonic-assisted wireless communication solutions see significant challenges in meeting this demand due to the limited bandwidths of individual devices and the intrinsically rigid nature of their system architectures. Here, we demonstrate adaptive wireless communications over an unprecedented frequency range spanning over 100 GHz, driven by a universal thin-film lithium niobate (TFLN) photonic wireless engine. Leveraging the strong Pockels effect and excellent scalability of the TFLN platform, we achieve monolithic integration of essential functional elements, including baseband modulation, broadband wireless-photonic conversion, and reconfigurable carrier/local signal generation. Powered by broadband tunable optoelectronic oscillators, our signal sources operate across a record-wide frequency range from 0.5 GHz to 115 GHz with high frequency stability and consistent coherence. Based on the broadband and reconfigurable integrated photonic solution, we realize, for the first time, full-link wireless communication across 9 consecutive bands, achieving record lane speeds of up to 100 Gbps. The real-time reconfigurability further enables adaptive frequency allocation, a crucial capability to ensure enhanced reliability in complex spectrum environments. Our proposed system marks a significant step towards future full-spectrum and omni-scenario wireless networks.

physics.optics

A LongFormer-Based Framework for Accurate and Efficient Medical Text Summarization

This paper proposes a medical text summarization method based on LongFormer, aimed at addressing the challenges faced by existing models when processing long medical texts. Traditional summarization methods are often limited by short-term memory, leading to information loss or reduced summary quality in long texts. LongFormer, by introducing long-range self-attention, effectively captures long-range dependencies in the text, retaining more key information and improving the accuracy and information retention of summaries. Experimental results show that the LongFormer-based model outperforms traditional models, such as RNN, T5, and BERT in automatic evaluation metrics like ROUGE. It also receives high scores in expert evaluations, particularly excelling in information retention and grammatical accuracy. However, there is still room for improvement in terms of conciseness and readability. Some experts noted that the generated summaries contain redundant information, which affects conciseness. Future research will focus on further optimizing the model structure to enhance conciseness and fluency, achieving more efficient medical text summarization. As medical data continues to grow, automated summarization technology will play an increasingly important role in fields such as medical research, clinical decision support, and knowledge management.

cs.CL

Enhancing Deep Learning with Optimized Gradient Descent: Bridging Numerical Methods and Neural Network Training

Optimization theory serves as a pivotal scientific instrument for achieving optimal system performance, with its origins in economic applications to identify the best investment strategies for maximizing benefits. Over the centuries, from the geometric inquiries of ancient Greece to the calculus contributions by Newton and Leibniz, optimization theory has significantly advanced. The persistent work of scientists like Lagrange, Cauchy, and von Neumann has fortified its progress. The modern era has seen an unprecedented expansion of optimization theory applications, particularly with the growth of computer science, enabling more sophisticated computational practices and widespread utilization across engineering, decision analysis, and operations research. This paper delves into the profound relationship between optimization theory and deep learning, highlighting the omnipresence of optimization problems in the latter. We explore the gradient descent algorithm and its variants, which are the cornerstone of optimizing neural networks. The chapter introduces an enhancement to the SGD optimizer, drawing inspiration from numerical optimization methods, aiming to enhance interpretability and accuracy. Our experiments on diverse deep learning tasks substantiate the improved algorithm's efficacy. The paper concludes by emphasizing the continuous development of optimization theory and its expanding role in solving intricate problems, enhancing computational capabilities, and informing better policy decisions.

cs.LG

Text classification optimization algorithm based on graph neural network

In the field of natural language processing, text classification, as a basic task, has important research value and application prospects. Traditional text classification methods usually rely on feature representations such as the bag of words model or TF-IDF, which overlook the semantic connections between words and make it challenging to grasp the deep structural details of the text. Recently, GNNs have proven to be a valuable asset for text classification tasks, thanks to their capability to handle non-Euclidean data efficiently. However, the existing text classification methods based on GNN still face challenges such as complex graph structure construction and high cost of model training. This paper introduces a text classification optimization algorithm utilizing graph neural networks. By introducing adaptive graph construction strategy and efficient graph convolution operation, the accuracy and efficiency of text classification are effectively improved. The experimental results demonstrate that the proposed method surpasses traditional approaches and existing GNN models across multiple public datasets, highlighting its superior performance and feasibility for text classification tasks.

cs.CL

Research on Deep Learning Model of Feature Extraction Based on Convolutional Neural Network

Neural networks with relatively shallow layers and simple structures may have limited ability in accurately identifying pneumonia. In addition, deep neural networks also have a large demand for computing resources, which may cause convolutional neural networks to be unable to be implemented on terminals. Therefore, this paper will carry out the optimal classification of convolutional neural networks. Firstly, according to the characteristics of pneumonia images, AlexNet and InceptionV3 were selected to obtain better image recognition results. Combining the features of medical images, the forward neural network with deeper and more complex structure is learned. Finally, knowledge extraction technology is used to extract the obtained data into the AlexNet model to achieve the purpose of improving computing efficiency and reducing computing costs. The results showed that the prediction accuracy, specificity, and sensitivity of the trained AlexNet model increased by 4.25 percentage points, 7.85 percentage points, and 2.32 percentage points, respectively. The graphics processing usage has decreased by 51% compared to the InceptionV3 mode.

eess.IV

Research on Optimization of Natural Language Processing Model Based on Multimodal Deep Learning

This project intends to study the image representation based on attention mechanism and multimodal data. By adding multiple pattern layers to the attribute model, the semantic and hidden layers of image content are integrated. The word vector is quantified by the Word2Vec method and then evaluated by a word embedding convolutional neural network. The published experimental results of the two groups were tested. The experimental results show that this method can convert discrete features into continuous characters, thus reducing the complexity of feature preprocessing. Word2Vec and natural language processing technology are integrated to achieve the goal of direct evaluation of missing image features. The robustness of the image feature evaluation model is improved by using the excellent feature analysis characteristics of a convolutional neural network. This project intends to improve the existing image feature identification methods and eliminate the subjective influence in the evaluation process. The findings from the simulation indicate that the novel approach has developed is viable, effectively augmenting the features within the produced representations.

cs.CL

Advancements in Feature Extraction Recognition of Medical Imaging Systems Through Deep Learning Technique

This study introduces a novel unsupervised medical image feature extraction method that employs spatial stratification techniques. An objective function based on weight is proposed to achieve the purpose of fast image recognition. The algorithm divides the pixels of the image into multiple subdomains and uses a quadtree to access the image. A technique for threshold optimization utilizing a simplex algorithm is presented. Aiming at the nonlinear characteristics of hyperspectral images, a generalized discriminant analysis algorithm based on kernel function is proposed. In this project, a hyperspectral remote sensing image is taken as the object, and we investigate its mathematical modeling, solution methods, and feature extraction techniques. It is found that different types of objects are independent of each other and compact in image processing. Compared with the traditional linear discrimination method, the result of image segmentation is better. This method can not only overcome the disadvantage of the traditional method which is easy to be affected by light, but also extract the features of the object quickly and accurately. It has important reference significance for clinical diagnosis.

eess.IV

Theoretical Analysis of Meta Reinforcement Learning: Generalization Bounds and Convergence Guarantees

This research delves deeply into Meta Reinforcement Learning (Meta RL) through a exploration focusing on defining generalization limits and ensuring convergence. By employing a approach this article introduces an innovative theoretical framework to meticulously assess the effectiveness and performance of Meta RL algorithms. We present an explanation of generalization limits measuring how well these algorithms can adapt to learning tasks while maintaining consistent results. Our analysis delves into the factors that impact the adaptability of Meta RL revealing the relationship, between algorithm design and task complexity. Additionally we establish convergence assurances by proving conditions under which Meta RL strategies are guaranteed to converge towards solutions. We examine the convergence behaviors of Meta RL algorithms across scenarios providing a comprehensive understanding of the driving forces behind their long term performance. This exploration covers both convergence and real time efficiency offering a perspective, on the capabilities of these algorithms.

cs.LG

Adapting LLMs for Efficient Context Processing through Soft Prompt Compression

The rapid advancement of Large Language Models (LLMs) has inaugurated a transformative epoch in natural language processing, fostering unprecedented proficiency in text generation, comprehension, and contextual scrutiny. Nevertheless, effectively handling extensive contexts, crucial for myriad applications, poses a formidable obstacle owing to the intrinsic constraints of the models' context window sizes and the computational burdens entailed by their operations. This investigation presents an innovative framework that strategically tailors LLMs for streamlined context processing by harnessing the synergies among natural language summarization, soft prompt compression, and augmented utility preservation mechanisms. Our methodology, dubbed SoftPromptComp, amalgamates natural language prompts extracted from summarization methodologies with dynamically generated soft prompts to forge a concise yet semantically robust depiction of protracted contexts. This depiction undergoes further refinement via a weighting mechanism optimizing information retention and utility for subsequent tasks. We substantiate that our framework markedly diminishes computational overhead and enhances LLMs' efficacy across various benchmarks, while upholding or even augmenting the caliber of the produced content. By amalgamating soft prompt compression with sophisticated summarization, SoftPromptComp confronts the dual challenges of managing lengthy contexts and ensuring model scalability. Our findings point towards a propitious trajectory for augmenting LLMs' applicability and efficiency, rendering them more versatile and pragmatic for real-world applications. This research enriches the ongoing discourse on optimizing language models, providing insights into the potency of soft prompts and summarization techniques as pivotal instruments for the forthcoming generation of NLP solutions.

cs.LG

An Evolution Kernel Method for Graph Classification through Heat Diffusion Dynamics

Autonomous individuals establish a structural complex system through pairwise connections and interactions. Notably, the evolution reflects the dynamic nature of each complex system since it recodes a series of temporal changes from the past, the present into the future. Different systems follow distinct evolutionary trajectories, which can serve as distinguishing traits for system classification. However, modeling a complex system's evolution is challenging for the graph model because the graph is typically a snapshot of the static status of a system, and thereby hard to manifest the long-term evolutionary traits of a system entirely. To address this challenge, we suggest utilizing a heat-driven method to generate temporal graph augmentation. This approach incorporates the physics-based heat kernel and DropNode technique to transform each static graph into a sequence of temporal ones. This approach effectively describes the evolutional behaviours of the system, including the retention or disappearance of elements at each time point based on the distributed heat on each node. Additionally, we propose a dynamic time-wrapping distance GDTW to quantitatively measure the distance between pairwise evolutionary systems through optimal matching. The resulting approach, called the Evolution Kernel method, has been successfully applied to classification problems in real-world structural graph datasets. The results yield significant improvements in supervised classification accuracy over a series of baseline methods.

cs.LG

Anonymous Pattern Molecular Fingerprint and its Applications on Property Identification

Molecular fingerprints are significant cheminformatics tools to map molecules into vectorial space according to their characteristics in diverse functional groups, atom sequences, and other topological structures. In this paper, we set out to investigate a novel molecular fingerprint \emph{Anonymous-FP} that possesses abundant perception about the underlying interactions shaped in small, medium, and large molecular scale links. In detail, the possible inherent atom chains are sampled from each molecule and are extended in a certain anonymous pattern. After that, the molecular fingerprint \emph{Anonymous-FP} is encoded in virtue of the Natural Language Processing technique \emph{PV-DBOW}. \emph{Anonymous-FP} is studied on molecular property identification and has shown valuable advantages such as rich information content, high experimental performance, and full structural significance. During the experimental verification, the scale of the atom chain or its anonymous manner matters significantly to the overall representation ability of \emph{Anonymous-FP}. Generally, the typical scale $r = 8$ enhances the performance on a series of real-world molecules, and specifically, the accuracy could level up to above $93\%$ on all NCI datasets.

stat.AP

Metric Distribution to Vector: Constructing Data Representation via Broad-Scale Discrepancies

Graph embedding provides a feasible methodology to conduct pattern classification for graph-structured data by mapping each data into the vectorial space. Various pioneering works are essentially coding method that concentrates on a vectorial representation about the inner properties of a graph in terms of the topological constitution, node attributions, link relations, etc. However, the classification for each targeted data is a qualitative issue based on understanding the overall discrepancies within the dataset scale. From the statistical point of view, these discrepancies manifest a metric distribution over the dataset scale if the distance metric is adopted to measure the pairwise similarity or dissimilarity. Therefore, we present a novel embedding strategy named $\mathbf{MetricDistribution2vec}$ to extract such distribution characteristics into the vectorial representation for each data. We demonstrate the application and effectiveness of our representation method in the supervised prediction tasks on extensive real-world structural graph datasets. The results have gained some unexpected increases compared with a surge of baselines on all the datasets, even if we take the lightweight models as classifiers. Moreover, the proposed methods also conducted experiments in Few-Shot classification scenarios, and the results still show attractive discrimination in rare training samples based inference.

cs.LG

Optimizing Counterdiabaticity by Variational Quantum Circuits

Utilizing counterdiabatic (CD) driving - aiming at suppression of diabatic transition - in digitized adiabatic evolution have garnered immense interest in quantum protocols and algorithms. However, improving the approximate CD terms with a nested commutator ansatz is a challenging task. In this work, we propose a technique of finding optimal coefficients of the CD terms using a variational quantum circuit. By classical optimizations routines, the parameters of this circuit are optimized to provide the coefficients corresponding to the CD terms. Then their improved performance is exemplified in Greenberger-Horne-Zeilinger state preparation on nearest-neighbor Ising model. Finally, we also show the advantage over the usual quantum approximation optimization algorithm, in terms of fidelity with bounded time.

quant-ph

High-temperature superconductivity in hydrides: experimental evidence and details

Since the discovery of superconductivity at 200 K in H3S [1] similar or higher transition temperatures, Tcs, have been reported for various hydrogen-rich compounds under ultra-high pressures [2]. Superconductivity was experimentally proved by different methods, including electrical resistance, magnetic susceptibility, optical infrared, and nuclear resonant scattering measurements. The crystal structures of superconducting phases were determined by X-ray diffraction. Numerous electrical transport measurements demonstrate the typical behaviour of a conventional phonon-mediated superconductor: zero resistance below Tc, the shift of Tc to lower temperatures under external magnetic fields, and pronounced isotope effect. Remarkably, the results are in good agreement with the theoretical predictions, which describe superconductivity in hydrides within the framework of the conventional BCS theory. However, despite this acknowledgment, experimental evidence for the superconducting state in these compounds has recently been treated with criticism [3, 4], which apparently stems from misunderstanding and misinterpretation of complicated experiments performed under very high pressures. Here, we describe in greater detail the experiments revealing high-temperature superconductivity in hydrides under high pressures. We show that the arguments against superconductivity [3, 4] can be either refuted or explained. The experiments on the high-temperature superconductivity in hydrides clearly contradict the theory of hole superconductivity [4] and eliminate it [3].

cond-mat.supr-con

Heisenberg spins on an anisotropic triangular lattice: PdCrO2 under uniaxial stress

When Heisenberg spins interact antiferromagnetically on a triangular lattice and nearest-neighbor interactions dominate, the ground state is 120$^{\circ}$ antiferromagnetism. In this work, we probe the response of this state to lifting the triangular symmetry, through investigation of the triangular antiferromagnet PdCrO$_2$ under uniaxial stress by neutron diffraction and resistivity measurements. The periodicity of the magnetic order is found to change rapidly with applied stress; the rate of change indicates that the magnetic anisotropy is roughly forty times the stress-induced bond length anisotropy. At low stress, the incommensuration period becomes extremely long, on the order of 1000 lattice spacings; no locking of the magnetism to commensurate periodicity is detected. Separately, the magnetic structure is found to undergo a first-order transition at a compressive stress of $\sim$0.4 GPa, at which the interlayer ordering switches from a double- to a single-q structure.

cond-mat.str-el

A Graph Data Augmentation Strategy with Entropy Preservation

The Graph Convolutional Networks (GCN) proposed by Kipf and Welling is an effective model for semi-supervised learning, but faces the obstacle of over-smoothing, which will weaken the representation ability of GCN. Recently some works are proposed to tackle above limitation by randomly perturbing graph topology or feature matrix to generate data augmentations as input for training. However, these operations inevitably do damage to the integrity of information structures and have to sacrifice the smoothness of feature manifold. In this paper, we first introduce a novel graph entropy definition as a measure to quantitatively evaluate the smoothness of a data manifold and then point out that this graph entropy is controlled by triangle motif-based information structures. Considering the preservation of graph entropy, we propose an effective strategy to generate randomly perturbed training data but maintain both graph topology and graph entropy. Extensive experiments have been conducted on real-world datasets and the results verify the effectiveness of our proposed method in improving semi-supervised node classification accuracy compared with a surge of baselines. Beyond that, our proposed approach could significantly enhance the robustness of training process for GCN.

cs.LG

Graph Classification Based on Skeleton and Component Features

Most existing popular methods for learning graph embedding only consider fixed-order global structural features and lack structures hierarchical representation. To address this weakness, we propose a novel graph embedding algorithm named GraphCSC that realizes classification based on skeleton information using fixed-order structures learned in anonymous random walks manner, and component information using different size subgraphs. Two graphs are similar if their skeletons and components are both similar, thus in our model, we integrate both of them together into embeddings as graph homogeneity characterization. We demonstrate our model on different datasets in comparison with a comprehensive list of up-to-date state-of-the-art baselines, and experiments show that our work is superior in real-world graph classification tasks.

cs.LG

High-temperature superconductivity on the verge of a structural instability in lanthanum superhydride

A possibility of high, room-temperature superconductivity was predicted for metallic hydrogen in the 1960s. However, metallization and superconductivity of hydrogen are yet to be unambiguously demonstrated in the laboratory and may require pressures as high as 5 million atmospheres. Rare earth based "superhydrides" such as LaH10 can be considered a close approximation of metallic hydrogen even though they form at moderately lower pressures. In superhydrides the predominance of H-H metallic bonds and high superconducting transition temperatures bear the hallmarks of metallic hydrogen. Still, experimental studies revealing the key factors controlling their superconductivity are scarce. Here, we report on the pressure and magnetic field response of the superconducting order observed in LaH10. For LaH10 we find a correlation between superconductivity and a structural instability, strongly affecting the lattice vibrations responsible for the superconductivity.

cond-mat.supr-con