Searcharxiv⌕ Search

arXiv subjects

Xue Han

Publications and source records attributed to Xue Han.

At least 37 records · Page 2Linked to original sources

Investigating and Scaling up Code-Switching for Multilingual Language Model Pre-Training

Large language models (LLMs) exhibit remarkable multilingual capabilities despite the extreme language imbalance in the pre-training data. In this paper, we closely examine the reasons behind this phenomenon, focusing on the pre-training corpus. We find that the existence of code-switching, alternating between different languages within a context, is key to multilingual capabilities. We conduct an analysis to investigate code-switching in the pre-training corpus, examining its presence and categorizing it into four types within two quadrants. We then assess its impact on multilingual performance. These types of code-switching data are unbalanced in proportions and demonstrate different effects on facilitating language transfer. To better explore the power of code-switching for language alignment during pre-training, we investigate the strategy of synthetic code-switching. We continuously scale up the synthetic code-switching data and observe remarkable improvements in both benchmarks and representation space. Extensive experiments indicate that incorporating synthetic code-switching data enables better language alignment and generalizes well to high, medium, and low-resource languages with pre-training corpora of varying qualities.

cs.CL↗

Large Language Models Are Cross-Lingual Knowledge-Free Reasoners

Large Language Models have demonstrated impressive reasoning capabilities across multiple languages. However, the relationship between capabilities in different languages is less explored. In this work, we decompose the process of reasoning tasks into two separated components: knowledge retrieval and knowledge-free reasoning, and analyze the relationship between cross-lingual transferability and these two components. With adapted commonsense reasoning datasets and constructed knowledge-free reasoning datasets, we show that the knowledge-free reasoning capability can be nearly perfectly transferred across various source-target language directions despite the secondary impact of resource in some specific target languages, while cross-lingual knowledge retrieval significantly hinders the transfer. Moreover, by analyzing the hidden states and feed-forward network neuron activation during the reasoning, we show that higher similarity of hidden representations and larger overlap of activated neurons could explain the better cross-lingual transferability of knowledge-free reasoning than knowledge retrieval. Thus, we hypothesize that knowledge-free reasoning shares similar neurons in different languages for reasoning, while knowledge is stored separately in different languages. Our code and data is available at: https://github.com/NJUNLP/Knowledge-Free-Reasoning.

cs.CL↗

A new framework for X-ray absorption spectroscopy data analysis based on machine learning: XASDAML

X-ray absorption spectroscopy (XAS) is a powerful technique to probe the electronic and structural properties of materials. With the rapid growth in both the volume and complexity of XAS datasets driven by advancements in synchrotron radiation facilities, there is an increasing demand for advanced computational tools capable of efficiently analyzing large-scale data. To address these needs, we introduce XASDAML,a flexible, machine learning based framework that integrates the entire data-processing workflow-including dataset construction for spectra and structural descriptors, data filtering, ML modeling, prediction, and model evaluation-into a unified platform. Additionally, it supports comprehensive statistical analysis, leveraging methods such as principal component analysis and clustering to reveal potential patterns and relationships within large datasets. Each module operates independently, allowing users to modify or upgrade modules in response to evolving research needs or technological advances. Moreover, the platform provides a user-friendly interface via Jupyter Notebook, making it accessible to researchers at varying levels of expertise. The versatility and effectiveness of XASDAML are exemplified by its application to a copper dataset, where it efficiently manages large and complex data, supports both supervised and unsupervised machine learning models, provides comprehensive statistics for structural descriptors, generates spectral plots, and accurately predicts coordination numbers and bond lengths. Furthermore, the platform streamlining the integration of XAS with machine learning and lowering the barriers to entry for new users.

physics.comp-ph↗

SCSC: A Novel Standards-Compatible Semantic Communication Framework for Image Transmission

Joint source-channel coding (JSCC) is a promising paradigm for next-generation communication systems, particularly in challenging transmission environments. In this paper, we propose a novel standard-compatible JSCC framework for the transmission of images over multiple-input multiple-output (MIMO) channels. Different from the existing end-to-end AI-based DeepJSCC schemes, our framework consists of learnable modules that enable communication using conventional separate source and channel codes (SSCC), which makes it amenable for easy deployment on legacy systems. Specifically, the learnable modules involve a preprocessing-empowered network (PPEN) for preserving essential semantic information, and a precoder \& combiner-enhanced network (PCEN) for efficient transmission over a resource-constrained MIMO channel. We treat existing compression and channel coding modules as non-trainable blocks. Since the parameters of these modules are non-differentiable, we employ a proxy network that mimics their operations when training the learnable modules. Numerical results demonstrate that our scheme can save more than 29\% of the channel bandwidth, and requires lower complexity compared to the constrained baselines. We also show its generalization capability to unseen datasets and tasks through extensive experiments.

cs.IT↗

MoE-LPR: Multilingual Extension of Large Language Models through Mixture-of-Experts with Language Priors Routing

Large Language Models (LLMs) are often English-centric due to the disproportionate distribution of languages in their pre-training data. Enhancing non-English language capabilities through post-pretraining often results in catastrophic forgetting of the ability of original languages. Previous methods either achieve good expansion with severe forgetting or slight forgetting with poor expansion, indicating the challenge of balancing language expansion while preventing forgetting. In this paper, we propose a method called MoE-LPR (Mixture-of-Experts with Language Priors Routing) to alleviate this problem. MoE-LPR employs a two-stage training approach to enhance the multilingual capability. First, the model is post-pretrained into a Mixture-of-Experts (MoE) architecture by upcycling, where all the original parameters are frozen and new experts are added. In this stage, we focus improving the ability on expanded languages, without using any original language data. Then, the model reviews the knowledge of the original languages with replay data amounting to less than 1% of post-pretraining, where we incorporate language priors routing to better recover the abilities of the original languages. Evaluations on multiple benchmarks show that MoE-LPR outperforms other post-pretraining methods. Freezing original parameters preserves original language knowledge while adding new experts preserves the learning ability. Reviewing with LPR enables effective utilization of multilingual knowledge within the parameters. Additionally, the MoE architecture maintains the same inference overhead while increasing total model parameters. Extensive experiments demonstrate MoE-LPR's effectiveness in improving expanded languages and preserving original language proficiency with superior scalability. Code and scripts are freely available at https://github.com/zjwang21/MoE-LPR.git.

cs.CL↗

Getting More from Less: Large Language Models are Good Spontaneous Multilingual Learners

Recently, Large Language Models (LLMs) have shown impressive language capabilities. While most of the existing LLMs have very unbalanced performance across different languages, multilingual alignment based on translation parallel data is an effective method to enhance the LLMs' multilingual capabilities. In this work, we discover and comprehensively investigate the spontaneous multilingual alignment improvement of LLMs. We find that LLMs instruction-tuned on the question translation data (i.e. without annotated answers) are able to encourage the alignment between English and a wide range of languages, even including those unseen during instruction-tuning. Additionally, we utilize different settings and mechanistic interpretability methods to analyze the LLM's performance in the multilingual scenario comprehensively. Our work suggests that LLMs have enormous potential for improving multilingual alignment efficiently with great language and task generalization.

cs.CL↗

Real-time Neuron Segmentation for Voltage Imaging

In voltage imaging, where the membrane potentials of individual neurons are recorded at from hundreds to thousand frames per second using fluorescence microscopy, data processing presents a challenge. Even a fraction of a minute of recording with a limited image size yields gigabytes of video data consisting of tens of thousands of frames, which can be time-consuming to process. Moreover, millisecond-level short exposures lead to noisy video frames, obscuring neuron footprints especially in deep-brain samples where noisy signals are buried in background fluorescence. To address this challenge, we propose a fast neuron segmentation method able to detect multiple, potentially overlapping, spiking neurons from noisy video frames, and implement a data processing pipeline incorporating the proposed segmentation method along with GPU-accelerated motion correction. By testing on existing datasets as well as on new datasets we introduce, we show that our pipeline extracts neuron footprints that agree well with human annotation even from cluttered datasets, and demonstrate real-time processing of voltage imaging data on a single desktop computer for the first time.

eess.IV↗

High-Accuracy Prediction of Metal-Insulator-Metal Metasurface with Deep Learning

Deep learning prediction of electromagnetic software calculation results has been a widely discussed issue in recent years. But the prediction accuracy was still one of the challenges to be solved. In this work, we proposed that the ResNets-10 model was used for predicting plasmonic metasurface S11 parameters. The two-stage training was performed by the k-fold cross-validation and small learning rate. After the training was completed, the prediction loss for aluminum, gold, and silver metal-insulator-metal metasurfaces was -48.45, -46.47, and -35.54, respectively. Due to the ultralow error value, the proposed network can replace the traditional electromagnetic computing method for calculation within a certain structural range. Besides, this network can finish the training process less than 1,100 epochs. This means that the network training process can effectively lower the design process time. The ResNets-10 model we proposed can also be used to design meta-diffractive devices and biosensors, thereby reducing the time required for the calculation process. The ultralow error of the network indicates that this work contributes to the development of future artificial intelligence electromagnetic computing software.

cs.LG↗

Two Kinds Hybrid Power Mean Involving Two-Term Exponential Sums and Dedekind Sums

The main purpose of this article is using the analytic mathods and the quadratic residual transformation technique, and properties of Dedekind sums to study the calculating problem of two kinds hybrid power mean involving the two-term exponential sums and Dedekind sums, and give two asymptotic formulas for it. This work is a generalization for existing conclusions.

math.NT↗

ESCL: Equivariant Self-Contrastive Learning for Sentence Representations

Previous contrastive learning methods for sentence representations often focus on insensitive transformations to produce positive pairs, but neglect the role of sensitive transformations that are harmful to semantic representations. Therefore, we propose an Equivariant Self-Contrastive Learning (ESCL) method to make full use of sensitive transformations, which encourages the learned representations to be sensitive to certain types of transformations with an additional equivariant learning task. Meanwhile, in order to improve practicability and generality, ESCL simplifies the implementations of traditional equivariant contrastive methods to share model parameters from the perspective of multi-task learning. We evaluate our ESCL on semantic textual similarity tasks. The proposed method achieves better results while using fewer learning parameters compared to previous methods.

cs.CL↗

Selective surface modification and layer thinning of MoS2 via ultraviolet light irradiation in ionic solution

The electrical and optoelectronic properties of transition-metal dichalcogenides (TMDs), such as MoS2, are highly dependent on carrier doping and layer thickness. The ability to selectively control these two critical characteristics is of great importance to develop TMD-based multifunctional device applications, which remains challenging. Here, we report a strategy for controllable surface modification and layer thinning of MoS2 via ultraviolet (UV) light irradiation in a silver ionic solution environment. The results show that by adjusting UV irradiation time, nanostructured silver ultrathin films (~2.9 nm) are uniformly deposited on monolayer MoS2 and can lead to controllable p-type doping effect, while the thickness of MoS2 from few-layer to bulk crystals could be thinned down to the atomic monolayer limit. Both silver nanostructure deposition and layer thinning process have been evidenced to initiate from the edges of MoS2, and independent of the edge type, thus revealing a unique UV light-assisted defect-induced surface modification and layer thinning mechanism. Overall, this study provides a new methodology for selective control of doping and layer thickness in TMDs, paving the way for developing novel 2D nanoelectronics and integrated optoelectronics.

cond-mat.mtrl-sci↗

Response of open two-band systems to a momentum-carrying single-mode quantized field

As a new quantum state, topological insulators have become the focus of condensed matter and material science. The open system research of topological insulators has aroused the interest of many researchers. Recently, many aspects, especially experimental aspects, have been developed rapidly, such as prediction and discovery of many novel quantum effects and applications of topological properties of new materials, but the theoretical research is slightly tough. In this paper, we study the response of topological insulator driven by momentum-carrying single-mode field. We solve the ground state of the system after the addition of a single mode light field with adjustable photon momentum. Specifically, We show that from the analytical solution of hall conductance compared with the closed system, there is an extra correction term, and hall conductance can no longer be expressed in terms of the chern number or the weighted sum of the chern number. Furthermore, the topological properties are analyzed and discussed through the results of different instance with their illustration. Such as, the phase transition point of topological phase is robust to the environment, and the system still has topological phase transition. It is expected to be realized or controlled by experiments, and our observations may contribute to its application and extension in condensed matter physics and quantum statistical physics.

cond-mat.mes-hall↗

AI Based Digital Twin Model for Cattle Caring

In this paper, we developed innovative digital twins of cattle status that are powered by artificial intelligence (AI). The work was built on a farm IoT system that remotely monitors and tracks the state of cattle. A digital twin model of cattle health based on Deep Learning (DL) was generated using the sensor data acquired from the farm IoT system. The health and physiological cycle of cattle can be monitored in real time, and the state of the next physiological cycle of cattle can be anticipated using this model. The basis of this work is the vast amount of data which is required to validate the legitimacy of the digital twins model. In terms of behavioural state, it was found that the cattle treated with a combination of topical anaesthetic and meloxicam exhibits the least pain reaction. The digital twins model developed in this work can be used to monitor the health of cattle

cs.AI↗

Automated Performance Tuning for Highly-Configurable Software Systems

Performance is an important non-functional aspect of the software requirement. Modern software systems are highly-configurable and misconfigurations may easily cause performance issues. A software system that suffers performance issues may exhibit low program throughput and long response time. However, the sheer size of the configuration space makes it challenging for administrators to manually select and adjust the configuration options to achieve better performance. In this paper, we propose ConfRL, an approach to tune software performance automatically. The key idea of ConfRL is to use reinforcement learning to explore the configuration space by a trial-and-error approach and to use the feedback received from the environment to tune configuration option values to achieve better performance. To reduce the cost of reinforcement learning, ConfRL employs sampling, clustering, and dynamic state reduction techniques to keep states in a large configuration space manageable. Our evaluation of four real-world highly-configurable server programs shows that ConfRL can efficiently and effectively guide software systems to achieve higher long-term performance.

cs.SE↗

Plasmonic tweezers based on connected nanoring apertures

The manipulation of microparticles using optical forces has led to many applications in the life and physical sciences. To extend optical trapping towards the nano-regime, in this work we demonstrate trapping of single nanoparticles in arrays of plasmonic coaxial nano-apertures with various inner disk configurations and theoretically estimate the associated forces. A high normalised experimental trap stiffness of 3.50fN/nm/mW for 20nm polystyrene particles is observed for an optimum design of 149nm for the nanodisk diameter at a trapping wavelength of 980nm. Theoretical simulations are used to interpret the enhancement of the observed trap stiffness. A quick particle trapping time of less than 8sec is obtained at a concentration of 14$\times$10$^{11}$ particles/ml with low incident laser intensity of 0.59mW/$μ$m$^{2}$. This good trapping performance with fast delivery of nanoparticles to multiple trapping sites emerges from a combination of the enhanced electromagnetic near-field and spatial temperature increase. This work has applications in nanoparticle delivery and trapping with high accuracy, and bridges the gap between optical manipulation and nanofluidics.

physics.optics↗

Enhanced photon blockade in an optomechanical system with parametric amplification

We propose a scheme to enhance the single- and two-photon blockade effect significantly in a standard optomechanical system (OMS) via optical parametric amplification (OPA). The scheme does not rely on the strong single-photon optomechanical coupling and can eliminate the disadvantages of suppressing multi-photon excitation incompletely. Through analyzing the single-photon blockade (1PB) mechanism and optimizing the system parameters, we obtain a perfect 1PB with a high occupancy probability of single-photon excitation, which means that a high quality and efficient single-photon source can be generated. Moreover, we find that not only the two-photon blockade (2PB) effect is significantly enhanced but also the region of 2PB occurring is widened when the OPA exists, where we also derive the optimal parameter condition to maximize the two-photon emission and the higher photon excitations are intensely suppressed at the same time.

quant-ph↗

Automatic Business Process Structure Discovery using Ordered Neurons LSTM: A Preliminary Study

Automatic process discovery from textual process documentations is highly desirable to reduce time and cost of Business Process Management (BPM) implementation in organizations. However, existing automatic process discovery approaches mainly focus on identifying activities out of the documentations. Deriving the structural relationships between activities, which is important in the whole process discovery scope, is still a challenge. In fact, a business process has latent semantic hierarchical structure which defines different levels of detail to reflect the complex business logic. Recent findings in neural machine learning area show that the meaningful linguistic structure can be induced by joint language modeling and structure learning. Inspired by these findings, we propose to retrieve the latent hierarchical structure present in the textual business process documents by building a neural network that leverages a novel recurrent architecture, Ordered Neurons LSTM (ON-LSTM), with process-level language model objective. We tested the proposed approach on data set of Process Description Documents (PDD) from our practical Robotic Process Automation (RPA) projects. Preliminary experiments showed promising results.

cs.CL↗