SearcharxivSearch

arXiv subjects

Chunlan Ma

Publications and source records attributed to Chunlan Ma.

At least 19 recordsLinked to original sources

Modeling Bond-Dependent Kitaev-like interaction in 2D Edge-Sharing Tetrahedral Magnets: FeX (X=Te, Se)

Bond-dependent magnetic interactions, exemplified by the Kitaev model, are known to arise from the interplay between spin-orbit coupling (SOC) and specific coordination geometries, but have so far been almost exclusively identified in edge-sharing octahedral systems. Whether such interactions persist in edge-sharing tetrahedral environments, characteristic of the parent compounds of iron-based superconductors, remains an open question. Here, we construct a Kitaev-like model for monolayer FeTe and FeSe and demonstrate the presence of a previously unrecognized bond-dependent Ising-type interaction, induced jointly by chalcogen-mediated SOC and the tetrahedral crystal-field geometry. A microscopic spin model for these bond-dependent interactions is derived via strong-coupling perturbation theory, and the strengths of the individual exchange terms are extracted by partitioning the magnetic anisotropy energy calculated using density functional theory across various collinear magnetic orders. We reveal that the Kitaev-like interaction dominates the magnetic anisotropy in FeTe, whereas in FeSe, it strongly competes with a single-ion anisotropy of opposite sign. The resulting noncollinear local anisotropy axes generate intrinsic single-site spin frustration, providing a microscopic mechanism for magnetic disorder that transcends isotropic exchange models. Our results establish edge-sharing tetrahedral magnets as a new platform for bond-dependent interactions and extend the scope of Kitaev physics beyond octahedral coordination.

physics.comp-ph

Ah-SCDFT:A general approach for superconductivity with an-harmonic corrections

First-principles studies of superconductivity often neglect anharmonic effects (AHE), despite their crucial role in achieving quantitative accuracy in many materials. To bridge this gap, we introduce a general computational approach, termed anharmonic superconducting density functional theory (ah-SCDFT) which systematically incorporates anharmonic corrections into standard SCDFT. This approach allows for high-fidelity predictions of superconducting properties with only a modest increase in computational cost for a limited number of superconducting calculation convergence steps. We demonstrate the effectiveness and reliability of ah-SCDFT by applying it to the prototypical superconductor MgB2, accurately reproducing its superconducting behavior under both ambient conditions and applied pressure in excellent agreement with experiment. Our results establish ah-SCDFT as a powerful, efficient, and broadly applicable approach for quantitatively reliable studies of superconductivity and a promising tool for the prediction of new superconducting materials.

cond-mat.supr-con

Nature of point defects in bulk hexagonal diamond

Hexagonal diamond (HD), an exotic carbon allotrope recently synthesized in bulk form, exhibits superior mechanical properties compared to cubic diamond (CD) and holds promise for advanced industrial and quantum applications. Using first-principles calcu-lations, we systematically investigate intrinsic defects, extrinsic dopants, and defect complexes in HD. Our study shows that VC dominates intrinsic conductivity, while Ci is unstable. Among extrinsic dopants, boron acts as a benign acceptor enhancing p-type conductivity, whereas nitrogen and phosphorus serve as effective donors for n-type conductivity. Group II and Group IV dopants, however, introduce high formation energies or neutral charge states with limited impact. Furthermore, VC, MgC and XV defect com-plexes display multiple spin and charge states within the HD band gap, highlighting their potential as color centers for hosting qubits. These results not only clarify the defect physics of HD but also demonstrate its broader implications for conductivity engineering and quantum technologies.

cond-mat.mtrl-sci

CulturALL: Benchmarking Multilingual and Multicultural Competence of LLMs on Grounded Tasks

Large language models (LLMs) are now deployed worldwide, inspiring a surge of benchmarks that measure their multilingual and multicultural abilities. However, these benchmarks prioritize generic language understanding or superficial cultural trivia, leaving the evaluation of grounded tasks -- where models must reason within real-world, context-rich scenarios -- largely unaddressed. To fill this gap, we present CulturALL, a comprehensive and challenging benchmark to assess LLMs' multilingual and multicultural competence on grounded tasks. CulturALL is built via a human--AI collaborative framework: expert annotators ensure appropriate difficulty and factual accuracy, while LLMs lighten the manual workload. By incorporating diverse sources, CulturALL ensures comprehensive scenario coverage. Each item is carefully designed to present a high level of difficulty, making CulturALL challenging. CulturALL contains 2,610 samples in 14 languages from 51 regions, distributed across 16 topics to capture the full breadth of grounded tasks. Experiments show that the best LLM achieves 44.48% accuracy on CulturALL, underscoring substantial room for improvement.

cs.CL

Stabilization of zigzag order in NiPS$_3$ via positive biquadratic interaction

Despite extensive research, the precise spin Hamiltonian of the van der Waals antiferromagnet NiPS$_3$ -- which hosts a zigzag-ordered ground state -- remains debated. While consensus has emerged on ferromagnetic nearest-neighbor ($J_1$) and antiferromagnetic third-nearest-neighbor ($J_3$) Heisenberg interactions, recent studies suggest a biquadratic ($B$) exchange term may also play a role, though its estimated magnitude varies widely. To address this controversy, we perform density functional theory calculations and extract a positive biquadratic interaction with $B/J_3 \approx 0.44$. Within the minimal $J_1$-$J_3$-$B$ model, we show that these parameters naturally stabilize zigzag ordering using minimally augmented spin-wave theory. Density-matrix renormalization group calculations further validate our extracted parameters as a reasonable description of the ground state. Although fully resolving the spin Hamiltonian of NiPS$_3$ requires further investigation, our findings provide new insights into its biquadratic interaction.

cond-mat.str-el

Effect of superconductivity by Nb and V substitution in kagome CaPd5

Materials featuring kagome lattices have attracted significant research interest due to their unique geometric frustration, which gives rise to rich physical phenomena such as non-trivial topology, spin fluctuations, and superconductivity. In this work, using CaPd5 as the prototype structure, we discover and systematically investigate a new class of kagome superconductors, CaMxPd5-x (M = Nb and V) alloys. First-principles calculations confirm that these compounds are non-magnetic metals, among which four are dynamically stable: CaNb5, CaV5, CaNb2Pd3, and CaV2Pd3. CaNb5 is identified as a strong electron-phonon coupling (EPC) superconductor with the highest superconducting transition temperature (Tc) of 10.1 K, which can be further increased to 12.8 K under external pressure. In contrast, CaV5, CaNb2Pd3, and CaV2Pd3 exhibit weaker EPC and correspondingly lower Tc values. Furthermore, by applying the method of symmetry indicators, we systematically classify the topological and nodal characteristics of CaNb5, providing valuable insights for determining its superconducting pairing symmetry. Our findings demonstrate that Nb and V substitution in kagome CaPd5 provides an effective route for designing a new type of kagome superconductor with relatively high Tc. This study also offers new perspectives on topological superconductivity in kagome systems and establishes a useful guideline for discovering other superconducting materials with unique properties.

cond-mat.supr-con

Extraordinary cation-replace-cation antisite defect predominate in Bi2SeO5

As a newly identified single-crystalline van der Waals dielectric with a high dielectric constant, Bi2SeO5 plays a pivotal role in advancing 2D electronic devices. In this work, we systematically investigate the defect properties of Bi2SeO5 using first-principles calculations based on a hybrid functional. Although Bi2SeO5 is a chemically ternary compound, each constituent element occupies several crystallographically nonequivalent sites, rendering its defect chemistry highly complex. Due to the anomalous +4 cationic valence state of Se, the defect formation energies of same main group anion antisite defects (SeO and OSe) are prohibitively high, and their concentrations can therefore be neglected. In contrast, the extraordinary cation-cation antisite defects BiSe and SeBi emerge as the dominant defects. The pronounced variability in the formation energies of the six types of VO defects demonstrates that identical defect types located on nonequivalent atomic sites can exhibit markedly different properties. Under O-rich and Se/Bi-poor conditions, Bi2SeO5 shows relatively robust p-type behavior. Conversely, under O-poor and Se/Bi-rich conditions, or at intermediate O, Se, and Bi partial pressures, Bi2SeO5 behaves as an intrinsic semiconductor or displays very weak n-type conductivity due to strong donor-acceptor compensation. This study provides theoretical insights to guide the design and development of high-performance Bi2SeO5-based electronic devices.

cond-mat.mtrl-sci

Symmetry-Constrained Anomalous Transport in the Altermagnetic Material CuX$_2$ (X=F,Cl)

Recently discovered, altermagnetism represents a third class of collinear magnets. These materials exhibit zero net magnetization, similar to antiferromagnets, but display anomalous transport properties resembling those of ferromagnets. Altermagnetic materials manifest various anomalous electronic transport phenomena, including the anomalous Hall effect, anomalous Nernst effect, and anomalous thermal Hall effect. Additionally, they exhibit magneto-optical Kerr and Faraday effects, previously considered exclusive to ferromagnetic materials. These anomalous transport phenomena are constrained by symmetry, as revealed by density functional theory (DFT) calculations. However, an effective model-based approach to verify these symmetry constraints remains unavailable. In this Letter, we construct a $k\cdot p$ model for $d$-wave altermagnets CuX$_2$ (X=F,Cl) using spin space group representations and apply it to calculate the anomalous Hall effect. The symmetry-imposed transport properties predicted by the model are in agreement with the DFT results, providing a foundation for further investigation into symmetry-restricted transport phenomena in altermagnetic materials.

cond-mat.mtrl-sci

MPd5 kagome superconductors studied by density functional calculations

Kagome materials, which are composed of hexagons tiled with a shared triangle, have inspired enormous interest due to their unique structures and rich physical properties; exploring superconducting material systems with new kagome structures is still an important research direction. Here, we predict a type of kagome superconductor, MPd5 (M is a group-IIA metal element), and identify that it exhibits coexistence of superconductivity and nontrivial topological properties. We uncover its phonon-mediated superconductivity by the density functional theory for superconductors, predicting the superconducting transition temperatures (Tc) of 2.64, 2.03, and 1.50 K for CaPd5, SrPd5, and BaPd5, respectively. These Tc can be effectively tuned through the application of external pressure and electron doping. The present results also demonstrate that MPd5 have topological properties; e.g., CaPd5 shows topological nontrivial intersection near the Fermi level (EF). Our results indicate that MPd5 materials can be an emerging material platform with rich exotic physics in their kagome structures, and render themselves excellent candidates for superconducting and advanced functional materials that could be utilized in topological quantum computing and information technology.

cond-mat.supr-con

Origin of Oxygen Partial Pressure-Dependent Conductivity in SrTiO3

SrTiO3 (STO) displays a broad spectrum of physical properties, including superconductivity, ferroelectricity, and photoconductivity, making it a standout semiconductor material. Despite extensive researches, the oxygen partial pressure-dependent conductivity in STO has re-mained elusive. This study leverages first-principles calculations, and systematically investigates the intrinsic defect properties of STO. The results reveal that VO, VSr, and TiSr are the dominant intrinsic defects, influencing STO's conductivity under varying O chemical potentials (oxygen partial pressures). Under O-poor condition, VO is the predominant donor, while VSr is the main acceptor. As the oxygen pressure increases, TiSr emerges as a critical donor defect under O-rich condition, significantly affecting conductivity. Additionally, the study elucidates the abnormal phenomenon where VTi, typically an acceptor, exhibits donor-like behavior due to the formation of O-trimer. This work offers a comprehensive understanding of how intrinsic defects tune the Fermi level, thereby altering STO's conductivity from metallic to n-type, and eventually to p-type across different O chemical potentials. These insights resolve the long-standing issue of oxygen partial pressure-dependent conductivity and explain the observed metallic conductivity in oxygen-deficient STO.

cond-mat.mtrl-sci

LangSAMP: Language-Script Aware Multilingual Pretraining

Recent multilingual pretrained language models (mPLMs) often avoid using language embeddings -- learnable vectors assigned to individual languages. However, this places a significant burden on token representations to encode all language-specific information, which may hinder language neutrality. To address this limitation, we propose Language-Script Aware Multilingual Pretraining (LangSAMP), a method that incorporates both language and script embeddings to enhance representation learning. Specifically, we integrate these embeddings into the output of the Transformer blocks before passing the final representations to the language modeling head for prediction. We apply LangSAMP to the continual pretraining of XLM-R on a highly multilingual corpus covering more than 500 languages. The resulting model consistently outperforms the baseline in zero-shot crosslingual transfer across diverse downstream tasks. Extensive analysis reveals that language and script embeddings capture language- and script-specific nuances, which benefits more language-neutral representations, proven by improved pairwise cosine similarity. In our case study, we also show that language and script embeddings can be used to select better source languages for crosslingual transfer. We make our code and models publicly available at https://github.com/cisnlp/LangSAMP.

cs.CL

How Transliterations Improve Crosslingual Alignment

Recent studies have shown that post-aligning multilingual pretrained language models (mPLMs) using alignment objectives on both original and transliterated data can improve crosslingual alignment. This improvement further leads to better crosslingual transfer performance. However, it remains unclear how and why a better crosslingual alignment is achieved, as this technique only involves transliterations, and does not use any parallel data. This paper attempts to explicitly evaluate the crosslingual alignment and identify the key elements in transliteration-based approaches that contribute to better performance. For this, we train multiple models under varying setups for two pairs of related languages: (1) Polish and Ukrainian and (2) Hindi and Urdu. To assess alignment, we define four types of similarities based on sentence representations. Our experimental results show that adding transliterations alone improves the overall similarities, even for random sentence pairs. With the help of auxiliary transliteration-based alignment objectives, especially the contrastive objective, the model learns to distinguish matched from random pairs, leading to better crosslingual alignment. However, we also show that better alignment does not always yield better downstream performance, suggesting that further research is needed to clarify the connection between alignment and performance. The code implementation is based on \url{https://github.com/cisnlp/Transliteration-PPA}.

cs.CL

Exploring the Role of Transliteration in In-Context Learning for Low-resource Languages Written in Non-Latin Scripts

Decoder-only large language models (LLMs) excel in high-resource languages across various tasks through few-shot or even zero-shot in-context learning (ICL). However, their performance often does not transfer well to low-resource languages, especially those written in non-Latin scripts. Inspired by recent work that leverages transliteration in encoder-only models, we investigate whether transliteration is also effective in improving LLMs' performance for low-resource languages written in non-Latin scripts. To this end, we propose three prompt templates, where the target-language text is represented in (1) its original script, (2) Latin script, or (3) both. We apply these methods to several representative LLMs of different sizes on various tasks including text classification and sequential labeling. Our findings show that the effectiveness of transliteration varies by task type and model size. For instance, all models benefit from transliterations for sequential labeling (with increases of up to 25%).

cs.CL

TransMI: A Framework to Create Strong Baselines from Multilingual Pretrained Language Models for Transliterated Data

Transliterating related languages that use different scripts into a common script is effective for improving crosslingual transfer in downstream tasks. However, this methodology often makes pretraining a model from scratch unavoidable, as transliteration brings about new subwords not covered in existing multilingual pretrained language models (mPLMs). This is undesirable because it requires a large computation budget. A more promising way is to make full use of available mPLMs. To this end, this paper proposes a simple but effective framework: Transliterate-Merge-Initialize (TransMI). TransMI can create strong baselines for data that is transliterated into a common script by exploiting an existing mPLM and its tokenizer without any training. TransMI has three stages: (a) transliterate the vocabulary of an mPLM into a common script; (b) merge the new vocabulary with the original vocabulary; and (c) initialize the embeddings of the new subwords. We apply TransMI to three strong recent mPLMs. Our experiments demonstrate that TransMI not only preserves the mPLM's ability to handle non-transliterated data, but also enables it to effectively process transliterated data, thereby facilitating crosslingual transfer across scripts. The results show consistent improvements of 3% to 34% for different mPLMs and tasks. We make our code and models publicly available at \url{https://github.com/cisnlp/TransMI}.

cs.CL

TransliCo: A Contrastive Learning Framework to Address the Script Barrier in Multilingual Pretrained Language Models

The world's more than 7000 languages are written in at least 293 scripts. Due to various reasons, many closely related languages use different scripts, which poses a difficulty for multilingual pretrained language models (mPLMs) in learning crosslingual knowledge through lexical overlap. As a consequence, mPLMs are faced with a script barrier: representations from different scripts are located in different subspaces, which can result in crosslingual transfer involving languages of different scripts performing suboptimally. To address this problem, we propose TransliCo, a framework that optimizes the Transliteration Contrastive Modeling (TCM) objective to fine-tune an mPLM by contrasting sentences in its training data and their transliterations in a unified script (in our case Latin), which enhances uniformity in the representation space for different scripts. Using Glot500-m, an mPLM pretrained on over 500 languages, as our source model, we fine-tune it on a small portion (5%) of its training data, and refer to the resulting model as Furina. We show that Furina not only better aligns representations from distinct scripts but also outperforms the original Glot500-m on various zero-shot crosslingual transfer tasks. Additionally, we achieve consistent improvement in a case study on the Indic group where the languages exhibit areal features but use different scripts. We make our code and models publicly available.

cs.CL

MoSECroT: Model Stitching with Static Word Embeddings for Crosslingual Zero-shot Transfer

Transformer-based pre-trained language models (PLMs) have achieved remarkable performance in various natural language processing (NLP) tasks. However, pre-training such models can take considerable resources that are almost only available to high-resource languages. On the contrary, static word embeddings are easier to train in terms of computing resources and the amount of data required. In this paper, we introduce MoSECroT Model Stitching with Static Word Embeddings for Crosslingual Zero-shot Transfer), a novel and challenging task that is especially relevant to low-resource languages for which static word embeddings are available. To tackle the task, we present the first framework that leverages relative representations to construct a common space for the embeddings of a source language PLM and the static word embeddings of a target language. In this way, we can train the PLM on source-language training data and perform zero-shot transfer to the target language by simply swapping the embedding layer. However, through extensive experiments on two classification datasets, we show that although our proposed framework is competitive with weak baselines when addressing MoSECroT, it fails to achieve competitive results compared with some strong baselines. In this paper, we attempt to explain this negative result and provide several thoughts on possible improvement.

cs.CL

Evidence of Kitaev interaction in the monolayer 1T-CrTe$_2$

The two-dimensional 1T-CrTe$_2$ has been an attractive room-temperature van der Waals magnet which has a potential application in spintronic devices. Although it was recognized as a ferromagnetism in the past, the monolayer 1T-CrTe$_2$ was recently found to exhibit zigzag antiferromagnetism with the easy axis oriented at $70^\circ$ to the perpendicular direction of the plane. Therefore, the origin of the intricate anisotropic magnetic behavior therein is well worthy of thorough exploration. Here, by applying density functional theory with spin spiral method, we demonstrate that the Kitaev interaction, together with the single-ion anisotropy and other off-diagonal exchanges, is amenable to explain the magnetic orientation in the metallic 1T-CrTe$_2$. Moreover, the Ruderman-Kittle-Kasuya-Yosida interaction can also be extracted from the dispersion calculations, which explains the metallic behavior of 1T-CrTe$_2$. Our results demonstrate that 1T-CrTe$_2$ is potentially a rare metallic Kitaev material.

cond-mat.str-el

Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages

The NLP community has mainly focused on scaling Large Language Models (LLMs) vertically, i.e., making them better for about 100 languages. We instead scale LLMs horizontally: we create, through continued pretraining, Glot500-m, an LLM that covers 511 predominantly low-resource languages. An important part of this effort is to collect and clean Glot500-c, a corpus that covers these 511 languages and allows us to train Glot500-m. We evaluate Glot500-m on five diverse tasks across these languages. We observe large improvements for both high-resource and low-resource languages compared to an XLM-R baseline. Our analysis shows that no single factor explains the quality of multilingual LLM representations. Rather, a combination of factors determines quality including corpus size, script, "help" from related languages and the total capacity of the model. Our work addresses an important goal of NLP research: we should not limit NLP to a small fraction of the world's languages and instead strive to support as many languages as possible to bring the benefits of NLP technology to all languages and cultures. Code, data and models are available at https://github.com/cisnlp/Glot500.

cs.CL