SearcharxivSearch

arXiv subjects

Yao Yang

Publications and source records attributed to Yao Yang.

At least 19 recordsLinked to original sources

Probing three-dimensional structures of complex colloidal quantum dots at the single-atomic level

Colloidal quantum dots (QDs) are promising optoelectronic materials due to their size-tunable properties, yet their three-dimensional (3D) quantum confinement makes electronic states highly sensitive to structural and chemical heterogeneity, which critically impacts their optoelectronic performance. Accurately resolving the 3D atomic structure with sub-angstrom precision is thus essential for rational design. Here, we applied atomic electron tomography (AET) to determine, for the first time, the 3D atomic structure of complex core/shell QDs, resolving over 14,000 atoms per particle. Our reconstructions reveal surface morphology, eccentric cores, and nearly atomically abrupt heterovalent interfaces and identify anisotropic shell growth directed by twin boundaries. Utilizing an AET-derived atomic structure, we performed large-scale quantum mechanical calculations to uncover an orientation-dependent strain accommodation mechanism where the heterogeneous strain is compensated at interfaces and twin boundaries. Furthermore, our results reveal strain-induced localized states near the band edge, which contribute to the key features of the experimental ensemble absorption spectrum. This work sets a new benchmark for atomic-level characterization, establishing a powerful framework for the rational design of next-generation nanomaterials.

cond-mat.mtrl-sci

QAROO: AI-Driven Online Task Offloading for Energy-Efficient and Sustainable MEC Networks

With the rapid advancement of artificial intelligence (AI) and intelligent science, intelligent edge computing has been widely adopted. However, the limitations of traditional methods, such as poor adaptability and the slow convergence of heuristic algorithms, are becoming increasingly evident. To enable sustainable and resource-efficient edge applications, this paper proposes an online task offloading framework for wireless powered mobile edge computing (MEC) networks, called Quantum Attention-based Reinforcement learning for Online Offloading (QAROO). The system employs a binary offloading strategy with the aim of co-optimizing computing and energy resources in dynamic channel environments. In response to the issues of poor adaptability in traditional approaches and the slow convergence of heuristic algorithms, the framework integrates quantum neural networks and attention mechanisms, introducing three key improvements: using recurrent neural networks to enhance temporal modeling capability, proposing an uncertainty-guided quantization method to improve exploration efficiency, and incorporating attention mechanisms into quantum networks to strengthen feature representation. Experiments demonstrate that the proposed method outperforms comparative schemes in terms of normalized computation speed and processing time, offering an efficient and stable solution for online task offloading in large-scale Internet of Things (IoT) dynamic environments.

cs.AI

A Transformer-based Model for Rapid Microstructure Inference from Four-Dimensional Scanning Transmission Electron Microscopy Data

Properties of crystalline materials are closely linked to microstructure arising from the spatial arrangement, orientation, and phase of nanocrystals. Rapid characterization of crystalline microstructure can accelerate the identification of these links and the development of materials with desired properties. Here, we combine a machine learning framework with four-dimensional scanning transmission electron microscopy (4D-STEM) to enable fast inference of crystalline microstructure over large fields of view. The framework employs a transformer-based architecture to predict crystallographic orientations and phases from 4D-STEM diffraction patterns, yielding spatially resolved maps of microstructural features at the nanoscale. With this framework, crystallographic orientations are inferred up to two orders of magnitude faster than widely used correlative template-matching approaches. This capability enables high-throughput characterization of complex crystalline materials and facilitates the establishment of structure-property relationships central to materials design and optimization.

cond-mat.mtrl-sci

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models

Spatial reasoning is a core component of human cognition, enabling individuals to perceive, comprehend, and interact with the physical world. It relies on a nuanced understanding of spatial structures and inter-object relationships, serving as the foundation for complex reasoning and decision-making. To investigate whether current vision-language models (VLMs) exhibit similar capability, we introduce Jigsaw-Puzzles, a novel benchmark consisting of 1,100 carefully curated real-world images with high spatial complexity. Based on this dataset, we design five tasks to rigorously evaluate VLMs' spatial perception, structural understanding, and reasoning capabilities, while deliberately minimizing reliance on domain-specific knowledge to better isolate and assess the general spatial reasoning capability. We conduct a comprehensive evaluation across 24 state-of-the-art VLMs. The results show that even the strongest model, Gemini-2.5-Pro, achieves only 77.14% overall accuracy and performs particularly poorly on the Order Generation task, with only 30.00% accuracy, far below the performance exceeding 90% achieved by human participants. This persistent gap underscores the need for continued progress, positioning Jigsaw-Puzzles as a challenging and diagnostic benchmark for advancing spatial reasoning research in VLMs. Our project page is at https://zesen01.github.io/jigsaw-puzzles.

cs.AI

LLM-based Automated Theorem Proving Hinges on Scalable Synthetic Data Generation

Recent advancements in large language models (LLMs) have sparked considerable interest in automated theorem proving and a prominent line of research integrates stepwise LLM-based provers into tree search. In this paper, we introduce a novel proof-state exploration approach for training data synthesis, designed to produce diverse tactics across a wide range of intermediate proof states, thereby facilitating effective one-shot fine-tuning of LLM as the policy model. We also propose an adaptive beam size strategy, which effectively takes advantage of our data synthesis method and achieves a trade-off between exploration and exploitation during tree search. Evaluations on the MiniF2F and ProofNet benchmarks demonstrate that our method outperforms strong baselines under the stringent Pass@1 metric, attaining an average pass rate of $60.74\%$ on MiniF2F and $21.18\%$ on ProofNet. These results underscore the impact of large-scale synthetic data in advancing automated theorem proving.

cs.AI

Crystal nucleation and growth in high-entropy alloys revealed by atomic electron tomography

High-entropy alloys (HEAs) balance mixing entropy and intermetallic phase formation enthalpy, creating a vast compositional space for structural and functional materials (1-6). They exhibit exceptional strength-ductility trade-offs in metallurgy (4-10) and near-continuum adsorbate binding energies in catalysis (11-16). A deep understanding of crystal nucleation and growth in HEAs is essential for controlling their formation and optimizing their structural and functional properties. However, atomic-scale nucleation in HEAs challenges traditional theories based on one or two principal elements (17-23). The intricate interplay of structural and chemical orders among multiple principal elements further obscures our understanding of nucleation pathways (5,24-27). Due to the lack of direct three-dimensional (3D) atomic-scale observations, previous studies have relied on simulations and indirect measurements (28-32), leaving HEA nucleation and growth fundamentally elusive. Here, we advance atomic electron tomography (33,34) to resolve the 3D atomic structure and chemical composition of 7,662 HEA and 498 medium-entropy alloy nuclei at different nucleation stages. We observe local structural order that decreases from core to boundary, correlating with local chemical order. As nuclei grow, structural order improves. At later stages, most nuclei coalesce without misorientation, while some form coherent twin boundaries. To explain these experimental observations, we propose the gradient nucleation pathways model, in which the nucleation energy barrier progressively increases through multiple evolving intermediate states. We expect these findings to not only provide fundamental insights into crystal nucleation and growth in HEAs, but also offer a general framework for understanding nucleation mechanisms in other materials.

cond-mat.mtrl-sci

Giant spin shift current in two-dimensional altermagnetic multiferroics VOX$\mathrm{_2}$

Altermagnets represent a novel class of magnetic materials that integrate the advantages of both ferromagnets and antiferromagnets, providing a rich platform for exploring the physical properties of multiferroic materials.This work demonstrates that $\mathrm{VOX_2}$ monolayers ($\mathrm{X = Cl, Br, I}$) are two-dimensional ferroelectric altermagnets, as confirmed by symmetry analysis and first-principles calculations. $\mathrm{VOI_2}$ monolayer exhibits a strong magnetoelectric coupling coefficient ($\alpha_S \approx 1.208 \times 10^{-6}~\mathrm{s/m}$), with spin splitting in the electronic band structure tunable by both electric and magnetic fields. Additionally, the absence of inversion symmetry in noncentrosymmetric crystals enables significant nonlinear optical effects, such as shift current (SC). The $x$-direction component of SC exhibits a ferroicity-driven switching behavior. Moreover, the $\sigma^{yyy}$ component exhibits an exceptionally large spin SC of $330.072~\mathrm{\mu A/V^2}$. These findings highlight the intricate interplay between magnetism and ferroelectricity, offering versatile tunability of electronic and optical properties. $\mathrm{VOX_2}$ monolayers provide a promising platform for advancing two-dimensional multiferroics, paving the way for energy-efficient memory devices, nonlinear optical applications and opto-spintronics.

cond-mat.mtrl-sci

From Rational Answers to Emotional Resonance: The Role of Controllable Emotion Generation in Language Models

Purpose: Emotion is a fundamental component of human communication, shaping understanding, trust, and engagement across domains such as education, healthcare, and mental health. While large language models (LLMs) exhibit strong reasoning and knowledge generation capabilities, they still struggle to express emotions in a consistent, controllable, and contextually appropriate manner. This limitation restricts their potential for authentic human-AI interaction. Methods: We propose a controllable emotion generation framework based on Emotion Vectors (EVs) - latent representations derived from internal activation shifts between neutral and emotion-conditioned responses. By injecting these vectors into the hidden states of pretrained LLMs during inference, our method enables fine-grained, continuous modulation of emotional tone without any additional training or architectural modification. We further provide theoretical analysis proving that EV steering enhances emotional expressivity while maintaining semantic fidelity and linguistic fluency. Results: Extensive experiments across multiple LLM families show that the proposed approach achieves consistent emotional alignment, stable topic adherence, and controllable affect intensity. Compared with existing prompt-based and fine-tuning-based baselines, our method demonstrates superior flexibility and generalizability. Conclusion: Emotion Vector (EV) steering provides an efficient and interpretable means of bridging rational reasoning and affective understanding in large language models, offering a promising direction for building emotionally resonant AI systems capable of more natural human-machine interaction.

cs.CL

Dual-Label Learning With Irregularly Present Labels

In multi-task learning, labels are often missing irregularly across samples, which can be fully labeled, partially labeled or unlabeled. The irregular label presence often appears in scientific studies due to experimental limitations. It triggers a demand for a new training and inference mechanism that could accommodate irregularly present labels and maximize their utility. This work focuses on the two-label learning task and proposes a novel training and inference framework, Dual-Label Learning (DLL). The DLL framework formulates the problem into a dual-function system, in which the two functions should simultaneously satisfy standard supervision, structural duality and probabilistic duality. DLL features a dual-tower model architecture that allows for explicit information exchange between labels, aimed at maximizing the utility of partially available labels. During training, missing labels are imputed as part of the forward propagation process, while during inference, labels are predicted jointly as unknowns of a bivariate system of equations. Our theoretical analysis guarantees the feasibility of DLL, and extensive experiments are conducted to verify that by explicitly modeling label correlation and maximizing label utility, our method makes consistently better prediction than baseline approaches by up to 9.6% gain in F1-score or 10.2% reduction in MAPE. Remarkably, DLL maintains robust performance at a label missing rate of up to 60%, achieving even better results than baseline approaches at lower missing rates down to only 10%.

cs.LG

Executing Arithmetic: Fine-Tuning Large Language Models as Turing Machines

Large Language Models (LLMs) have demonstrated remarkable capabilities across a wide range of natural language processing and reasoning tasks. However, their performance in the foundational domain of arithmetic remains unsatisfactory. When dealing with arithmetic tasks, LLMs often memorize specific examples rather than learning the underlying computational logic, limiting their ability to generalize to new problems. In this paper, we propose a Composable Arithmetic Execution Framework (CAEF) that enables LLMs to learn to execute step-by-step computations by emulating Turing Machines, thereby gaining a genuine understanding of computational logic. Moreover, the proposed framework is highly scalable, allowing composing learned operators to significantly reduce the difficulty of learning complex operators. In our evaluation, CAEF achieves nearly 100% accuracy across seven common mathematical operations on the LLaMA 3.1-8B model, effectively supporting computations involving operands with up to 100 digits, a level where GPT-4o falls short noticeably in some settings.

cs.AI

NPAT Null-Space Projected Adversarial Training Towards Zero Deterioration

To mitigate the susceptibility of neural networks to adversarial attacks, adversarial training has emerged as a prevalent and effective defense strategy. Intrinsically, this countermeasure incurs a trade-off, as it sacrifices the model's accuracy in processing normal samples. To reconcile the trade-off, we pioneer the incorporation of null-space projection into adversarial training and propose two innovative Null-space Projection based Adversarial Training(NPAT) algorithms tackling sample generation and gradient optimization, named Null-space Projected Data Augmentation (NPDA) and Null-space Projected Gradient Descent (NPGD), to search for an overarching optimal solutions, which enhance robustness with almost zero deterioration in generalization performance. Adversarial samples and perturbations are constrained within the null-space of the decision boundary utilizing a closed-form null-space projector, effectively mitigating threat of attack stemming from unreliable features. Subsequently, we conducted experiments on the CIFAR10 and SVHN datasets and reveal that our methodology can seamlessly combine with adversarial training methods and obtain comparable robustness while keeping generalization close to a high-accuracy model.

cs.LG

Preserving Surface Strain in Nanocatalysts via Morphology Control

Engineering strain critically affects the properties of materials and has extensive applications in semiconductors and quantum systems. However, the deployment of strain-engineered nanocatalysts faces challenges, particularly in maintaining highly strained nanocrystals under reaction conditions. Here, we introduce a morphology-dependent effect that stabilizes surface strain even under harsh reaction conditions. Employing four-dimensional scanning transmission electron microscopy (4D-STEM), we discovered that core-shell Au@Pd nanoparticles with sharp-edged morphologies sustain coherent heteroepitaxial interfaces with designated surface strain. This configuration inhibits dislocation due to reduced shear stress at corners, as molecular dynamics simulations indicate. Demonstrated in a Suzuki-type cross-coupling reaction, our approach achieves a fourfold increase in activity over conventional nanocatalysts, owing to the enhanced stability of surface strain. These findings contribute to advancing the development of advanced nanocatalysts and indicate broader applications for strain engineering in various fields.

cond-mat.mtrl-sci

Vertical Federated Learning Hybrid Local Pre-training

Vertical Federated Learning (VFL), which has a broad range of real-world applications, has received much attention in both academia and industry. Enterprises aspire to exploit more valuable features of the same users from diverse departments to boost their model prediction skills. VFL addresses this demand and concurrently secures individual parties from exposing their raw data. However, conventional VFL encounters a bottleneck as it only leverages aligned samples, whose size shrinks with more parties involved, resulting in data scarcity and the waste of unaligned data. To address this problem, we propose a novel VFL Hybrid Local Pre-training (VFLHLP) approach. VFLHLP first pre-trains local networks on the local data of participating parties. Then it utilizes these pre-trained networks to adjust the sub-model for the labeled party or enhance representation learning for other parties during downstream federated learning on aligned data, boosting the performance of federated models. The experimental results on real-world advertising datasets, demonstrate that our approach achieves the best performance over baseline methods by large margins. The ablation study further illustrates the contribution of each technique in VFLHLP to its overall performance.

cs.LG

Enhancing the "Immunity" of Mixture-of-Experts Networks for Adversarial Defense

Recent studies have revealed the vulnerability of Deep Neural Networks (DNNs) to adversarial examples, which can easily fool DNNs into making incorrect predictions. To mitigate this deficiency, we propose a novel adversarial defense method called "Immunity" (Innovative MoE with MUtual information \& positioN stabilITY) based on a modified Mixture-of-Experts (MoE) architecture in this work. The key enhancements to the standard MoE are two-fold: 1) integrating of Random Switch Gates (RSGs) to obtain diverse network structures via random permutation of RSG parameters at evaluation time, despite of RSGs being determined after one-time training; 2) devising innovative Mutual Information (MI)-based and Position Stability-based loss functions by capitalizing on Grad-CAM's explanatory power to increase the diversity and the causality of expert networks. Notably, our MI-based loss operates directly on the heatmaps, thereby inducing subtler negative impacts on the classification performance when compared to other losses of the same type, theoretically. Extensive evaluation validates the efficacy of the proposed approach in improving adversarial robustness against a wide range of attacks.

cs.LG

BlockEcho: Retaining Long-Range Dependencies for Imputing Block-Wise Missing Data

Block-wise missing data poses significant challenges in real-world data imputation tasks. Compared to scattered missing data, block-wise gaps exacerbate adverse effects on subsequent analytic and machine learning tasks, as the lack of local neighboring elements significantly reduces the interpolation capability and predictive power. However, this issue has not received adequate attention. Most SOTA matrix completion methods appeared less effective, primarily due to overreliance on neighboring elements for predictions. We systematically analyze the issue and propose a novel matrix completion method ``BlockEcho" for a more comprehensive solution. This method creatively integrates Matrix Factorization (MF) within Generative Adversarial Networks (GAN) to explicitly retain long-distance inter-element relationships in the original matrix. Besides, we incorporate an additional discriminator for GAN, comparing the generator's intermediate progress with pre-trained MF results to constrain high-order feature distributions. Subsequently, we evaluate BlockEcho on public datasets across three domains. Results demonstrate superior performance over both traditional and SOTA methods when imputing block-wise missing data, especially at higher missing rates. The advantage also holds for scattered missing data at high missing rates. We also contribute on the analyses in providing theoretical justification on the optimality and convergence of fusing MF and GAN for missing block data.

cs.LG

The Oxygen Reduction Pathway for Spinel Metal Oxides in Alkaline Media: An Experimentally Supported Ab Initio Study

Precious-metal-free spinel oxide electrocatalysts are promising candidates for catalyzing the oxygen reduction reaction (ORR) in alkaline fuel cells. In this theory-driven study, we use joint density-functional theory in tandem with supporting electrochemical measurements to identify a novel theoretical pathway for the ORR on cubic Co3O4 nanoparticle electrocatalysts. This pathway aligns more closely with experimental results than previous models. The new pathway employs the cracked adsorbates *(OH)(O) and *(OH)(OH), which, through hydrogen bonding, induce spectator surface *H. This results in an onset potential closely matching experimental values, in stark contrast to the traditional ORR pathway, which keeps adsorbates intact and overestimates the onset potential by 0.7 V. Finally, we introduce electrochemical strain spectroscopy (ESS), a groundbreaking strain analysis technique. ESS combines ab initio calculations with experimental measurements to validate proposed reaction pathways and pinpoint rate-limiting steps.

cond-mat.mtrl-sci

Decentralized Graph Neural Network for Privacy-Preserving Recommendation

Building a graph neural network (GNN)-based recommender system without violating user privacy proves challenging. Existing methods can be divided into federated GNNs and decentralized GNNs. But both methods have undesirable effects, i.e., low communication efficiency and privacy leakage. This paper proposes DGREC, a novel decentralized GNN for privacy-preserving recommendations, where users can choose to publicize their interactions. It includes three stages, i.e., graph construction, local gradient calculation, and global gradient passing. The first stage builds a local inner-item hypergraph for each user and a global inter-user graph. The second stage models user preference and calculates gradients on each local device. The third stage designs a local differential privacy mechanism named secure gradient-sharing, which proves strong privacy-preserving of users' private data. We conduct extensive experiments on three public datasets to validate the consistent superiority of our framework.

cs.IR

Multimodal Operando X-ray Mechanistic Studies of a Bimetallic Oxide Electrocatalyst in Alkaline Media

Furthering the understanding of the catalytic mechanisms in the oxygen reduction reaction (ORR) is critical to advancing and enabling fuel cell technology. In this work, we use multimodal operando synchrotron X-ray diffraction (XRD) and resonant elastic X-ray scattering (REXS) to investigate the interplay between the structure and oxidation state of a Co-Mn spinel oxide electrocatalyst, which has previously shown ORR activity that rivals Pt in alkaline fuel cells. During cyclic voltammetry, the electrocatalyst exhibited a reversible and rapid increase in tensile strain at low potentials, suggesting robust structural reversibility and stability of Co-Mn oxide electrocatalysts during normal fuel cell operating conditions. At low potential holds, exploring the limit of structural stability, an irreversible tetragonal-to-cubic phase transition was observed, which may be correlated to reduction in both Co and Mn valence states. Meanwhile, joint density-functional theory (JDFT) calculations provide insight into how reactive adsorbates induce strain in spinel oxide nanoparticles. Through this work, strain and oxidation state changes that are possible sources of degradation during the ORR in Co-Mn oxide electrocatalysts are uncovered, and the unique capabilities of combining structural and chemical characterization of electrocatalysts in multimodal operando X-ray studies are demonstrated.

cond-mat.mtrl-sci