SearcharxivSearch

arXiv subjects

Yang Lv

Publications and source records attributed to Yang Lv.

At least 19 recordsLinked to original sources

CRAM-ER: Error-Resilient Spintronic Computational Random Access Memory for Scalable In-Memory Computation

Deep neural networks (DNNs) have achieved state-of-the-art performance across diverse domains. However, typical Von Neumann compute paradigms face severe memory bottlenecks. Emerging near-memory and compute-in-memory approaches alleviate this but incur significant peripheral overhead. Computational Random Access Memory (CRAM) based on MRAM enables in-situ logic without peripheral overhead, offering a dense, energy-efficient solution. However, probabilistic MRAM switching induces gate-level errors that limit the scalability and reliability of CRAM for accelerating DNN. Moreover, the large number of sequential MRAM writes severely constrains CRAM throughput. To address these challenges, we propose an error-resilient CRAM (CRAM-ER) architecture for scalable in-memory matrix-vector multiplications (MVMs). Our error-aware hardware-software co-design framework leverages a hybrid spintronic-CRAM + CMOS adder-tree architecture to mitigate the impact of device-level errors, demonstrating MVM functionality with high area and energy efficiency. We further develop an error-aware model fine-tuning and fine-grained error correction for enhanced error resilience. Evaluations of the CMOS+spintronic hybrid architecture on DNN benchmarks show near-lossless accuracy while reducing CRAM latency by up to 2 orders of magnitude, outperforming CPU/GPU+high-bandwidth DRAM in both energy efficiency and energy-delay product.

cs.AR

Electronic excitation of ultrafast collective amorphous-amorphous transitions in glassy phase-change material

The intrinsic nature of glass states and glass transitions remain a fundamental open question in condensed-matter physics and materials science. The key to solving the glass transition problem lies in achieving a complete understanding of the physics governing the structural relaxation. Nonetheless, directly probing dynamic atomic-scale structural changes in order to identify the precise local structural motifs and establish quantitative structure-property relationships remains an outstanding challenge. By combining femtosecond electron diffraction with time-dependent density-functional theory molecular dynamics simulations, we directly capture ultrafast amorphous-amorphous transitions indicated by collective bond stretching (0.2 ps) and angle bending (0.5-2 ps) in glassy phase-change material GeTe. The ultrafast bond stretching is accompanied by localized oscillation modes with the frequency of 3.10 THz, unambiguously signaling the local Peierls-like bonding structure and the flexibility of these polarized bonds. These ultrafast collective atomic motions, captured across timescales ranging from femtoseconds to picoseconds, directly reveals the structural origin of the boson peak and provide compelling evidence for many-body interactions in amorphous materials. Furthermore, the ultrafast amorphous-amorphous transitions induce a drastic insulator-metal transition, directly revealing both the underlying switching mechanism and the fundamental speed limit of the ovonic threshold switch. These insights establish a fundamental framework for rationally engineering relaxation pathways and phase-change/threshold switch in amorphous materials. Femtosecond electron diffraction provides a powerful novel approach to deciphering the structural complexity and functional mechanisms of amorphous materials by resolving collective atomic motions from random diffusion dynamics in the time domain.

cond-mat.mtrl-sci

Sub-angstrom many-body localization driven by phononic flat bands in real quantum materials

Defects, fluctuations, degenerate states and correlated interactions facilitate the emergence of exotic properties in condensed matter systems while also inducing atomic-scale local correlated structures that deviate from the average long-range order. Establishing the structure-property relationship from the perspective of these atomic-scale local correlated structures remains ambiguous and controversial due to the lack of direct methods for identifying such local correlated structures. In this work, based on the photoexcited ultrafast structural response, we propose a Bragg scattering phase breaking regime to identify sub-angstrom local correlated structures in quantum materials. With this regime, we unambiguously identify the many-body-interaction driven local correlated structures in the low temperature ground state of AgCrSe2, characterized by static off-center displacements of Ag atoms ranging from 0 to 0.5 angstrom. The competition between Ag-Ag Coulomb correlations and potential wells induced by CrSe2 layers, leading to phononic flat bands and driving the system into a many body localization (MBL) regime. As temperature rising, these static local correlated structures transform to a dynamic state where the thermal fluctuations overwhelm the multiple localized states. These distinctive local correlated structures constitute the first experimental observation of MBL with vortex-like topological characteristic in a real material system. Emergent vibrational modes arising from MBL have been confirmed and show excellent agreement with inelastic neutron scattering experiments. Our work not only offers a universal approach to characterize sub-angstrom local correlated structures across a wide range of quantum materials but also deepens our understanding of the fundamental mechanism behind exotic properties from the perspective of atomic-scale local correlated structures.

cond-mat.mtrl-sci

MBD: A Model-Based Debiasing Framework Across User, Content, and Model Dimensions

Modern recommendation systems rank candidates by aggregating multiple behavioral signals through a value model. However, many commonly used signals are inherently affected by heterogeneous biases. For example, watch time naturally favors long-form content, loop rate favors short - form content, and comment probability favors videos over images. Such biases introduce two critical issues: (1) value model scores may be systematically misaligned with users' relative preferences - for instance, a seemingly low absolute like probability may represent exceptionally strong interest for a user who rarely engages; and (2) changes in value modeling rules can trigger abrupt and undesirable ecosystem shifts. In this work, we ask a fundamental question: can biased behavioral signals be systematically transformed into unbiased signals, under a user - defined notion of ``unbiasedness'', that are both personalized and adaptive? We propose a general, model-based debiasing (MBD) framework that addresses this challenge by augmenting it with distributional modeling. By conditioning on a flexible subset of features (partial feature set), we explicitly estimate the contextual mean and variance of the engagement distribution for arbitrary cohorts (e.g., specific video lengths or user regions) directly alongside the main prediction. This integration allows the framework to convert biased raw signals into unbiased representations, enabling the construction of higher-level, calibrated signals (such as percentiles or z - scores) suitable for the value model. Importantly, the definition of unbiasedness is flexible and controllable, allowing the system to adapt to different personalization objectives and modeling preferences. Crucially, this is implemented as a lightweight, built-in branch of the existing MTML ranking model, requiring no separate serving infrastructure.

cs.LG

Unlock Anionic Behavior of Calcium Through Pressure Engineering

An isolated calcium (Ca) atom has empty d-orbitals under ambient conditions. However, s-d band hybridization has been observed in both elemental Ca and compounds by manipulating thermodynamic conditions. Here, we reveal that the Ca 3d-band can even capture electrons from halogen atoms under pressure, exhibiting anionic behaviors in iodides. We predict a CsCl-type monovalent CaI at above 50 GPa by employing first-principles structural searching and successfully identified the phase at 84 GPa using in situ X-ray diffraction. We further reveal that, due to the effect of orbital broadening, unusual charge transfer from the 5p orbitals of I to the 3d orbitals of Ca in CaI, gradually reverses the ionicity of Ca and becomes the anionic ICa at 485 GPa. Multivalent Ca stabilizes a set of metallic iodides with eight- to ten-fold iodine hyper-coordination. Our findings demonstrate that the valence states of Ca can vary from negative to +2, suggesting much greater complexity of Ca chemistry under ultrahigh pressures.

cond-mat.mtrl-sci

DPFNAS: Differential Privacy-Enhanced Federated Neural Architecture Search for 6G Edge Intelligence

The Sixth-Generation (6G) network envisions pervasive artificial intelligence (AI) as a core goal, enabled by edge intelligence through on-device data utilization. To realize this vision, federated learning (FL) has emerged as a key paradigm for collaborative training across edge devices. However, the sensitivity and heterogeneity of edge data pose key challenges to FL: parameter sharing risks data reconstruction, and a unified global model struggles to adapt to diverse local distributions. In this paper, we propose a novel federated learning framework that integrates personalized differential privacy (DP) and adaptive model design. To protect training data, we leverage sample-level representations for knowledge sharing and apply a personalized DP strategy to resist reconstruction attacks. To ensure distribution-aware adaptation under privacy constraints, we develop a privacy-aware neural architecture search (NAS) algorithm that generates locally customized architectures and hyperparameters. To the best of our knowledge, this is the first personalized DP solution tailored for representation-based FL with theoretical convergence guarantees. Our scheme achieves strong privacy guarantees for training data while significantly outperforming state-of-the-art methods in model performance. Experiments on benchmark datasets such as CIFAR-10 and CIFAR-100 demonstrate that our scheme improves accuracy by 6.82\% over the federated NAS method PerFedRLNAS, while reducing model size to 1/10 and communication cost to 1/20.

cs.LG

Ultrasound Tomography of Musculoskeletal Tissues with Generative Neural Physics

Ultrasound Tomography (UT) is a radiation-free, high-resolution modality, but remains limited for musculoskeletal imaging due to the high computational cost and instability of full-waveform inversion in strongly scattering media. We propose a generative neural physics framework that couples generative networks with physics-informed neural simulation for fast, high-fidelity 3D UT. By learning a compact surrogate of ultrasonic wave propagation from a limited set of cross-modality images, our method merges the accuracy of wave modeling with the efficiency and stability of deep learning. This enables accurate quantitative imaging of in vivo musculoskeletal tissues, producing spatial maps of acoustic properties beyond reflection-mode images. On synthetic and in vivo data of breasts, arms, and legs, we reconstruct 3D maps of tissue parameters in under ten minutes, with sensitivity to acoustic variations in musculoskeletal tissues and resolution comparable to MRI. By overcoming computational bottlenecks in strongly scattering regimes, this approach demonstrates the feasibility of quantitative UT for musculoskeletal imaging and advances its development toward future routine clinical use.

cs.CV

A Survey on Medical Image Compression: From Traditional to Learning-Based Approaches

The exponential growth of medical imaging has created significant challenges in data storage, transmission, and management for healthcare systems. In this vein, efficient compression becomes increasingly important. Unlike natural image compression, medical image compression prioritizes preserving diagnostic details and structural integrity, imposing stricter quality requirements and demanding fast, memory-efficient algorithms that balance computational complexity with clinically acceptable reconstruction quality. Meanwhile, the medical imaging family includes a plethora of modalities, each possessing different requirements. For example, 2D medical image (e.g., X-rays, histopathological images) compression focuses on exploiting intra-slice spatial redundancy, while volumetric medical image faces require handling intra-slice and inter-slice spatial correlations, and 4D dynamic imaging (e.g., time-series CT/MRI, 4D ultrasound) additionally demands processing temporal correlations between consecutive time frames. Traditional compression methods, grounded in mathematical transforms and information theory principles, provide solid theoretical foundations, predictable performance, and high standardization levels, with extensive validation in clinical environments. In contrast, deep learning-based approaches demonstrate remarkable adaptive learning capabilities and can capture complex statistical characteristics and semantic information within medical images. This comprehensive survey establishes a two-facet taxonomy based on data structure (2D vs 3D/4D) and technical approaches (traditional vs learning-based), thereby systematically presenting the complete technological evolution, analyzing the unique technical challenges, and prospecting future directions in medical image compression.

eess.IV

COMS-Integrated Atomic Vapor Cells with Ultra-long Optical Access for Highly Sensitive and Scalable Quantum Sensors

The most appealing features of chip-scale quantum sensors are their capability to maintain extreme sensitivity while enabling large-scale batch manufacturing. This necessitates high-level integration and wafer-level fabrication of atomic vapor cells. In this paper, we describe a micromachining paradigm for wafer-level atomic vapor cells functionalized by CMOS-compatible non-magnetic heaters and temperature sensors and demonstrate several innovative applications. Leveraging standard micro-nanofabrication technology, the integrated vapor cells achieved an ultra-long optical access of 5 mm, nearly four time that of previously microfabricated vapor cells. The feasibility of the integrated atomic vapor cells fabrication process was verified by a consecutive 30-day aging test in a harsh environment (operating temperature of 473 K and vacuum of approximately 1 Pa). Benefiting from the ultra-long optical path, we observed several typical quantum effects, including the saturation absorption and spin fluctuations, a regime previously inaccessible with conventional micromachined vapor cells. Finally, a zero-field quantum magnetometry with an ultra-high magnetic sensitivity of 12 fT/Hz1/2 was also demonstrated. Our achievements broaden the potential applications of microfabricated atomic vapor cells and pave the way for scalable manufacturing of ultrasensitive, chip-scale quantum sensors.

quant-ph

HGFormer: A Hierarchical Graph Transformer Framework for Two-Stage Colonel Blotto Games via Reinforcement Learning

Two-stage Colonel Blotto game represents a typical adversarial resource allocation problem, in which two opposing agents sequentially allocate resources in a network topology across two phases: an initial resource deployment followed by multiple rounds of dynamic reallocation adjustments. The sequential dependency between game stages and the complex constraints imposed by the graph topology make it difficult for traditional approaches to attain a globally optimal strategy. To address these challenges, we propose a hierarchical graph Transformer framework called HGformer. By incorporating an enhanced graph Transformer encoder with structural biases and a two-agent hierarchical decision model, our approach enables efficient policy generation in large-scale adversarial environments. Moreover, we design a layer-by-layer feedback reinforcement learning algorithm that feeds the long-term returns from lower-level decisions back into the optimization of the higher-level strategy, thus bridging the coordination gap between the two decision-making stages. Experimental results demonstrate that, compared to existing hierarchical decision-making or graph neural network methods, HGformer significantly improves resource allocation efficiency and adversarial payoff, achieving superior overall performance in complex dynamic game scenarios.

cs.AI

Q-learning-based Hierarchical Cooperative Local Search for Steelmaking-continuous Casting Scheduling Problem

The steelmaking continuous casting scheduling problem (SCCSP) is a critical and complex challenge in modern steel production, requiring the coordinated assignment and sequencing of steel charges across multiple production stages. Efficient scheduling not only enhances productivity but also significantly reduces energy consumption. However, both traditional heuristics (e.g., two-stage local search) and recent metaheuristics often struggle to adapt to the dynamic characteristics of practical SCCSP instances. To address these limitations, this paper introduces a novel Q learning based hierarchical cooperative local search framework, termed HierC_Q, aimed at minimizing the weighted sum of the maximum completion time and the average waiting time in SCCSP. The core contributions of HierC_Q are twofold. First, considering the intrinsic coupling properties of the SCCSP, a dedicated reward function is proposed based on a novel coupling measure (CM), guiding the exploration process towards promising regions of the solution space. Second, a hierarchical architecture is devised, comprising two distinct tiers: the learn to improve (L2I) tier and the "disturb to renovate" (D2R) tier. The L2I tier performs deep exploitation within promising regions using two independent Q-learning-based local search frameworks (QLSFs) tailored for subproblems, along with a synergy QLSF designed for the main problem. To enhance the effectiveness of local search, a validity evaluation approach and a speed-up evaluation method are also intro-duced, grounded in a detailed study of the problem's structure. Meanwhile, the D2R tier incorporates a perturbation and construction based solution renewal strategy to mitigate the risk of premature convergence. The superiority and effectiveness of HierC_Q are demonstrated through extensive comparisons with eleven local search frameworks and nine state-of-the-art algorithms.

eess.SY

Modulation of switching dynamics in magnetic tunnel junctions for low-error-rate computational random-access memory

The conventional computer architecture has been facing challenges answering the ever-increasing demands from emerging applications, such as AI, for energy-efficient computation and memory hardware systems. Computational Random Access Memory (CRAM) represents a true in-memory computing paradigm that integrates logic and memory functions within the same array. At its core, CRAM relies on Magnetic Tunnel Junctions (MTJs), which serve as the foundational building blocks for implementing both memory storage and logic operations. However, a key challenge in CRAM lies in the non-ideal error rates associated with switching dynamics of MTJs, necessitating innovative approaches to reduce errors and optimize logic margins. This work proposes a novel approach of utilizing the voltage-controlled magnetic anisotropy (VCMA) to steepen the switching probability transfer curve (SPTC), thereby significantly reducing the logic operation error rate in CRAM. Using several numerical modeling tools, we validate the effectiveness of VCMA in modulating the energy barrier and switching dynamics in MTJs. It is revealed that the VCMA effect significantly reduces the error rate of CRAM by 61.43% at a VCMA coefficient of 200 fJ/V/m compared to CRAM without VCMA. The reduction of error rate is further rapidly amplified with an increasing TMR ratio. Furthermore, the introduction of the VCMA effect decreases the logic voltage (Vlogic) required for logic operations in CRAM and results in reduction of energy consumption. Our work serves as a first exploration in reducing the error rate in CRAM by modifying SPTC in MTJs.

cs.ET

Demonstration of Electron-Mediated Voltage-Controlled Exchange Coupling in Perpendicular Magnetic Tunnel Junctions

Electron-mediated voltage control of exchange coupling (EM-VCEC) has been proposed as a mechanism for magnetization switching via modulation of spin-dependent electron reflection. However, its experimental verification has been challenging due to the coexistence of slower, voltage-induced ionic effects. Here, we fabricate magnetic tunnel junction (MTJ) devices that enable nanosecond timescale voltage application. Our results reveal rapid exchange coupling modulation on the nanosecond timescale, consistent with an electronic origin. The observed enhancement and saturation at low temperatures further rule out ionic migration, conclusively confirming the electronic nature of the mechanism. These results establish EM-VCEC as a viable mechanism for fast and energy-efficient voltage-driven magnetic switching.

physics.app-ph

Energy Efficient Stochastic Signal Manipulation in Superparamagnetic Tunnel Junctions via Voltage-Controlled Exchange Coupling

Superparamagnetic tunnel junctions (sMTJs) are emerging as promising components for stochastic units in neuromorphic computing, owing to their tunable random switching behavior. Conventional MTJ control methods, such as spin-transfer torque (STT) and spin-orbit torque (SOT), often require substantial power. Here, we introduce the voltage-controlled exchange coupling (VCEC) mechanism, enabling switching between antiparallel and parallel states in sMTJs with an ultralow power consumption of only 40 nW, approximately two orders of magnitude lower than conventional STT-based sMTJs. This mechanism yields a sigmoid-shaped output response, making it ideally suited for neuromorphic computing applications. Furthermore, we validate the feasibility of integrating VCEC with the SOT current control, offering an additional dimension for magnetic state manipulation. This work marks the first practical demonstration of VCEC effect in sMTJs, highlighting its potential as a low-power control solution for probabilistic bits in advanced computing systems.

physics.app-ph

A Local Information Aggregation based Multi-Agent Reinforcement Learning for Robot Swarm Dynamic Task Allocation

In this paper, we explore how to optimize task allocation for robot swarms in dynamic environments, emphasizing the necessity of formulating robust, flexible, and scalable strategies for robot cooperation. We introduce a novel framework using a decentralized partially observable Markov decision process (Dec_POMDP), specifically designed for distributed robot swarm networks. At the core of our methodology is the Local Information Aggregation Multi-Agent Deep Deterministic Policy Gradient (LIA_MADDPG) algorithm, which merges centralized training with distributed execution (CTDE). During the centralized training phase, a local information aggregation (LIA) module is meticulously designed to gather critical data from neighboring robots, enhancing decision-making efficiency. In the distributed execution phase, a strategy improvement method is proposed to dynamically adjust task allocation based on changing and partially observable environmental conditions. Our empirical evaluations show that the LIA module can be seamlessly integrated into various CTDE-based MARL methods, significantly enhancing their performance. Additionally, by comparing LIA_MADDPG with six conventional reinforcement learning algorithms and a heuristic algorithm, we demonstrate its superior scalability, rapid adaptation to environmental changes, and ability to maintain both stability and convergence speed. These results underscore LIA_MADDPG's outstanding performance and its potential to significantly improve dynamic task allocation in robot swarms through enhanced local collaboration and adaptive strategy execution.

cs.AI

Staircase Cascaded Fusion of Lightweight Local Pattern Recognition and Long-Range Dependencies for Structural Crack Segmentation

Accurately segmenting structural cracks at the pixel level remains a major hurdle, as existing methods fail to integrate local textures with pixel dependencies, often leading to fragmented and incomplete predictions. Moreover, their high parameter counts and substantial computational demands hinder practical deployment on resource-constrained edge devices. To address these challenges, we propose CrackSCF, a Lightweight Cascaded Fusion Crack Segmentation Network designed to achieve robust crack segmentation with exceptional computational efficiency. We design a lightweight convolutional block (LRDS) to replace all standard convolutions. This approach efficiently captures local patterns while operating with a minimal computational footprint. For a holistic perception of crack structures, a lightweight Long-range Dependency Extractor (LDE) captures global dependencies. These are then intelligently unified with local patterns by our Staircase Cascaded Fusion Module (SCFM), ensuring the final segmentation maps are both seamless in continuity and rich in fine-grained detail. To comprehensively evaluate our method, this paper created the challenging TUT benchmark dataset and evaluated it alongside five other public datasets. The experimental results show that the CrackSCF method consistently outperforms the existing methods, and it demonstrates greater robustness in dealing with complex background noise. On the TUT dataset, CrackSCF achieved 0.8382 on F1 score and 0.8473 on mIoU, and it only required 4.79M parameters.

cs.CV

EnviroExam: Benchmarking Environmental Science Knowledge of Large Language Models

In the field of environmental science, it is crucial to have robust evaluation metrics for large language models to ensure their efficacy and accuracy. We propose EnviroExam, a comprehensive evaluation method designed to assess the knowledge of large language models in the field of environmental science. EnviroExam is based on the curricula of top international universities, covering undergraduate, master's, and doctoral courses, and includes 936 questions across 42 core courses. By conducting 0-shot and 5-shot tests on 31 open-source large language models, EnviroExam reveals the performance differences among these models in the domain of environmental science and provides detailed evaluation standards. The results show that 61.3% of the models passed the 5-shot tests, while 48.39% passed the 0-shot tests. By introducing the coefficient of variation as an indicator, we evaluate the performance of mainstream open-source large language models in environmental science from multiple perspectives, providing effective criteria for selecting and fine-tuning language models in this field. Future research will involve constructing more domain-specific test sets using specialized environmental science textbooks to further enhance the accuracy and specificity of the evaluation.

cs.CL

Experimental demonstration of magnetic tunnel junction-based computational random-access memory

Conventional computing paradigm struggles to fulfill the rapidly growing demands from emerging applications, especially those for machine intelligence, because much of the power and energy is consumed by constant data transfers between logic and memory modules. A new paradigm, called "computational random-access memory (CRAM)" has emerged to address this fundamental limitation. CRAM performs logic operations directly using the memory cells themselves, without having the data ever leave the memory. The energy and performance benefits of CRAM for both conventional and emerging applications have been well established by prior numerical studies. However, there lacks an experimental demonstration and study of CRAM to evaluate its computation accuracy, which is a realistic and application-critical metrics for its technological feasibility and competitiveness. In this work, a CRAM array based on magnetic tunnel junctions (MTJs) is experimentally demonstrated. First, basic memory operations as well as 2-, 3-, and 5-input logic operations are studied. Then, a 1-bit full adder with two different designs is demonstrated. Based on the experimental results, a suite of modeling has been developed to characterize the accuracy of CRAM computation. Scalar addition, multiplication, and matrix multiplication, which are essential building blocks for many conventional and machine intelligence applications, are evaluated and show promising accuracy performance. With the confirmation of MTJ-based CRAM's accuracy, there is a strong case that this technology will have a significant impact on power- and energy-demanding applications of machine intelligence.

cs.ET