Searcharxiv⌕ Search

arXiv subjects

Song Chen

Publications and source records attributed to Song Chen.

At least 37 records · Page 2Linked to original sources

Synergy and Competition of Dual Chirality in the Chirality-Induced Spin Selectivity of Supramolecular Helices

Recent progress in constructing supramolecular assemblies with hierarchical chirality offers new opportunities to investigate the chirality-induced spin selectivity (CISS) effect and its potential applications. In this work, we systematically examine the CISS effect in such multichiral systems by designing a class of multilayer helical architectures constructed of stacked and interfaced individual helical rings, each possessing well-defined local chirality. Through controlled interlayer twisting, a global helical handedness is further imposed, forming a multichiral tubular helix. Theoretical calculations reveal that these two distinct chiral hierarchies lead to several unprecedented CISS phenomena, such as enhanced spin polarization arising from cooperative dual chirality, along with the simultaneous emergence of transverse and longitudinal CISS signals. Moreover, interlayer torsional competition modulates the system's response to external fields. The dual-chiral geometry breaks the conventional symmetry of single helices, inducing an anomalous angular phase shift in magnetoresistance. Furthermore, Floquet analysis reveals that the interplay between local and global chirality enables controlled spin polarization switching under circularly polarized light. These findings provide a basic theoretical framework for studying the CISS in multichiral superstructures and establish design principles for coupled optical, magnetic, and spin manipulations, thereby facilitating the development of multichiral spintronic devices.

cond-mat.mes-hall↗

SASQ: Static Activation Scaling for Quantization-Aware Training in Large Language Models

Large language models (LLMs) excel at natural language tasks but face deployment challenges due to their growing size outpacing GPU memory advancements. Model quantization mitigates this issue by lowering weight and activation precision, but existing solutions face fundamental trade-offs: dynamic quantization incurs high computational overhead and poses deployment challenges on edge devices, while static quantization sacrifices accuracy. Existing approaches of quantization-aware training (QAT) further suffer from weight training costs. We propose SASQ: a lightweight QAT framework specifically tailored for activation quantization factors. SASQ exclusively optimizes only the quantization factors (without changing pre-trained weights), enabling static inference with high accuracy while maintaining deployment efficiency. SASQ adaptively truncates some outliers, thereby reducing the difficulty of quantization while preserving the distributional characteristics of the activations. SASQ not only surpasses existing SOTA quantization schemes but also outperforms the corresponding FP16 models. On LLaMA2-7B, it achieves 5.2% lower perplexity than QuaRot and 4.7% lower perplexity than the FP16 model on WikiText2.

cs.CL↗

VeriGRAG: Enhancing LLM-Based Verilog Code Generation with Structure-Aware Soft Prompts

Large language models (LLMs) have demonstrated strong capabilities in generating Verilog code from natural language descriptions. However, Verilog code inherently encodes structural information of hardware circuits. Effectively leveraging this structural information to enhance the functional and syntactic correctness of LLM-generated Verilog code remains a significant challenge. To address this challenge, we propose VeriGRAG , a novel framework that extracts structural graph embeddings from Verilog code using graph neural networks (GNNs). A multimodal retriever then selects the graph embeddings most relevant to the given generation task, which are aligned with the code modality through the VeriFormer module to generate structure-aware soft prompts. Our experiments demonstrate that VeriGRAG substantially improves the correctness of Verilog code generation, achieving state-of-the-art or superior performance across both VerilogEval and RTLLM benchmarks.

cs.AR↗

YOCO: A Hybrid In-Memory Computing Architecture with 8-bit Sub-PetaOps/W In-Situ Multiply Arithmetic for Large-Scale AI

In this paper, we further explore the potential of analog in-memory computing (AiMC) and introduce an innovative artificial intelligence (AI) accelerator architecture named YOCO, featuring three key proposals: (1) YOCO proposes a novel 8-bit in-situ multiply arithmetic (IMA) achieving 123.8 TOPS/W energy-efficiency and 34.9 TOPS throughput through efficient charge-domain computation and timedomain accumulation mechanism. (2) YOCO employs a hybrid ReRAM-SRAM memory structure to balance computational efficiency and storage density. (3) YOCO tailors an IMC-friendly attention computing flow with an efficient pipeline to accelerate the inference of transformer-based AI models. Compared to three SOTA baselines, YOCO on average improves energy efficiency by up to 3.9x-19.9x and throughput by up to 6.8x-33.6x across 10 CNN/transformer models.

cs.AR↗

Reversible magneto ionics in crystallized W Co20Fe60B20 MgO HfO2 ultra-thin films with perpendicular magnetic anisotropy

We have investigated electric field (E-field) induced modulation of perpendicular magnetic anisotropy (PMA) in both amorphous and crystalline W/CoFeB/MgO/HfO2 ultra-thin films. We find that in the amorphous state, the E-field effect is volatile and reversible, which is consistent with the conventional electrostatic effect through charge accumulation and depletion. In the crystallized system annealed at 370°C, we find that two effects are at play, a non-volatile and reversible voltage-induced effect on PMA and an electrostatic response. We discuss these results in terms of higher oxygen mobility at the crystallized CoFeB-MgO interface, which induces a non-volatile magnetoionic response. Modulating PMA in crystallized CoFeB-MgO materials through ionic migration opens the path to integrating magneto-ionics in full magnetic tunnel junctions.

cond-mat.mtrl-sci↗

Accelerated Distributed Aggregative Optimization

This paper delves into the investigation of a distributed aggregative optimization problem within a network. In this scenario, each agent possesses its own local cost function, which relies not only on the local state variable but also on an aggregated function of state variables from all agents. To expedite the optimization process, we amalgamate the heavy ball and Nesterovs accelerated method with distributed aggregative gradient tracking, resulting in the proposal of two innovative algorithms, aimed at resolving the distributed aggregative optimization problem. Our analysis demonstrates that the proposed algorithms can converge to an optimal solution at a global linear convergence rate when the objective function is strongly convex with the Lipschitz-continuous gradient, and when the parameters (e.g., step size and momentum coefficients) are chosen within specific ranges. Additionally, we present several numerical experiments to verify the effectiveness, robustness and superiority of our proposed algorithms.

math.OC↗

Baichuan-Omni-1.5 Technical Report

We introduce Baichuan-Omni-1.5, an omni-modal model that not only has omni-modal understanding capabilities but also provides end-to-end audio generation capabilities. To achieve fluent and high-quality interaction across modalities without compromising the capabilities of any modality, we prioritized optimizing three key aspects. First, we establish a comprehensive data cleaning and synthesis pipeline for multimodal data, obtaining about 500B high-quality data (text, audio, and vision). Second, an audio-tokenizer (Baichuan-Audio-Tokenizer) has been designed to capture both semantic and acoustic information from audio, enabling seamless integration and enhanced compatibility with MLLM. Lastly, we designed a multi-stage training strategy that progressively integrates multimodal alignment and multitask fine-tuning, ensuring effective synergy across all modalities. Baichuan-Omni-1.5 leads contemporary models (including GPT4o-mini and MiniCPM-o 2.6) in terms of comprehensive omni-modal capabilities. Notably, it achieves results comparable to leading models such as Qwen2-VL-72B across various multimodal medical benchmarks.

cs.CL↗

Ocean-OCR: Towards General OCR Application via a Vision-Language Model

Multimodal large language models (MLLMs) have shown impressive capabilities across various domains, excelling in processing and understanding information from multiple modalities. Despite the rapid progress made previously, insufficient OCR ability hinders MLLMs from excelling in text-related tasks. In this paper, we present \textbf{Ocean-OCR}, a 3B MLLM with state-of-the-art performance on various OCR scenarios and comparable understanding ability on general tasks. We employ Native Resolution ViT to enable variable resolution input and utilize a substantial collection of high-quality OCR datasets to enhance the model performance. We demonstrate the superiority of Ocean-OCR through comprehensive experiments on open-source OCR benchmarks and across various OCR scenarios. These scenarios encompass document understanding, scene text recognition, and handwritten recognition, highlighting the robust OCR capabilities of Ocean-OCR. Note that Ocean-OCR is the first MLLM to outperform professional OCR models such as TextIn and PaddleOCR.

cs.CV↗

Baichuan-Omni Technical Report

The salient multimodal capabilities and interactive experience of GPT-4o highlight its critical role in practical applications, yet it lacks a high-performing open-source counterpart. In this paper, we introduce Baichuan-omni, the first open-source 7B Multimodal Large Language Model (MLLM) adept at concurrently processing and analyzing modalities of image, video, audio, and text, while delivering an advanced multimodal interactive experience and strong performance. We propose an effective multimodal training schema starting with 7B model and proceeding through two stages of multimodal alignment and multitask fine-tuning across audio, image, video, and text modal. This approach equips the language model with the ability to handle visual and audio data effectively. Demonstrating strong performance across various omni-modal and multimodal benchmarks, we aim for this contribution to serve as a competitive baseline for the open-source community in advancing multimodal understanding and real-time interaction.

cs.AI↗

SparseMap: Loop Mapping for Sparse CNNs on Streaming Coarse-grained Reconfigurable Array

Streaming coarse-grained reconfgurable array (CGRA) is a promising architecture for data/computing-intensive applications because of its fexibility, high throughput and efcient memory system. However,when accelerating sparse CNNs, the irregular input data demands inside sparse CNNs would cause excessive caching operations (COPs) and multi-cycle internal dependencies (MCIDs) between operations, declining the throughput of the streaming CGRA. We propose a mapping method for sparse CNNs onto streaming CGRA, SparseMap, which incorporates an efcient I/O data management along with operation scheduling and binding, to reduce the COPs and MCIDs, thereby ensuring the optimal throughput of streaming CGRA.The experimental results show SparseMap reduces 92.5% COPs and 46.0 % MCIDs while achieves the same or even smaller initiation interval (II) compared to previous works.

cs.DC↗

Tuning of perpendicular magnetic anisotropy in Bi-substituted yttrium iron garnet films by He$^+$ ion irradiation

We report the continuous tuning of magnetic anisotropy in perpendicularly magnetized bismuth-substituted yttrium iron garnet (Bi-YIG) films via He$^+$ ion irradiation. Our findings indicate that the magnetization direction of epitaxial Bi-YIG films on sGGG substrates transitions from out-of-plane in the as-grown state to in-plane after He+ ion irradiation at a fluence exceeding $2\times 10^{14}$ ions/cm$^2$. The reorientation is attributed to the relaxation of tensile film strain, which reduces the perpendicular magnetic anisotropy without affecting the saturation magnetization. The Gilbert damping parameter and the inhomogeneous broadening of the ferromagnetic resonance linewidth show only minimal increases with ion irradiation. Additionally, at a fluence of $5\times 10^{13}$ ions/cm$^2$, we observe the formation of magnetic bubble domains in the Bi-YIG films. Micromagnetic simulations estimate a Dzyaloshinskii-Moriya interaction of 0.006 mJ/m$^2$, which is insufficient for stabilizing Néel-type skyrmions. Finally, we demonstrate that the effects of He$^+$ ion irradiation can be largely reversed through thermal annealing in an oxygen atmosphere.

cond-mat.mtrl-sci↗

FullAnno: A Data Engine for Enhancing Image Comprehension of MLLMs

Multimodal Large Language Models (MLLMs) have shown promise in a broad range of vision-language tasks with their strong reasoning and generalization capabilities. However, they heavily depend on high-quality data in the Supervised Fine-Tuning (SFT) phase. The existing approaches aim to curate high-quality data via GPT-4V, but they are not scalable due to the commercial nature of GPT-4V and the simplicity of the prompts used to instruct the model. To this end, we devised the FullAnno system, which is a data engine that can generate large-scale, high-quality, and fine-grained image annotations consisting of the category and position of objects, region descriptions, text information, as well as image dense captions. This engine is characterized by its cascade annotation process, which involves multiple expert models and employs rich prompts to instruct LLMs in generating dense image captions. We re-annotated the COCO and Visual Genome datasets using our FullAnno system, tripling the number of object annotations and increasing the length of the original image captions by a factor of 15. Experiments show that the regenerated annotation can significantly enhance the capabilities of LLaVA-v1.5 on several benchmarks. The re-annotated data are available at: https://arcana-project-page.github.io

cs.CV↗

Allspark: Workload Orchestration for Visual Transformers on Processing In-Memory Systems

The advent of Transformers has revolutionized computer vision, offering a powerful alternative to convolutional neural networks (CNNs), especially with the local attention mechanism that excels at capturing local structures within the input and achieve state-of-the-art performance. Processing in-memory (PIM) architecture offers extensive parallelism, low data movement costs, and scalable memory bandwidth, making it a promising solution to accelerate Transformer with memory-intensive operations. However, the crucial issue lies in efficiently deploying an entire model onto resource-limited PIM system while parallelizing each transformer block with potentially many computational branches based on local-attention mechanisms. We present Allspark, which focuses on workload orchestration for visual Transformers on PIM systems, aiming at minimizing inference latency. Firstly, to fully utilize the massive parallelism of PIM, Allspark employs a fine-grained partitioning scheme for computational branches, and formats a systematic layout and interleaved dataflow with maximized data locality and reduced data movement. Secondly, Allspark formulates the scheduling of the complete model on a resource-limited distributed PIM system as an integer linear programming (ILP) problem. Thirdly, as local-global data interactions exhibit complex yet regular dependencies, Allspark provides a two-stage placement method, which simplifies the challenging placement of computational branches on the PIM system into the structured layout and greedy-based binding, to minimize NoC communication costs. Extensive experiments on 3D-stacked DRAM-based PIM systems show that Allspark brings 1.2x-24.0x inference speedup for various visual Transformers over baselines. Compared to Nvidia V100 GPU, Allspark-enriched PIM system yields average speedups of 2.3x and energy savings of 20x-55x.

cs.AR↗

Chirality-Induced Majorana Polarization

To realize Majorana fermions having novel physical features has been developed as a key while difficult task in topological superconductor. Here we have proposed another platform to generate Majorana zero modes (MZMs), which is constructed by a single opened circular helix molecules (CHM) coupled with a s-wave superconductor (with magnetic field) or by an interlinked-CHMs chain coupled with a phase-bias s-wave superconducting heterostructure (without any magnetic field). The MZMs achieved here are tightly associated with the structural chirality in CHMs. Importantly, the left and right handedness may result in completely opposite Majorana polarization (MP), and the local MP is associated to the chiraliy-induced spin polarization. These properties provides us multiple effective ways to detect and regulate the MZMs by using the chirality-induced spin selectivity (CISS) effect and the related spin-polarized currents in chiral materials.

cond-mat.mes-hall↗

Graph Attention-Based Symmetry Constraint Extraction for Analog Circuits

In recent years, analog circuits have received extensive attention and are widely used in many emerging applications. The high demand for analog circuits necessitates shorter circuit design cycles. To achieve the desired performance and specifications, various geometrical symmetry constraints must be carefully considered during the analog layout process. However, the manual labeling of these constraints by experienced analog engineers is a laborious and time-consuming process. To handle the costly runtime issue, we propose a graph-based learning framework to automatically extract symmetric constraints in analog circuit layout. The proposed framework leverages the connection characteristics of circuits and the devices' information to learn the general rules of symmetric constraints, which effectively facilitates the extraction of device-level constraints on circuit netlists. The experimental results demonstrate that compared to state-of-the-art symmetric constraint detection approaches, our framework achieves higher accuracy and F1-score.

cs.LG↗

PiRD: Physics-informed Residual Diffusion for Flow Field Reconstruction

The use of machine learning in fluid dynamics is becoming more common to expedite the computation when solving forward and inverse problems of partial differential equations. Yet, a notable challenge with existing convolutional neural network (CNN)-based methods for data fidelity enhancement is their reliance on specific low-fidelity data patterns and distributions during the training phase. In addition, the CNN-based method essentially treats the flow reconstruction task as a computer vision task that prioritizes the element-wise precision which lacks a physical and mathematical explanation. This dependence can dramatically affect the models' effectiveness in real-world scenarios, especially when the low-fidelity input deviates from the training data or contains noise not accounted for during training. The introduction of diffusion models in this context shows promise for improving performance and generalizability. Unlike direct mapping from a specific low-fidelity to a high-fidelity distribution, diffusion models learn to transition from any low-fidelity distribution towards a high-fidelity one. Our proposed model - Physics-informed Residual Diffusion, demonstrates the capability to elevate the quality of data from both standard low-fidelity inputs, to low-fidelity inputs with injected Gaussian noise, and randomly collected samples. By integrating physics-based insights into the objective function, it further refines the accuracy and the fidelity of the inferred high-quality data. Experimental results have shown that our approach can effectively reconstruct high-quality outcomes for two-dimensional turbulent flows from a range of low-fidelity input conditions without requiring retraining.

physics.flu-dyn↗

IOPS: An Unified SpMM Accelerator Based on Inner-Outer-Hybrid Product

Sparse matrix multiplication (SpMM) is widely applied to numerous domains, such as graph processing, machine learning, and data analytics. However, inner product based SpMM induces redundant zero-element computing for mismatched nonzero operands, while outer product based approach lacks input reuse across Process Elements (PEs) and poor output locality for accumulating partial sum (psum) matrices. Besides, current works only focus on sparse-sparse matrix multiplication (SSMM) or sparse-dense matrix multiplication (SDMM), rarely performing efficiently for both. To address these problems, this paper proposes an unified SpMM accelerator, called IOPS, hybridizing inner with outer products. It reuses the input matrix among PEs with inner product dataflow, and removes zero-element calculations with outer product approach in each PE, which can efficiently process SSMM and SDMM. Moreover, an address mapping method is designed to accumulate the irregular sparse psum matrices, reducing the latency and DRAM access of psum accumulating. Furthermore, an adaptive partition strategy is proposed to tile the input matrices based on their sparsity ratios, effectively utilizing the storage of architecture and reducing DRAM access. Compared with the SSMM accelerator, SpArch, we achieve 1.7x~6.3x energy efficiency and 1.2x~4.4x resource efficiency, with 1.4x~2.1x DRAM access saving.

cs.AR↗

Continuous Cross-resolution Remote Sensing Image Change Detection

Most contemporary supervised Remote Sensing (RS) image Change Detection (CD) approaches are customized for equal-resolution bitemporal images. Real-world applications raise the need for cross-resolution change detection, aka, CD based on bitemporal images with different spatial resolutions. Given training samples of a fixed bitemporal resolution difference (ratio) between the high-resolution (HR) image and the low-resolution (LR) one, current cross-resolution methods may fit a certain ratio but lack adaptation to other resolution differences. Toward continuous cross-resolution CD, we propose scale-invariant learning to enforce the model consistently predicting HR results given synthesized samples of varying resolution differences. Concretely, we synthesize blurred versions of the HR image by random downsampled reconstructions to reduce the gap between HR and LR images. We introduce coordinate-based representations to decode per-pixel predictions by feeding the coordinate query and corresponding multi-level embedding features into an MLP that implicitly learns the shape of land cover changes, therefore benefiting recognizing blurred objects in the LR image. Moreover, considering that spatial resolution mainly affects the local textures, we apply local-window self-attention to align bitemporal features during the early stages of the encoder. Extensive experiments on two synthesized and one real-world different-resolution CD datasets verify the effectiveness of the proposed method. Our method significantly outperforms several vanilla CD methods and two cross-resolution CD methods on the three datasets both in in-distribution and out-of-distribution settings. The empirical results suggest that our method could yield relatively consistent HR change predictions regardless of varying bitemporal resolution ratios. Our code is available at \url{https://github.com/justchenhao/SILI_CD}.

cs.CV↗