SearcharxivSearch

arXiv subjects

Yao Chen

Publications and source records attributed to Yao Chen.

At least 19 recordsLinked to original sources

Memory Augmentation Unlocks Efficient Chain-of-Thought Reasoning

Large language models often rely on Chain-of-Thought (CoT) reasoning to solve complex tasks, but verbose reasoning traces introduce substantial inference overhead. CoT compression shortens generation, yet aggressive compression may disrupt logical coherence and degrade performance. We formalize this trade-off as the Context-Generation Substitution Law, where explicit reasoning context substitutes for part of decode-time generation. Based on this principle, we propose Memory-Augmented Compression, a training-free framework that constructs reusable reasoning memories from historical traces and retrieves them as prefill-side scaffolds. Rather than using raw demonstrations, these memories summarize reusable reasoning patterns, key constraints, and critical operations to compensate for information lost during compression. Experiments show that Memory consistently improves prompt-based Chain-of-Draft (CoD) compression across mathematical reasoning, complex reasoning, and science question answering tasks, yielding accuracy gains of 21.4, 28.0, 29.5, and 6.61 points over CoD on GSM8K, MATH, BBH, and MMLU-Sci, while achieving a 1.14-1.49x latency speedup latency speedup over standard CoT. Memory is also compatible with token-level, reasoning-trace-level, and inference-state compression mechanisms.

cs.CL

Empirical Simulation of Survival and Mixed-Type Data for Clinical Trial Design

Simulating realistic time-to-event data is essential for planning and evaluating complex clinical trial designs. Conventional approaches often sample event times from parametric families, such as Weibull or log-normal distributions, which restrict hazard shapes and may poorly represent observed survival data. We propose an empirical copula-based framework for simulating multivariate data containing continuous, binary, count, and right-censored time-to-event variables. The method completes censored historical survival data using a two-zone procedure that combines conditional Kaplan-Meier imputation with a parametric tail. It matches a target survival distribution through a log-scale location-scale transformation and a power distortion of the empirical percentile function, while preserving historical dependence through a Gaussian copula fitted to rank correlations. In an oncology trial of previously treated non-small-cell lung cancer, the method reconstructs overall survival and progression-free survival curves for the experimental arm using control-arm data and a small set of target percentiles. Simulations preserve rank correlations among baseline covariates and the dependence between progression-free and overall survival, with censored Kendall's tau of 0.522 compared with 0.549 in the observed data. The method is implemented in the R package EmpiricalSim.

stat.ME

Alfv\'enic Motions in a Stratified Open Flux Tube: Transition from Propagating to Locally Standing Motions and Implications for the Kelvin-Helmholtz Instability

Standing transverse waves in closed coronal structures have been widely studied as a possible route to energy dissipation, with resonant absorption transferring kink wave energy to localized Alfv\'enic motions and the Kelvin-Helmholtz instability (KHI) accelerating the formation of small dissipative scales. However, it remains unclear whether the same mechanism applies to the open corona, given the long-standing consensus that the KHI tends to be prohibited for propagating Alfv\'enic waves. Within the framework of magnetohydrodynamics (MHD), we perform three-dimensional MHD simulations of boundary-driven kink waves in a gravitationally stratified open flux tube extending from the chromosphere into the corona. We find that propagating waves in open magnetic structures can also drive the system toward a turbulent state, with KH vortices clearly identifiable across the flux tube. This occurs because resonant absorption transfers energy from the propagating kink waves to azimuthal Alfv\'enic motions near the tube boundary, and wave reflection off the gradient of the Alfv\'en speed subsequently enables these boundary motions to acquire a locally standing character. Our results provide a possible answer to the long-standing question of whether and how propagating waves in open magnetic structures can generate nonlinear turbulent fine structures despite their globally propagating nature.

astro-ph.SR

CW-Ghost: Search-Free Granularity Selection for Helper-Thread Prefetching via Capacity Windows

Helper-thread prefetching hides the latency of irregular memory accesses by executing address dependency chains ahead of the main thread. However, its effectiveness depends on the range of future iterations covered by the helper thread. A fixed coverage range cannot consistently accommodate different workloads and processors, whereas exhaustively evaluating candidate configurations incurs substantial configuration cost. This paper presents CW-Ghost, which uses a single offline profiling run to estimate the average demand cache line fill volume generated per target iteration in a target region. CW-Ghost combines this estimate with a cache capacity budget to derive a Capacity Window, which determines the iteration granularity of each prefetch chunk. In addition, bounded chunk-level synchronization limits the number of chunks by which the helper thread may run ahead of the main thread. Across 14 workload instances evaluated on Intel and AMD CPU platforms, CW-Ghost achieves geometric mean speedups of 1.54x and 1.33x, respectively, over the original programs. Compared with Ghost Threading, it improves geometric mean performance by 15.8% and 10.8%, respectively, while achieving more than 99% of the empirically optimal performance within the candidate set on both platforms. These results demonstrate that cache capacity constraints can effectively guide the selection of granularity for helper-thread prefetching.

cs.DC

Dual-Perspective Microwave and Hard X-ray Constraints of Asymmetric Nonthermal Loops in an X-class Flare

We report dual-perspective microwave and HXR observations of an X-class flare on 2024 May 15, taking advantage of the unique geometry of a front-side view from the Earth and a back-side perspective from the Solar Orbiter (SolO). Using spatially resolved imaging spectroscopy from Siberian Radioheliograph (SRH) together with Chashan Broadband Solar millimeter spectrometer (CBSmm) and STIX data, we identify a set of nonthermal flaring loops in an asymmetric magnetic field, with microwave sources located near the loop top and HXR sources associated with the southern footpoint. Compared with HXR, the microwave emission shows an opposite ascending trend, an increasing time lag in time profile, and a distinctive ``SHH'' spectral pattern, which we attribute to energy-dependent trapping and precipitation of energetic electrons in an asymmetric magnetic configuration. Flux pulsations and their spectral and polarization signatures are consistent with intermittent particle acceleration rather than MHD wave modulation. Microwave magnetic diagnostics, corroborated by non-linear force free field (NLFFF) extrapolation, provide key constraints on the three-dimensional magnetic configuration. The dual-perspective flux profile comparison and consistent QPP signatures across wavelengths together support a self-consistent picture of energy-dependent electron trapping, precipitation, and transport in these asymmetric loops.

astro-ph.SR

Kinetic Processes to Radio Burst: First Observational-driven Study in Coronal Loops

Understanding the origin of coherent solar radio bursts requires linking macroscopic coronal structures with the kinetic processes responsible for wave generation. We investigate emissions from slowly positively drifting bursts (SPDBs), a specific type of solar radio emission. SPDBs provide observational constraints for modeling beam-plasma interactions in coronal loops, serving as a basis for a multiscale description. We employ a three--stage numerical framework that combines nonlinear force--free field (NLFFF) magnetic extrapolation, guiding--center simulations, and fully kinetic particle--in--cell (PIC) modeling. The background plasma density is described using a hydrostatic model consistent with active--region conditions, producing a plasma--frequency gradient comparable to that inferred from the observed SPDBs spectrum. Energetic electrons injected near the loop top evolve through magnetic mirroring, pitch--angle scattering, turbulence development, and partial precipitation. Evolved velocity distribution functions (EVDFs) are sampled after approximately one bounce period and used in PIC simulations to evaluate emission properties. The results show that the evolved beam distribution of energetic electrons predominantly excites beam--Langmuir waves and fundamental plasma emission along the loop, with the emission intensity gradually decreasing from the loop top toward the footpoint. The modest initial beam velocity of energetic electrons explains the inefficient generation of harmonic plasma emission. The temporal evolution of the modeled emission reproduces key SPDB characteristics, including the ~ 4s duration and frequency drift behavior. These results suggest plasma emission explains the mechanisms behind SPDB generation and demonstrate the feasibility of a unified model connecting coronal magnetic topology, particle transport, and radio emission.

astro-ph.SR

Advancing axion detection: Photon regeneration in high-sensitivity Penning trap experiments

The axion, a hypothetical particle proposed to solve the strong CP problem and considered a viable candidate for dark matter, has prompted extensive experimental efforts for its detection. This study presents a novel approach combining photon regeneration techniques ("light-shining-through-a-wall") with high-sensitivity Penning trap technologies to enhance the search for axions. Penning traps offer significant advantages, including precise electromagnetic field measurement, strong magnetic fields, and single-particle detection capabilities. By integrating these traps with resonantly enhanced photon regeneration using microwave cavities, our proposed method significantly increases sensitivity to axion-photon couplings. Preliminary calculations demonstrate an unprecedented achievable sensitivity, reaching an axion-photon coupling constant limit of $g_{a\gamma \gamma } \le 7.10\times 10^{-8}\mathrm{GeV ^{-1}}$ in just one day, specifically targeting axion energies below 1 MHz. This experimental setup presents a robust and controlled platform, circumventing astrophysical uncertainties, and represents a substantial advancement in laboratory searches for axions and our understanding of dark matter.

hep-ph

ConSA: Controllable Sparsity in Hybrid Attention via Learnable Allocation

Hybrid architectures combining full attention (FA) and sliding-window attention (SWA) are a promising paradigm for efficient LLM inference. However, existing methods typically rely on hand-crafted rules or simple post-hoc heuristics for FA/SWA allocation and offer limited analysis of the attention behaviors underlying these designs. We propose Controllable Sparsity in Hybrid Attention (ConSA), a framework that learns optimal FA/SWA assignment under a user-specified sparsity target. ConSA employs L0 regularization to learn binary masks selecting between FA and SWA for each attention unit, while an augmented Lagrangian constraint enforces the target sparsity at either layer or KV-head granularity. We evaluate ConSA on two LLMs at the 0.6B and 1.7B scales. Learned allocations consistently outperform rule-based baselines, with KV-head-wise allocation yielding clear gains over layer-wise allocation. The learned patterns place SWA in the bottom layers and concentrate FA into contiguous middle-layer blocks, diverging from evenly interleaved patterns in rule-based methods. This structure persists across model scales, sparsity levels, and allocation granularities, revealing a fine-grained spectrum of intrinsic attention behaviors that underlies the learned allocation.

cs.CL

Approaching Shannon Bound with Lossless LLM Weight Compression

Large language models (LLMs) now scale to trillions of parameters, driving weight storage into the terabyte regime and creating an acute mismatch with GPU memory capacity. Although lossless compression is widely effective in other domains, it remains underutilized in LLM systems. Through a comprehensive entropy study across models from 1.5B to 405B parameters and numeric formats ranging from bf16 to int4 and AWQ/SQ8, we find that LLM weights contain far less intrinsic randomness than their stored bitwidth implies, their effective entropy is 2-10x lower, indicating that up to a 10x footprint reduction is theoretically achievable without altering any weight values. Leveraging this insight, we introduce a tile-level, on-the-fly lossless decompression framework based on Asymmetric Numeral Systems that aligns decoding with the GEMM tiling pattern of GPU inference. Our design achieves bit-rates within 0.01-0.1 bits of the Shannon limit across a wide range of LLM numerical formats, demonstrating that nearly all statistical redundancy is eliminated. Integrated into the SGLang serving framework with multi-GPU support, our approach increases the maximum batch size of Qwen-14B from 47 to 75, improving throughput by up to 1.2x. On Mixtral-176B, the feasible batch size increases from 20 to 95 (4.8x), yielding up to 1.6x throughput improvement. Compared to state-of-the-art lossless compression approaches NeuZip and DFloat11, our design further improves throughput by up to 11x.

cs.AR

Structure-Guided Adaptive Propagation for Protein-Protein Interaction Site Prediction

Accurate prediction of protein-protein interaction sites (PPIS) is essential for understanding cellular processes, disease mechanisms, and therapeutic target discovery. Graph-based deep learning has advanced PPIS prediction by incorporating residue-level structural context. However, most graph-based models still rely on fixed propagation schemes that treat all residues similarly, despite the structural and functional heterogeneity of protein interfaces. Such propagation may limit the ability to adapt information diffusion to local geometric environments, making it difficult to distinguish true interaction sites from structurally similar non-interacting neighbors. We present SGAP-PPIS, a structure-guided adaptive propagation model for PPIS prediction. Rather than using a fixed propagation mechanism, SGAP-PPIS leverages multi-scale geometric states from an equivariant graph neural network to generate residue-wise propagation coefficients. This design allows each residue to adaptively balance local feature preservation and neighborhood diffusion according to its geometric microenvironment. Experimental results show that SGAP-PPIS achieves competitive performance among the state-of-the-art methods on Test\_60. Ablation studies show that geometry-conditioned adaptive propagation, scale-aligned geometric guidance, and multi-step propagation-state representation jointly drive these improvements.

cs.AI

MADS: Model-Aware Diverse Core Set Selection for Instruction Tuning

Instruction fine-tuning is employed to enhance the instruction-following ability of large language models (LLMs). As the amount of instruction fine-tuning data increases, selecting the optimal core set becomes particularly important. However, ensuring the diversity of the core set remains a significant challenge. Existing methods predominantly distinguish different training data based on the text features themselves, decoupled from LLMs' own understanding and representation of the data. To address this issue, we propose a Model-Aware Diverse Core Set Selection method, which distinguishes data features based on the neural activation states during LLM inference. This approach serves as an efficient instantiation of coverage-based selection using model-intrinsic activation features to ensure the diversity in the core set. We extensively evaluate our method on six benchmarks that cover five distinct tasks. In our method, the core set selected by the 3B-parameter LLM performs effectively when utilized to fine-tune larger models with 7B, 8B, and 13B parameters. Experimental results on the Alpaca-GPT4 dataset, which comprises 52K instruction-response pairs, show that the core set, sized at 15\% of the original dataset and selected by Llama-3.2-3B-Instruct, achieves an average improvement of 2.5\% when fine-tuning four larger base models compared with training on the full dataset. The experimental results demonstrate that our method enhances model performance on multiple downstream tasks while reducing data requirements.

cs.CL

Co-Designing Graph-based Approximate Nearest Neighbor Search at Billion Scale for Processing-in-Memory

Approximate Nearest Neighbor Search (ANNS) is a core primitive in modern AI systems, and graph-based methods currently offer the best accuracy-efficiency trade-off at scale. The workload is fundamentally memory-bound: graph traversal produces frequent, irregular memory accesses that cap CPU throughput at main-memory bandwidth, while GPUs lack the high-bandwidth memory capacity to host billion-scale indexes. Processing-in-Memory (PIM) is a natural candidate, as placing computation next to data unlocks the abundant internal bandwidth that such bandwidth-starved workloads demand. Porting graph-based ANNS to PIM, however, exposes several architectural mismatches: each processing unit has only a small local memory, inter-unit communication is costly, host coordination adds overhead, and in-memory compute units are relatively weak -- limitations that have forced prior PIM-based ANNS designs to fall back on cluster-based indexing, whose recall ceiling is far below that of graph methods. This paper presents an algorithm-architecture co-design that overcomes these obstacles through three components: a compacted index layout that shrinks the PIM-resident memory footprint by 14.5x; an asynchronous pipelined scheduler that keeps the host-to-PIM interconnect saturated; and a multiplication-free distance kernel that loses under 0.08% recall. Across three billion-scale benchmarks, the proposed design achieves up to 20x and 17.1x higher throughput than CPU and GPU baselines, respectively, outperforms prior PIM accelerators by 129x in the high-recall regime, and scales gracefully across multi-node deployments and emerging PIM architecture.

cs.AR

Type-III solar radio bursts with spike-like toppings

Spike-type III burst pairs represent a distinct class of solar radio emissions in which clusters of spike-like bursts appear atop the highfrequency onset of type III bursts. Using high time-frequency resolution data from the Chashan Broadband Solar radio spectrometer at meter wavelengths (CBSm), we present the largest statistical study to date of such events, comprising 502 spike-type III pairs from 35 events recorded between November 2023 and October 2025. We find that spike-like clusters systematically precede their associated type III bursts by 0.5-3 s in time (~87% of pairs) and by 3-30 MHz in frequency (~80%), a temporal and spectral offset that differs from earlier reports. The spike-like clusters exhibit diverse morphologies, including point-like, blob-like, drifting, and diffuse structures, with durations of ~0.5-5 s and bandwidths of 15-150 MHz. Bi-directional drifting structures with rates of ~20-100 MHz s-1 are observed, consistent with source motion both toward and away from the Sun. Furthermore, spike emission is predominantly strongly circularly polarized, with more than 64% of clusters showing maximum polarization exceeding 0.6, in stark contrast to the generally weak polarization of type III bursts. These findings point to an origin of the spike radiation in a multiscale, inhomogeneous, and highly dynamic electron-acceleration region, providing novel observational constraints on the mechanisms underlying coherent solar radio bursts.

astro-ph.SR

Integration by Parts Formulas of Mckean-Vlasov SDEs with Jumps and Some Applications

In this article, we establish integration by parts formulas for the solutions of McKean-Vlasov stochastic differential equations with jumps under elliptic coefficients. The derived formulas accommodate both derivatives with respect to real-valued variables and measure-valued variables, interpreted through the Lions' derivative. As applications, we obtain estimates for the derivatives of the density functions of the McKean-Vlasov SDEs, and relying on the integration by parts formulas, we subsequently prove the existence and uniqueness of classical solutions to the associated PDEs with irregular terminal conditions.

math.PR

A Workflow for Evaluating Regional Treatment Effect Heterogeneity in Multi-Regional Clinical Trials

Multi-regional clinical trials (MRCTs) enable efficient global drug development by assessing treatment effects across regions within a single protocol. While powered for overall efficacy, MRCTs are typically not designed to provide confirmatory evidence on regional differences, making an assessment of observed regional heterogeneity largely exploratory and susceptible to sampling variability. Despite this challenge, understanding regional heterogeneity remains important for interpretation and regulatory decision-making. This paper proposes a structured, question-driven framework to guide exploratory assessments of regional heterogeneity in MRCTs. We formulate four key questions to clarify the objectives of such analyses and propose a set of statistical methods to address them. Simulation studies evaluate performance under scenarios with no heterogeneity and heterogeneity driven by observed or unobserved treatment effect modifiers, illustrating how a structured approach can support transparent and cautious interpretation.

stat.AP

HLS-Seek: QoR-Aware Code Generation for High-Level Synthesis via Proxy Comparative Reward Reinforcement Learning

High-Level Synthesis (HLS) compiles algorithmic C/C++ descriptions into hardware, with Quality of Results (QoR)---latency and resource utilization---critically governed by pragma configurations and code structure. Existing natural-language-to-HLS (NL-to-HLS) training approaches prioritize functional correctness while largely ignoring QoR. We observe that reinforcement learning (RL) for HLS does not require absolute synthesis results---only relative comparisons between candidates. Based on this insight, we propose \textbf{HLS-Seek}, a QoR-aware NL-to-HLS framework that avoids full synthesis-in-the-loop RL via a comparative proxy reward model achieving 99.53\% Pareto-dominance accuracy. To prevent reward hacking, we introduce \textit{uncertainty-aware Monte Carlo (MC) dropout switching} that selectively invokes real Vitis HLS synthesis for low-confidence candidates and online updates the proxy, creating a self-improving reward system. HLS-Seek achieves 84.7\% syntax correctness pass@1 and 81.4\% functional correctness pass@5 on HLS-Eval~\cite{abikaram2025hlseval} with only 7B parameters, surpassing GPT-5.1 on functional pass@5, while achieving 8.5$\times$ faster training than real-reward RL. On QoR evaluation, HLS-Seek achieves the lowest latency on 19/30 kernels and Pareto-dominates HLS-specific baselines on 9 kernels.

cs.LG

XtraMAC: An Efficient MAC Architecture for Mixed-Precision LLM Inference on FPGA

The widespread adoption of mixed-precision quantization in large language models (LLMs) has created demand for hardware that can efficiently perform multiply-accumulate (MAC) operations across mixed datatypes and switch datatypes at runtime. Existing FPGA-based MAC solutions fall short due to limitations in fixed-datatype design, inefficient spatial or temporal resource sharing, and poor support for mixed-precision execution. These limitations collectively lead to under-utilization of DSP resources, limiting achievable parallelism and throughput. In this work, we present XtraMAC, a novel MAC architecture that unifies integer, floating-point, and mixed-precision operations within a single, datatype-adaptive microarchitecture. XtraMAC decomposes all supported MAC formats into a shared integer mantissa product with lightweight sign and exponent handling, enabling dynamic operand packing and efficient DSP resource sharing with constant latency and initiation interval of one across all datatypes. Evaluated on an AMD Xilinx U55c FPGA, XtraMAC achieves 1.4-2.0x higher compute density, reduces per-operation LUT, FF, and DSP consumption by 27-51%, and delivers up to 1.9x greater energy efficiency and 1.2x speedup on representative mixed-precision LLM workloads. The implementation of XtraMAC is open-sourced at https://github.com/Xtra-Computing/XtraMAC.

cs.AR

Data-Constrained Modeling of Electron Transport and Asymmetric Precipitation in the 2011 August 4 Solar Flare

Energetic electrons accelerated at coronal reconnection sites during solar flares precipitate into the lower solar atmosphere, generating nonthermal emissions and regulating energy deposition. However, how their transport and precipitation are jointly governed by the three-dimensional (3D) magnetic topology, turbulent scattering, and Coulomb collisions remains unclear. Here, we aim to disentangle these physical processes by using a data-constrained 3D particle transport model for the 2011 August 4 flare. The simulated distribution of precipitated electrons aligns closely with photospheric quasi-separatrix layers and reproduces the observed two-ribbon morphology in 1700~\AA. We reveal a strong polarity asymmetry, with the 10~s precipitation fraction about six times higher in the weak positive polarity. This arises primarily from distinct mirror ratios of different polarities under the 3D magnetic configuration and can be understood via a modified escape probability for an asymmetric magnetic bottle. Varying strengths of turbulent scattering lead to a rise-then-fall trend and a pronounced energy dependence in the precipitation fraction. Coulomb collisions globally suppress precipitation, especially at low energies, and further amplify the polarity asymmetry. This integrated modeling framework bridges detailed transport physics to observable flare emissions and advances the development of quantitative models for realistic solar flare events.

astro-ph.SR