SearcharxivSearch

arXiv subjects

Ziyi Guo

Publications and source records attributed to Ziyi Guo.

At least 19 recordsLinked to original sources

Mountain Muography for China Jinping Underground Laboratory

The China Jinping Underground Laboratory (CJPL), located $\sim 2,400$~m beneath Jinping Mountain, is one of the world's deepest and largest ($\sim 300{,}000~\mathrm{m}^3$) underground facilities, hosting dark matter, nuclear astrophysics, and neutrino experiments. We report the first muon radiography (muography) conducted at this extraordinary depth. Cosmic muons detected by a one-ton prototype developed for the Jinping Neutrino Experiment were used to perform non-invasive subsurface density mapping over a 3~km lateral range. The 1.3~m diameter detector provides nearly isotropic acceptance and an angular resolution of $\sim 4.5^\circ$. By correlating the predicted surface muon flux distributions with the underground measurements, we reconstruct a directional opacity map that constrains the density structure of the overburden and shows excellent agreement with satellite-derived terrain models. This work demonstrates the feasibility of muography at extreme depths with kilometer-scale overburden and establishes a robust methodology for future geophysical applications and large-scale facilities, such as the full Jinping Neutrino Experiment. Based on this validated overburden model, we further predict the total muon fluxes for the eight experimental halls in CJPL-II, providing essential input for their physics programs.

hep-ex

AutoSci: A Memory-Centric Agentic System for the Full Scientific Research Lifecycle

Scientific research has traditionally been human-intensive, requiring researchers to coordinate literature, ideas, experiments, manuscripts, and review responses across long project cycles. The rise of LLM-based scientific agents creates an opportunity to automate this process. Such a system must support the full research lifecycle, maintain structured persistent memory across projects, and improve its own research procedures over time. However, existing systems either partially satisfy or fail to satisfy these requirements, leaving a gap for a unified automated scientific research system. As a result, we present AutoSci, a memory-centric agentic system for the full scientific research lifecycle. AutoSci is organized around four modules. SciMem provides schema-governed research memory, separating Long-Term Knowledge Memory for reusable scientific knowledge from Active Research Memory for project-level artifacts such as ideas, experiments, manuscripts, and reviews. SciFlow executes a five-stage lifecycle from literature understanding to rebuttal through a harness that controls state, context, verification, feedback, and orchestration. SciDAG augments difficult skills with DAG-shaped multi-agent operators and reusable stage-specific templates. SciEvolve converts feedback signals from users, experiments, reviews, and external environments into versioned updates to SciMem organization, SciFlow skills, and SciDAG templates. Together, these modules make AutoSci a persistent research environment that can execute, remember, and evolve across research projects. The code repository is available at https://github.com/skyllwt/AutoSci.

cs.AI

Limited imprint of high-mass IMF variations on sodium abundances in main-sequence galaxies

Growing evidence suggests that the stellar initial mass function (IMF) varies systematically across galaxies, deviating from the canonical Milky Way form. Such variations would modify the integrated nucleosynthetic yields, and hence the abundance patterns used in stellar population synthesis studies. How these could impact, in particular, the sodium abundance (and sodium-to-oxygen ratios) in star-forming galaxies is not well understood. In this work, we systematically study how high-mass IMF variations affect sodium enrichment using a one-zone galactic chemical evolution model. The model incorporates star formation histories from semi-analytic simulations and is calibrated to match the observed galaxy mass--metallicity relation. We find that varying the IMF high-mass end (and the IMF slope) could only alter the sodium abundance by less than 0.1 dex, across galaxies with stellar masses from $10^9\,\mathrm{M}_\odot$ to $10^{11}\,\mathrm{M}_\odot$. This result is robust under different stellar models and galaxy evolution assumptions, primarily because sodium production is similar to that of oxygen. We conclude that sodium abundance is largely insensitive to changes in the high-mass IMF, unlikely to compromise the use of sodium indices as IMF diagnostics in stellar population studies.

astro-ph.GA

Vacuum-Sealed Thermal Treatment Regulates Trap States and Red Persistent Luminescence in CaTiO3: Pr3+,Al3+ Phosphors

Red persistent phosphors remain less mature than green and blue-green systems because their afterglow is often weak and decays rapidly. Here, CaTiO3:0.3%Pr3+,0.3%Al3+ phosphors were prepared by a high-temperature solid-state route under either air or vacuum-sealed quartz-tube (QT) conditions. The effects of processing atmosphere and sintering temperature on phase structure, microstructure, steady-state photoluminescence, afterglow, thermoluminescence, and excited-state decay were examined. X-ray diffraction and Raman spectra show that all samples retain the orthorhombic CaTiO3 perovskite phase, with no detectable secondary phase. SEM observations show particle coarsening at higher QT temperatures, while EDS mapping indicates a homogeneous distribution of Ca, Ti, O, Pr, and Al within the examined region. The QT-treated samples exhibit stronger Pr3+ red emission near 612 nm and markedly improved afterglow compared with the air-treated sample. The QT-1300 sample shows the best afterglow among the present samples, with a reported 3.6-fold higher intensity at 1200 s than Air-1200{\deg}C. Thermoluminescence results indicate that QT treatment increases the population of thermally active traps and enhances the deeper trap component. These results suggest that a low-oxygen sealed environment regulates defect-related traps, most likely involving oxygen-vacancy-associated centers, and improves carrier storage and release through the Pr3+ and Ti4+ intervalence charge-transfer pathway. This work provides a practical processing strategy for improving CaTiO3-based red persistent phosphors and offers insight into trap-state regulation under low-oxygen thermal treatment.

physics.optics

Can Large Language Models Revolutionize Survey Research? Experiments with Disaster Preparedness Responses

Survey research faces mounting structural challenges: declining response rates, sample bias, block-wise missingness among at-risk respondents, and AI-assisted fraudulent completions in online panels. Large language models (LLMs) have been proposed as a remedy, yet rigorous evaluations across the full survey workflow remain scarce, particularly in disaster contexts where data quality matters most. We present and evaluate a five-stage framework for LLM integration covering questionnaire design, sample selection, pilot testing, missing-data imputation, and post-collection analysis, using the 2024 Hurricane Milton preparedness survey of Florida residents (n=946) as a shared empirical testbed. We introduce a Protection Motivation Theory (PMT)-constrained co-occurrence knowledge graph and develop seven LLM configurations spanning zero-shot inference, retrieval-augmented baselines, and novel theory-informed variants. Our proposed Anchored Marginal Theory-Informed LLM (A-TLM) outperforms all three classical imputation baselines (IPW/MI, MICE+PMM, missForest) on RMSE under disaster-relevant block-wise MNAR conditions (S4 RMSE 1.439 vs. 1.496 for the next-best), while achieving near-zero signed bias (-0.121) where the random-forest imputer produces the largest absolute bias (-0.631). Organizing retrieval around PMT causal structure and integrating all evidence in a single model call outperforms unstructured retrieval and staged sequential inference (MAE 0.993 vs. 1.097 for standard RAG). We document that near-zero aggregate bias can mask opposing subgroup errors and propose subgroup-stratified bias auditing as a reporting standard. A retrieval-constrained knowledge-graph chatbot demonstrates that hallucination is architecturally manageable through grounded refusal.

cs.AI

NarraScore: Bridging Visual Narrative and Musical Dynamics via Hierarchical Affective Control

Synthesizing coherent soundtracks for long-form videos remains a formidable challenge, currently stalled by three critical impediments: computational scalability, temporal coherence, and, most critically, a pervasive semantic blindness to evolving narrative logic. To bridge these gaps, we propose NarraScore, a hierarchical framework predicated on the core insight that emotion serves as a high-density compression of narrative logic. Uniquely, we repurpose frozen Vision-Language Models (VLMs) as continuous affective sensors, distilling high-dimensional visual streams into dense, narrative-aware Valence-Arousal trajectories. Mechanistically, NarraScore employs a Dual-Branch Injection strategy to reconcile global structure with local dynamism: a \textit{Global Semantic Anchor} ensures stylistic stability, while a surgical \textit{Token-Level Affective Adapter} modulates local tension via direct element-wise residual injection. This minimalist design bypasses the bottlenecks of dense attention and architectural cloning, effectively mitigating the overfitting risks associated with data scarcity. Experiments demonstrate that NarraScore achieves state-of-the-art consistency and narrative alignment with negligible computational overhead, establishing a fully autonomous paradigm for long-video soundtrack generation.

cs.SD

DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI

The rapidly growing demand for high-quality data in Large Language Models (LLMs) has intensified the need for scalable, reliable, and semantically rich data preparation pipelines. However, current practices remain dominated by ad-hoc scripts and loosely specified workflows, which lack principled abstractions, hinder reproducibility, and offer limited support for model-in-the-loop data generation. To address these challenges, we present DataFlow, a unified and extensible LLM-driven data preparation framework. DataFlow is designed with system-level abstractions that enable modular, reusable, and composable data transformations, and provides a PyTorch-style pipeline construction API for building debuggable and optimizable dataflows. The framework consists of nearly 200 reusable operators and six domain-general pipelines spanning text, mathematical reasoning, code, Text-to-SQL, agentic RAG, and large-scale knowledge extraction. To further improve usability, we introduce DataFlow-Agent, which automatically translates natural-language specifications into executable pipelines via operator synthesis, pipeline planning, and iterative verification. Across six representative use cases, DataFlow consistently improves downstream LLM performance. Our math, code, and text pipelines outperform curated human datasets and specialized synthetic baselines, achieving up to +3\% execution accuracy in Text-to-SQL over SynSQL, +7\% average improvements on code benchmarks, and 1--3 point gains on MATH, GSM8K, and AIME. Moreover, a unified 10K-sample dataset produced by DataFlow enables base models to surpass counterparts trained on 1M Infinity-Instruct data. These results demonstrate that DataFlow provides a practical and high-performance substrate for reliable, reproducible, and scalable LLM data preparation, and establishes a system-level foundation for future data-centric AI development.

cs.LG

Paper2SysArch: Structure-Constrained System Architecture Generation from Scientific Papers

The manual creation of system architecture diagrams for scientific papers is a time-consuming and subjective process, while existing generative models lack the necessary structural control and semantic understanding for this task. A primary obstacle hindering research and development in this domain has been the profound lack of a standardized benchmark to quantitatively evaluate the automated generation of diagrams from text. To address this critical gap, we introduce a novel and comprehensive benchmark, the first of its kind, designed to catalyze progress in automated scientific visualization. It consists of 3,000 research papers paired with their corresponding high-quality ground-truth diagrams and is accompanied by a three-tiered evaluation metric assessing semantic accuracy, layout coherence, and visual quality. Furthermore, to establish a strong baseline on this new benchmark, we propose Paper2SysArch, an end-to-end system that leverages multi-agent collaboration to convert papers into structured, editable diagrams. To validate its performance on complex cases, the system was evaluated on a manually curated and more challenging subset of these papers, where it achieves a composite score of 69.0. This work's principal contribution is the establishment of a large-scale, foundational benchmark to enable reproducible research and fair comparison. Meanwhile, our proposed system serves as a viable proof-of-concept, demonstrating a promising path forward for this complex task.

cs.AI

Investigating Production of TeV-scale Muons in Extensive Air Shower at 2400 Meters Underground

Deep underground experiments present a new avenue to probe the first interactions in extensive air showers or hadronic interactions in the extreme forward phase space. The China Jinping Underground Laboratory, characterized by a vertical rock overburden of 2,400~m, provides an exceptionally effective shield against cosmic muons with energies below 3~TeV. The surviving high-energy muons, produced in the first interactions of extensive air showers, open a unique observational window into primary cosmic rays from tens of TeV up to the PeV scale and beyond. This distinctive feature also enables detailed studies of charged hadron production in the earliest stages of shower development. Using 1,338.6 live days of data collected with a one-ton prototype detector for the Jinping Neutrino Experiment, we measured the underground muon flux originating from air showers. The results show discrepancies of about 40\% corresponding to significances of more than 2$\sigma$, relative to predictions from several leading hadronic interaction models. We interpret these findings from two complementary perspectives: (i) by adopting the expected cosmic-ray spectra, we constrain the modeling of the first hadronic interactions in air showers and provide novel insights into resolving the long-standing \textit{muon puzzle}; and (ii) by assuming specific hadronic interaction models, we infer the mass composition of cosmic rays, and our data favor a lighter component in the corresponding energy range. Our study demonstrates the potential of deep underground laboratories to provide new experimental insights into air shower physics and cosmic rays.

hep-ex

Stellar population astrophysics (SPA) with the TNG. The Phosphorus abundance on the young side of MilkyWay

We present phosphorus abundance measurements for a total of 102 giant stars, including 82 stars in 24 open clusters and 20 Cepheids, based on high-resolution near-infrared spectra obtained with GIANO-B. Evolution of phosphorus abundance, despite its astrophysical and biological significance, remains poorly understood due to a scarcity of observational data. By combining precise stellar parameters from the optical, a robust line selection and measurement method, we measure phosphorus abundances using available P I lines. Our analysis confirms a declining trend in [P/Fe] with increasing [Fe/H] around solar metallicity for clusters and Cepheids, consistent with previous studies. We also report a [P/Fe]-age relation among open clusters older than 1 Gyr, indicating a time-dependent enrichment pattern. Such pattern can be explained by the different stellar formation history of their parental gas, with more efficient stellar formation in the gas of older clusters (thus with higher phosphorus abundances). [P/Fe] shows a flat trend among cepheids and clusters younger than 1 Gyr (along with three Cepheids inside open clusters), possibly hinting at the phosphorus contribution from the previous-generation low-mass stars. Such trend suggests that the young clusters share a nearly common chemical history, with a mild increase in phosphorus production by low-mass stars.

astro-ph.SR

Massive Star Formation at Supersolar Metallicities: Constraints on the Initial Mass Function

Metals enhance the cooling efficiency of molecular clouds, promoting fragmentation. Consequently, increasing the metallicity may boost the formation of low-mass stars. Within the integrated galaxy initial mass function (IGIMF) theory, this effect is empirically captured by a linear relation between the slope of the low-mass stellar IMF, $α_1$, and the metal mass fraction, $Z$. This linear $α_1$-$Z$ relation has been calibrated up to $\approx 2 \, Z_{\odot}$, though higher metallicity environments are known to exist. We show that if the linear $α_1$-$Z$ relation extends to higher metallicities ($[Z] \gtrsim 0.5$), massive star formation is suppressed entirely. Alternatively, fragmentation efficiency may saturate beyond some metallicity threshold if gravitational collapse cascades rapidly enough. To model this behavior, we propose a logistic function describing the transition from metallicity-sensitive to metallicity-insensitive fragmentation regimes. We provide a user-friendly public code, pyIGIMF, which enables the instantaneous computation of the IGIMF theory with the logistic $α_1$-$Z$ relation.

astro-ph.GA

Time Profile of U.S. Neighborhoods: Datasets of Time Use at Social Infrastructure Places

Social infrastructure plays a critical role in shaping neighborhood well-being by fostering social and cultural interaction, enabling service provision, and encouraging exposure to diverse environments. Despite the growing knowledge of its spatial accessibility, time use at social infrastructure places is underexplored due to the lack of a spatially resolved national dataset. We address this gap by developing scalable Social-Infrastructure Time Use measures (STU) that capture length and depth of engagement, activity diversity, and spatial inequality, supported by first-of-their-kind datasets spanning multiple geographic scales from census tracts to metropolitan areas. Our datasets leverage anonymized and aggregated foot traffic data collected between 2019 and 2024 across 49 continental U.S. states. The data description reveals variances in STU across time, space, and differing neighborhood sociodemographic characteristics. Validation demonstrates generally robust population representation, consistent with established national survey findings while revealing more nuanced patterns. Future analyses could link STU with public health outcomes and environmental factors to inform targeted interventions aimed at enhancing population well-being and guiding social infrastructure planning and usage.

cs.SI

Urban-STA4CLC: Urban Theory-Informed Spatio-Temporal Attention Model for Predicting Post-Disaster Commercial Land Use Change

Natural disasters such as hurricanes and wildfires increasingly introduce unusual disturbance on economic activities, which are especially likely to reshape commercial land use pattern given their sensitive to customer visitation. However, current modeling approaches are limited in capturing such complex interplay between human activities and commercial land use change under and following disturbances. Such interactions have been more effectively captured in current resilient urban planning theories. This study designs and calibrates a Urban Theory-Informed Spatio-Temporal Attention Model for Predicting Post-Disaster Commercial Land Use Change (Urban-STA4CLC) to predict both the yearly decline and expansion of commercial land use at census block level under cumulative impact of disasters on human activities over two years. Guided by urban theories, Urban-STA4CLC integrates both spatial and temporal attention mechanisms with three theory-informed modules. Resilience theory guides a disaster-aware temporal attention module that captures visitation dynamics. Spatial economic theory informs a multi-relational spatial attention module for inter-block representation. Diffusion theory contributes a regularization term that constrains land use transitions. The model performs significantly better than non-theoretical baselines in predicting commercial land use change under the scenario of recurrent hurricanes, with around 19% improvement in F1 score (0.8763). The effectiveness of the theory-guided modules was further validated through ablation studies. The research demonstrates that embedding urban theory into commercial land use modeling models may substantially enhance the capacity to capture its gains and losses. These advances in commercial land use modeling contribute to land use research that accounts for cumulative impacts of recurrent disasters and shifts in economic activity patterns.

cs.CY

Breaking the Bulkhead: Demystifying Cross-Namespace Reference Vulnerabilities in Kubernetes Operators

Kubernetes Operators, automated tools designed to manage application lifecycles within Kubernetes clusters, extend the functionalities of Kubernetes, and reduce the operational burden on human engineers. While Operators significantly simplify DevOps workflows, they introduce new security risks. In particular, Kubernetes enforces namespace isolation to separate workloads and limit user access, ensuring that users can only interact with resources within their authorized namespaces. However, Kubernetes Operators often demand elevated privileges and may interact with resources across multiple namespaces. This introduces a new class of vulnerabilities, the Cross-Namespace Reference Vulnerability. The root cause lies in the mismatch between the declared scope of resources and the implemented scope of the Operator logic, resulting in Kubernetes being unable to properly isolate the namespace. Leveraging such vulnerability, an adversary with limited access to a single authorized namespace may exploit the Operator to perform operations affecting other unauthorized namespaces, causing Privilege Escalation and further impacts. To the best of our knowledge, this paper is the first to systematically investigate Kubernetes Operator attacks. We present Cross-Namespace Reference Vulnerability with two strategies, demonstrating how an attacker can bypass namespace isolation. Through large-scale measurements, we found that over 14% of Operators in the wild are potentially vulnerable. Our findings have been reported to the relevant developers, resulting in 8 confirmations and 7 CVEs by the time of submission, affecting vendors including Red Hat and NVIDIA, highlighting the critical need for enhanced security practices in Kubernetes Operators. To mitigate it, we open-source the static analysis suite and propose concrete mitigation to benefit the ecosystem.

cs.CR

Take a Step Further: Understanding Page Spray in Linux Kernel Exploitation

Recently, a novel method known as Page Spray emerges, focusing on page-level exploitation for kernel vulnerabilities. Despite the advantages it offers in terms of exploitability, stability, and compatibility, comprehensive research on Page Spray remains scarce. Questions regarding its root causes, exploitation model, comparative benefits over other exploitation techniques, and possible mitigation strategies have largely remained unanswered. In this paper, we conduct a systematic investigation into Page Spray, providing an in-depth understanding of this exploitation technique. We introduce a comprehensive exploit model termed the \sys model, elucidating its fundamental principles. Additionally, we conduct a thorough analysis of the root causes underlying Page Spray occurrences within the Linux Kernel. We design an analyzer based on the Page Spray analysis model to identify Page Spray callsites. Subsequently, we evaluate the stability, exploitability, and compatibility of Page Spray through meticulously designed experiments. Finally, we propose mitigation principles for addressing Page Spray and introduce our own lightweight mitigation approach. This research aims to assist security researchers and developers in gaining insights into Page Spray, ultimately enhancing our collective understanding of this emerging exploitation technique and making improvements to the community.

cs.CR

Dual-branch PolSAR Image Classification Based on GraphMAE and Local Feature Extraction

The annotation of polarimetric synthetic aperture radar (PolSAR) images is a labor-intensive and time-consuming process. Therefore, classifying PolSAR images with limited labels is a challenging task in remote sensing domain. In recent years, self-supervised learning approaches have proven effective in PolSAR image classification with sparse labels. However, we observe a lack of research on generative selfsupervised learning in the studied task. Motivated by this, we propose a dual-branch classification model based on generative self-supervised learning in this paper. The first branch is a superpixel-branch, which learns superpixel-level polarimetric representations using a generative self-supervised graph masked autoencoder. To acquire finer classification results, a convolutional neural networks-based pixel-branch is further incorporated to learn pixel-level features. Classification with fused dual-branch features is finally performed to obtain the predictions. Experimental results on the benchmark Flevoland dataset demonstrate that our approach yields promising classification results.

cs.CV

PRTGS: Precomputed Radiance Transfer of Gaussian Splats for Real-Time High-Quality Relighting

We proposed Precomputed RadianceTransfer of GaussianSplats (PRTGS), a real-time high-quality relighting method for Gaussian splats in low-frequency lighting environments that captures soft shadows and interreflections by precomputing 3D Gaussian splats' radiance transfer. Existing studies have demonstrated that 3D Gaussian splatting (3DGS) outperforms neural fields' efficiency for dynamic lighting scenarios. However, the current relighting method based on 3DGS still struggles to compute high-quality shadow and indirect illumination in real time for dynamic light, leading to unrealistic rendering results. We solve this problem by precomputing the expensive transport simulations required for complex transfer functions like shadowing, the resulting transfer functions are represented as dense sets of vectors or matrices for every Gaussian splat. We introduce distinct precomputing methods tailored for training and rendering stages, along with unique ray tracing and indirect lighting precomputation techniques for 3D Gaussian splats to accelerate training speed and compute accurate indirect lighting related to environment light. Experimental analyses demonstrate that our approach achieves state-of-the-art visual quality while maintaining competitive training times and allows high-quality real-time (30+ fps) relighting for dynamic light and relatively complex scenes at 1080p resolution.

cs.CV

CAMP: Compiler and Allocator-based Heap Memory Protection

The heap is a critical and widely used component of many applications. Due to its dynamic nature, combined with the complexity of heap management algorithms, it is also a frequent target for security exploits. To enhance the heap's security, various heap protection techniques have been introduced, but they either introduce significant runtime overhead or have limited protection. We present CAMP, a new sanitizer for detecting and capturing heap memory corruption. CAMP leverages a compiler and a customized memory allocator. The compiler adds boundary-checking and escape-tracking instructions to the target program, while the memory allocator tracks memory ranges, coordinates with the instrumentation, and neutralizes dangling pointers. With the novel error detection scheme, CAMP enables various compiler optimization strategies and thus eliminates redundant and unnecessary check instrumentation. This design minimizes runtime overhead without sacrificing security guarantees. Our evaluation and comparison of CAMP with existing tools, using both real-world applications and SPEC CPU benchmarks, show that it provides even better heap corruption detection capability with lower runtime overhead.

cs.CR