SearcharxivSearch

arXiv subjects

Yunlong Zhang

Publications and source records attributed to Yunlong Zhang.

At least 19 recordsLinked to original sources

Development of Neutron Transmutation Doped Germanium (NTD-Ge) for Cryogenic Applications

This paper presents the systematic fabrication and characterization of cryogenic thermometers based on neutron transmutation-doped germanium (NTD-Ge). High-purity (10N) germanium samples were irradiated by thermal neutrons with different fluences at the China Advanced Research Reactor (CARR). After irradiation and a six-month cooling-down period, positron annihilation lifetime spectroscopy and temperature-dependent Hall effect measurements were performed to characterize irradiation-induced defects and carrier concentrations in the NTD-Ge samples. Utilizing standard semiconductor fabrication techniques, point electrodes were deposited onto the processed samples to fabricate functional NTD-Ge cryogenic thermometers. The low-temperature resistance performance of the devices was characterized down to 20 mK on a millikelvin range cryogenic test platform. The measured temperature dependence of resistance follows Mott's law, showing excellent agreement across the full measured range. The extracted T0 is consistent with expectations. These results collectively verified both the applicability of the thermometers in cryogenic system and the reliability of the fabrication procedure.

physics.ins-det

CausalGaze: Unveiling Hallucinations via Counterfactual Graph Intervention in Large Language Models

Despite the groundbreaking advancements made by large language models (LLMs), hallucination remains a critical bottleneck for their deployment in high-stakes domains. Existing classification-based methods mainly rely on static and passive signals from internal states, which often captures the noise and spurious correlations, while overlooking the underlying causal mechanisms. To address this limitation, we shift the paradigm from passive observation to active intervention by introducing CausalGaze, a novel hallucination detection framework based on structural causal models (SCMs). CausalGaze models LLMs' internal states as dynamic causal graphs and employs counterfactual interventions to disentangle causal reasoning paths from incidental noise, thereby enhancing model interpretability. Extensive experiments across four datasets and three widely used LLMs demonstrate the effectiveness of CausalGaze, especially achieving 3.3% improvement in AUROC on the TruthfulQA dataset compared to state-of-the-art baselines.

cs.LG

E-biofuels reduce the cost of achieving emissions targets in hard-to-electrify sectors

Renewable liquid fuels are essential for achieving emissions targets for hard-to-electrify sectors such as aviation and shipping. While biofuels and synthetic e-fuels have been well-studied, e-biofuels, produced by adding renewable hydrogen to biomass conversion to better utilise the biogenic carbon, remain understudied and lack a clear role in EU fuel regulations. In this paper, using a sector-coupled European energy system model, we find that e-biofuels are cost-effective to meet stringent emissions targets if biomass availability is limited and fossil fuels are ineligible, either due to limited carbon sequestration capacity or to high renewable fuel mandates. By directly increasing utilisation of biogenic carbon instead of synthesising fuels based on captured $CO_2$, there are savings from fuel production and carbon capture that reduce total system costs by up to 2.7% and liquid fuel costs by more than 10%. Our results highlight the role of e-biofuels as a potential hedge against uncertainty in biomass, hydrogen, and carbon storage availability, as well as evolving policy implementation.

physics.soc-ph

A Statistical Framework for Auditing Behavioral Dependence and Induced Bias in LLM Judges

The rapid growth of the large language model (LLM) ecosystem raises a critical question: are seemingly diverse models truly independent? Shared pretraining data, distillation, and alignment pipelines can induce hidden behavioral dependencies, or latent entanglement, that undermine multi-model systems such as LLM-as-a-judge pipelines and ensemble verification, which implicitly assume independent signals. In practice, this manifests as correlated reasoning patterns and synchronized failures, where apparent agreement reflects shared error modes rather than independent validation. To address this, we develop a statistical framework for auditing behavioral entanglement among black-box LLMs. Our approach introduces a multi-resolution hierarchy that characterizes the joint failure manifold through two information-theoretic metrics: (i) a Difficulty-Weighted Behavioral Entanglement Index (BEI), which amplifies synchronized failures on easy tasks, and (ii) a Cumulative Information Gain (CIG) metric, which captures directional alignment in erroneous responses. Through experiments on 18 LLMs from six model families, we identify statistically significant behavioral entanglement. Such behavioral dependence is associated with judge over-endorsement bias on a disjoint MMLU-Pro evaluation set (rho = 0.508 for BEI and rho = 0.520 for CIG; p < 0.01). The association further transfers to the MATH-500 benchmark (rho = 0.441 for BEI and rho = 0.457 for CIG; p < 0.05), providing cross-benchmark evidence that the identified dependency structure generalizes beyond the data and response format used for its estimation. Finally, we demonstrate a practical use case of entanglement through de-entangled verifier ensemble reweighting, achieving 3.5 and 2.6 percentage-point gains in accuracy and precision, respectively, over majority voting.

cs.AI

A Method for On-Orbit Calibration of the VLAST-P Electromagnetic Calorimeter

The Very Large Area Gamma-ray Space Telescope Pathfinder (VLAST-P), as the technology validation satellite for the VLAST mission, is designed to observe high-energy solar bursts on orbit. The CsI electromagnetic calorimeter (ECAL) is one of the key sub-detectors of VLAST-P. To investigate the on-orbit energy calibration method of the ECAL, a Geant4-based simulation of VLAST-P was carried out. The results show an energy resolution better than 10% in the 0.1 to 5 GeV range and a linearity deviation below 2%. A dedicated minimum-ionization-particle (MIP) calibration method was developed to ensure accurate energy reconstruction and to monitor detector stability throughout the in-orbit calibration period.

hep-ex

Towards Effective and Efficient Context-aware Nucleus Detection in Histopathology Whole Slide Images

Nucleus detection in histopathology whole slide images (WSIs) is crucial for a broad spectrum of clinical applications. The gigapixel size of WSIs necessitates the use of sliding window methodology for nucleus detection. However, mainstream methods process each sliding window independently, which overlooks broader contextual information and easily leads to inaccurate predictions. To address this limitation, recent studies additionally crop a large Filed-of-View (LFoV) patch centered on each sliding window to extract contextual features. However, such methods substantially increase whole-slide inference latency. In this work, we propose an effective and efficient context-aware nucleus detection approach. Specifically, instead of using LFoV patches, we aggregate contextual clues from off-the-shelf features of historically visited sliding windows, which greatly enhances the inference efficiency. Moreover, compared to LFoV patches used in previous works, the sliding window patches have higher magnification and provide finer-grained tissue details, thereby enhancing the classification accuracy. To develop the proposed context-aware model, we utilize annotated patches along with their surrounding unlabeled patches for training. Beyond exploiting high-level tissue context from these surrounding regions, we design a post-training strategy that leverages abundant unlabeled nucleus samples within them to enhance the model's context adaptability. Extensive experimental results on three challenging benchmarks demonstrate the superiority of our method.

eess.IV

Unveiling Traffic Wave of Linear Adaptive Cruise Control: A Second-order Macroscopic Traffic Flow Model

Traffic waves, the spatiotemporal propagation of congestion, are a key feature of traffic flow. As Adaptive Cruise Control (ACC) systems gain widespread adoption and show promise for improving both efficiency and safety, understanding how these waves evolve under ACC becomes increasingly important. Yet most existing analyses rely on steady-state metrics (e.g., equilibrium spacing) and neglect the ACC control-law parameters, such as feedback gains, that fundamentally shape higher-order traffic dynamics. To overcome this limitation, we embed the ACC control law directly into the momentum equation while retaining mass conservation law. The result is a higher-order macroscopic model whose dynamics are governed by a second-order partial differential equation equivalent to the linear ACC feedback law. Analyzing the flux Jacobian confirms that the system is strictly hyperbolic, thereby preserving anisotropy and ensuring physical consistency. The derivation also shows that traffic wave evolution depends on both the initial state and the ACC control parameters. We analyze wave-propagation characteristics, linear degeneracy, admissible discontinuities, and their connection to ACC string stability, with the corresponding derivations. Numerical experiments confirm that the second-order model yields markedly lower vehicle-pair speed deviations along wave paths than a first-order model subject to the same non-steady disturbances, underscoring both the necessity of a second-order treatment and the soundness of the proposed framework.

math.AP

The development of a high granular crystal calorimeter prototype of VLAST

Very Large Area gamma-ray Space Telescope (VLAST) is the next-generation flagship space observatory for high-energy gamma-ray detection proposed by China. The observation energy range covers from MeV to TeV and beyond, with acceptance of 10 m^2sr. The calorimeter serves as a crucial subdetector of VLAST, responsible for high-precision energy measurement and electron/proton discrimination. This discrimination capability is essential for accurately identifying gamma-ray events among the background of charged particles. To accommodate such an extensive energy range, a high dynamic range readout scheme employing dual avalanche photodiodes (APDs) has been developed, achieving a remarkable dynamic range of 10^6. Furthermore, a high granular prototype based on bismuth germanate (BGO) cubic scintillation crystals has been developed. This high granularity enables detailed imaging of the particle showers, improving both energy resolution and particle identification. The prototype's performance is evaluated through cosmic ray testing, providing valuable data for optimizing the final calorimeter design for VLAST.

physics.ins-det

Knowledge-Grounded Agentic Large Language Models for Multi-Hazard Understanding from Reconnaissance Reports

Post-disaster reconnaissance reports contain critical evidence for understanding multi-hazard interactions, yet their unstructured narratives make systematic knowledge transfer difficult. Large language models (LLMs) offer new potential for analyzing these reports, but often generate unreliable or hallucinated outputs when domain grounding is absent. This study introduces the Mixture-of-Retrieval Agentic RAG (MoRA-RAG), a knowledge-grounded LLM framework that transforms reconnaissance reports into a structured foundation for multi-hazard reasoning. The framework integrates a Mixture-of-Retrieval mechanism that dynamically routes queries across hazard-specific databases while using agentic chunking to preserve contextual coherence during retrieval. It also includes a verification loop that assesses evidence sufficiency, refines queries, and initiates targeted searches when information remains incomplete. We construct HazardRecQA by deriving question-answer pairs from GEER reconnaissance reports, which document 90 global events across seven major hazard types. MoRA-RAG achieves up to 94.5 percent accuracy, outperforming zero-shot LLMs by 30 percent and state-of-the-art RAG systems by 10 percent, while reducing hallucinations across diverse LLM architectures. MoRA-RAG also enables open-weight LLMs to achieve performance comparable to proprietary models. It establishes a new paradigm for transforming post-disaster documentation into actionable, trustworthy intelligence for hazard resilience.

cs.CL

CyPortQA: Benchmarking Multimodal Large Language Models for Cyclone Preparedness in Port Operation

As tropical cyclones intensify and track forecasts become increasingly uncertain, U.S. ports face heightened supply-chain risk under extreme weather conditions. Port operators need to rapidly synthesize diverse multimodal forecast products, such as probabilistic wind maps, track cones, and official advisories, into clear, actionable guidance as cyclones approach. Multimodal large language models (MLLMs) offer a powerful means to integrate these heterogeneous data sources alongside broader contextual knowledge, yet their accuracy and reliability in the specific context of port cyclone preparedness have not been rigorously evaluated. To fill this gap, we introduce CyPortQA, the first multimodal benchmark tailored to port operations under cyclone threat. CyPortQA assembles 2,917 realworld disruption scenarios from 2015 through 2023, spanning 145 U.S. principal ports and 90 named storms. Each scenario fuses multisource data (i.e., tropical cyclone products, port operational impact records, and port condition bulletins) and is expanded through an automated pipeline into 117,178 structured question answer pairs. Using this benchmark, we conduct extensive experiments on diverse MLLMs, including both open-source and proprietary model. MLLMs demonstrate great potential in situation understanding but still face considerable challenges in reasoning tasks, including potential impact estimation and decision reasoning.

cs.CL

Development of the CEPC analog hadron calorimeter prototype

The Circular Electron Positron Collider (CEPC) is a next-generation electron$-$positron collider proposed for the precise measurement of the properties of the Higgs boson. To emphasize boson separation and jet reconstruction, the baseline design of the CEPC detector was guided by the particle flow algorithm (PFA) concept. As one of the calorimeter options, the analogue hadron calorimeter (AHCAL) was proposed. The CEPC AHCAL comprises a 40-layer sandwich structure using steel plates as absorbers and scintillator tiles coupled with silicon photomultipliers (SiPM) as sensitive units. To validate the feasibility of the AHCAL option, a series of studies were conducted to develop a prototype. This AHCAL prototype underwent an electronic test and a cosmic ray test to assess its performance and ensure it was ready for three beam tests performed in 2022 and 2023. The test beam data is currently under analysis, and the results are expected to deepen our understanding of hadron showers, validate the concept of Particle Flow Algorithm (PFA), and ultimately refine the design of the CEPC detector.

physics.ins-det

Modeling Headway in Heterogeneous and Mixed Traffic Flow: A Statistical Distribution Based on a General Exponential Function

The ability of existing headway distributions to accurately reflect the diverse behaviors and characteristics in heterogeneous traffic (different types of vehicles) and mixed traffic (human-driven vehicles with autonomous vehicles) is limited, leading to unsatisfactory goodness of fit. To address these issues, we modified the exponential function to obtain a novel headway distribution. Rather than employing Euler's number (e) as the base of the exponential function, we utilized a real number base to provide greater flexibility in modeling the observed headway. However, the proposed is not a probability function. We normalize it to calculate the probability and derive the closed-form equation. In this study, we utilized a comprehensive experiment with five open datasets: highD, exiD, NGSIM, Waymo, and Lyft to evaluate the performance of the proposed distribution and compared its performance with six existing distributions under mixed and heterogeneous traffic flow. The results revealed that the proposed distribution not only captures the fundamental characteristics of headway distribution but also provides physically meaningful parameters that describe the distribution shape of observed headways. Under heterogeneous flow on highways (i.e., uninterrupted traffic flow), the proposed distribution outperforms other candidate distributions. Under urban road conditions (i.e., interrupted traffic flow), including heterogeneous and mixed traffic, the proposed distribution still achieves decent results.

stat.AP

Determination of the absolute energy scale of the DAMPE calorimeter with the geomagnetic rigidity cutoff method

The Dark Matter Particle Explorer (DAMPE) is a satellite-borne detector designed to detect high-energy cosmic ray particles with its core component being a BGO calorimeter capable of measuring energies from $\sim$GeV to $O(100)$ TeV. The 32 radiation lengths thickness of the calorimeter is designed to ensure full containment of showers produced by cosmic ray electrons and positrons (CREs) and $γ$-rays at energies below tens of TeV, providing high resolution in energy measurements. The absolute energy scale therefore becomes a crucial parameter for precise measurements of the CRE energy spectrum. The geomagnetic field induces a rapid drop in the low energy spectrum of electrons and positrons, a phenomenon that provides a method to determine the calorimeter's absolute energy scale. By comparing the cutoff energies of the measured spectra of CREs with those expected from the International Geomagnetic Reference Field model across 4 McIlwain $L$ bins - which cover most regions of the DAMPE orbit - we find that the calorimeter's absolute energy scale exceeds the calibration based on Geant4 simulation by $1.013\pm0.012_{\rm stat}\pm0.026_{\rm sys}$ for energies between 7 GeV and 16 GeV. The absolute energy scale should be taken into account when comparing the absolute CREs fluxes among different detectors.

hep-ex

U.S. Port Disruptions under Tropical Cyclones: Resilience Analysis by Harnessing Multiple-Source Dataset

This study introduces the CyPort Dataset, recording disruptions to 145 U.S. principal ports and freight network from 90 tropical cyclones (2015-2023). It addresses limitations of event specific resilience studies and provides a comprehensive dataset for broader analysis. To account for excess zeros and unobserved heterogeneity in disruption outcomes, the Random Parameter Negative Binomial Lindley (RPNB Lindley) model is employed to produce more reliable resilience insights. The model demonstrates improved fit over traditional methods and uncovers variation in how features such as wind speed, storm surge height, rainfall, and distance to cyclone influence disruption outcomes across ports. This analysis reveals a tipping point at Saffir Simpson Hurricane Category 4, where disruptions escalate sharply, causing greater impacts and prolonged recovery. Regionally, ports along the Gulf of America show greatest vulnerability. Within the freight network, ports with high betweenness centrality are more resilient, while transshipment and local hubs are more fragile.

stat.AP

AEM: Attention Entropy Maximization for Multiple Instance Learning based Whole Slide Image Classification

Multiple Instance Learning (MIL) effectively analyzes whole slide images but faces overfitting due to attention over-concentration. While existing solutions rely on complex architectural modifications or additional processing steps, we introduce Attention Entropy Maximization (AEM), a simple yet effective regularization technique. Our investigation reveals the positive correlation between attention entropy and model performance. Building on this insight, we integrate AEM regularization into the MIL framework to penalize excessive attention concentration. To address sensitivity to the AEM weight parameter, we implement Cosine Weight Annealing, reducing parameter dependency. Extensive evaluations demonstrate AEM's superior performance across diverse feature extractors, MIL frameworks, attention mechanisms, and augmentation techniques. Here is our anonymous code: https://github.com/dazhangyu123/AEM.

cs.CV

Simulating the Unseen: Crash Prediction Must Learn from What Did Not Happen

Traffic safety science has long been hindered by a fundamental data paradox: the crashes we most wish to prevent are precisely those events we rarely observe. Existing crash-frequency models and surrogate safety metrics rely heavily on sparse, noisy, and under-reported records, while even sophisticated, high-fidelity simulations undersample the long-tailed situations that trigger catastrophic outcomes such as fatalities. We argue that the path to achieving Vision Zero, i.e., the complete elimination of traffic fatalities and severe injuries, requires a paradigm shift from traditional crash-only learning to a new form of counterfactual safety learning: reasoning not only about what happened, but also about the vast set of plausible yet perilous scenarios that could have happened under slightly different circumstances. To operationalize this shift, our proposed agenda bridges macro to micro. Guided by crash-rate priors, generative scene engines, diverse driver models, and causal learning, near-miss events are synthesized and explained. A crash-focused digital twin testbed links micro scenes to macro patterns, while a multi-objective validator ensures that simulations maintain statistical realism. This pipeline transforms sparse crash data into rich signals for crash prediction, enabling the stress-testing of vehicles, roads, and policies before deployment. By learning from crashes that almost happened, we can shift traffic safety from reactive forensics to proactive prevention, advancing Vision Zero.

cs.LG

Signal Timing Optimization for Mixed Connected Automated Traffic Based on A Markov Delay Approximation

Connected Automated Vehicles (CAVs) offer unparalleled opportunities to revolutionize existing transportation systems. In the near future, CAVs and human-driven vehicles (HDVs) are expected to coexist, forming a mixed traffic system. Although several prototype traffic signal systems leveraging CAVs have been developed, a simple yet realistic approximation of mixed traffic delay and optimal signal timing at intersections remains elusive. This paper presents an analytical approximation for delay and optimal cycle length at an isolated intersection of mixed traffic using a stochastic framework that combines Markov chain analysis, a car following model, and queuing theory. Given the intricate nature of mixed traffic delay, the proposed framework systematically incorporates the impacts of multiple factors, such as the distinct arrival and departure behaviors and headway characteristics of CAVs and HDVs, through mathematical derivations to ensure both realism and analytical tractability. Subsequently, closed-form expressions for intersection delay and optimal cycle length are derived. Numerical experiments are then conducted to validate the model and provide insights into the dynamics of mixed traffic delays at signalized intersections.

eess.SY

Neural Networks Enabled Discovery On the Higher-Order Nonlinear Partial Differential Equation of Traffic Dynamics

Modeling the traffic dynamics is essential for understanding and predicting the traffic spatiotemporal evolution. However, deriving the partial differential equation (PDE) models that capture these dynamics is challenging due to their potential high order property and nonlinearity. In this paper, we introduce a novel deep learning framework, "TRAFFIC-PDE-LEARN", designed to discover hidden PDE models of traffic network dynamics directly from measurement data. By harnessing the power of the neural network to approximate a spatiotemporal fundamental diagram that facilitates smooth estimation of partial derivatives with low-resolution loop detector data. Furthermore, the use of automatic differentiation enables efficient computation of the necessary partial derivatives through the chain and product rules, while sparse regression techniques facilitate the precise identification of physically interpretable PDE components. Tested on data from a real-world traffic network, our model demonstrates that the underlying PDEs governing traffic dynamics are both high-order and nonlinear. By leveraging the learned dynamics for prediction purposes, the results underscore the effectiveness of our approach and its potential to advance intelligent transportation systems.

eess.SY