SearcharxivSearch

arXiv subjects

Zhiyi Liu

Publications and source records attributed to Zhiyi Liu.

15 recordsLinked to original sources

Reproducible macroscopic dynamics in a closed-loop human-AI learning system

Closed-loop human-AI systems generate high-dimensional behavioural trajectories whose collective dynamics remain obscure. Using 297,915 learners' adaptive-tutoring histories, we define semantic order variables before model fitting and test them in user-disjoint cohorts. The state exhibits reproducible basin-like flow and operationally defined, state-heterogeneous metastable-like kinetics. A construction-matched null distinguishes normalised-memory relaxation from a reproducible excess field. A four-term conditional mechanism recovers population drift (r = 0.946; learner-bootstrap 95% CI, 0.935-0.955). Predictive event-level self-supervised learning recovers the state and learned-plane flow; null-referenced corrections retain directional, partial-amplitude excess-field structure without full calibration. Shuffled-order training reverses learned-plane flow on ordered trajectories; support-alignment randomisation selectively reduces inward transport. Both axes remain linearly accessible without state supervision. Without cross-model fitting, the models share leading population drift (r = 0.866; learner-bootstrap 95% CI, 0.857-0.875) and persistence ordering; residual directions remain model-specific. These results identify an externally anchored leading-order effective field linking empirical dynamics, an interpretable mechanism and neural computation.

cs.LG

The Research and Development of New Electronics System and its Testing on the JNE-1ton Prototype Detector

The Jinping Neutrino Experiment (JNE), a next-generation neutrino observatory under construction at the China Jinping Underground Laboratory II (CJPL-II), requires high-precision waveform-based event reconstruction, imposing stringent demands on its readout electronics. To meet these requirements, we have developed a high-performance readout system featuring 1 GSa/s real-time sampling, 14-bit physical resolution with an effective number of bits (ENOB) of 10.6, a total data throughput of 64 Gbps, and a deterministic zero-delay clock distribution architecture. The new single-crate 64-channel system (PDS1500) was validated through bench tests and deployment on the upgraded JNE-1ton prototype detector. Its performance was further evaluated against a commercial reference system. The results demonstrate that all key metrics meet the JNE experimental requirements: zero data loss within a 1000 ns acquisition window, baseline noise reduced to one-third of the reference level, timing drift limited to 0.3 ns across power cycles, and an energy threshold as low as 0.1 MeV, enabling the detection of low-energy solar neutrinos. While the 14-bit physical resolution provides significantly higher waveform fidelity, the overall energy resolution in this test remains dominated by the intrinsic limitations of the JNE-1ton detector, as expected. Furthermore, the modular architecture provides the throughput and scalability required to support the full-scale 3000-channel JNE detector. These results collectively demonstrate that the newly developed electronics system fully satisfies the technical requirements of the future JNE experiment.

physics.ins-det

Thresholds for the Frankl-Wang $3/7$ conjecture on maximum-degree ratios

Let $\mathcal{F}\subset\binom{[n]}{k}$ be an intersecting family, $\Delta(\mathcal{F})=\max_{x\in[n]}|\{F\in\mathcal{F}:x\in F\}|$, and $\varrho(\mathcal{F})=\Delta(\mathcal{F})/|\mathcal{F}|$. Frankl and Wang conjectured that if $n>100k$ and $|\mathcal{F}|>\binom{n-3}{k-3}$, then $\varrho(\mathcal{F})\ge 3/7$; the constant $3/7$ is sharp because of the Fano-plane construction. In this note we obtain three results. First, we show that no linear threshold $n>Ck$ can be sufficient: using a truncated Fano-plane construction we exhibit, for every constant $C$ and all large $k$, an intersecting family with $n>Ck$, $|\mathcal{F}|>\binom{n-3}{k-3}$, yet $\varrho(\mathcal{F})<3/7$. In particular, the original condition $n>100k$ does not guarantee the conclusion. Second, for $k=3$ we prove that $\varrho(\mathcal{F})\ge 3/7$ holds for every nonempty intersecting $3$-uniform family; the proof is nontrivial and does not rely on any assumption on $n$ or $|\mathcal{F}|$. Third, using the classical pseudo-sunflower bound $|\mathcal{F}|\le t^k$ (for families containing no pseudo-sunflower of size $t+1$), we obtain a completely explicit polynomial threshold for all $k\ge4$: if $n>(k-3)(7k^4+k)+3$ and $|\mathcal{F}|>\binom{n-3}{k-3}$, then $\varrho(\mathcal{F})\ge 3/7$. In particular, the simplified bound $n>7k^5$ is sufficient for every $k\ge4$.

math.CO

SIMAX: A Scalable and Interpretable Framework for Multi-Fidelity and Annotated Clinician-Patient Dialogue Simulation

Background. The widespread deployment of ambient digital scribes is driving large-scale capture of clinician-patient dialogues. Human coding of clinical communication data remains costly, inconsistent, and difficult to scale, motivating AI-driven communication coding systems. However, evaluating these systems requires real-world dialogues and human-coded labels, both hard to obtain at scale. Methods. We developed SIMAX (Scalable and Interpretable Framework for Multi-Fidelity and Annotated Clinician-Patient Dialogue Simulation), a framework for generating controlled clinical dialogue data with reference behavioral annotations. SIMAX generates clinician-patient dialogues from predefined clinical scenarios, personas and voice conditions, and target communication behaviors. Behaviors are controlled using two codebooks: the Global Codebook for overall communication quality and the WISER Codebook for specific countable behaviors. We evaluated SIMAX using automated and human quality assessments and an example communication coding system. Results. SIMAX generated 3,388 simulated dialogues across three specialties, multiple visit stages, persona characteristics, and accent conditions. Automated assessment showed mean UTMOS and WV-MOS scores of 3.03 and 2.61, WER and CER of 0.07 and 0.05, and CLAP cosine similarity of 0.41, suggesting reasonable speech naturalness, high transcription fidelity, and positive text-audio correspondence. Human evaluation showed a median MOS of 4.67 and a median clinical realism score of 3.00. Downstream evaluation suggests that SIMAX can assess how a communication coding system responds to behavioral targets and reveal insufficient sensitivity in some dimensions. Conclusions. SIMAX generates controlled and reproducible simulated clinician-patient dialogues, providing a data foundation for developing, validating, and refining communication coding systems.

cs.CL

Development and Performance Study of Vertical GaN $\alpha$-Particle Detector with High Energy Resolution

High-energy-resolution GaN $\alpha$-particle detectors have significant potential for space radiation, nuclear instrumentation, and harsh-environment applications. However, existing GaN $\alpha$-particle detectors still face several key challenges, including reducing the dead-layer thickness, suppressing leakage current under high reverse bias, improving energy resolution, and clarifying the physical mechanism underlying the low-energy tail phenomenon. This study presents a vertical homoepitaxial GaN $\alpha$-particle detector integrating a 20-nm ultrathin dead layer and a guard-ring structure. The detector exhibits an ultralow leakage current of 2.195 nA at -200 V and an intrinsic energy resolution of 2.69% with a charge collection efficiency (CCE) of 95.9% at -260 V. More importantly, this work demonstrates for the first time through Geant4 simulations that depletion-width nonuniformity is the dominant source of partial energy leakage, leading to an extended low-energy tail in the energy spectrum. We establish a depletion-width nonuniformity model and observe good agreement between simulation and experiment. This finding provides practical guidance for the design and optimization of high-performance GaN-based radiation detectors.

physics.ins-det

Disagreement as Signals: Dual-view Calibration for Sequential Recommendation Denoising

Sequential recommendation seeks to model the evolution of user interests by capturing temporal user intent and item-level transition patterns. Transformer-based recommenders demonstrate a strong capacity for learning long-range and interpretable dependencies, yet remain vulnerable to behavioral noise that is misaligned with users' true preferences. Recent large language model (LLM)-based approaches attempt to denoise interaction histories through static semantic editing. Such methods neglect the learning dynamics of recommendation models and fail to account for the evolving nature of user interests. To address this limitation, we propose a Dual-view Calibration framework for Sequential Recommendation denoising (DC4SR). Specifically, we introduce a semantic prior, derived from an LLM fine-tuned via labeled historical interactions, to estimate the noise distribution from a semantic perspective. From the learning perspective, we further employ a model-side posterior that infers the noise distribution based on the model's learning dynamics. The disagreement between the two distributions is then leveraged to jointly refine semantic understanding and learning-aware model-side representations. Through iterative updates, dynamic dual-view calibration is achieved for both the global semantic prior and the model-side posterior, enabling consistent alignment with evolving user interests. Extensive experiments demonstrate that DC4SR consistently outperforms strong Transformer-based recommenders and LLM-based denoising methods, exhibiting enhanced robustness across training stages and noise conditions.

cs.IR

Bollob\'{a}s-type inequalities for subspaces via weight invariance

Let $V$ be an $n$-dimension real vector space with a direct sum decomposition $V = V_1 \oplus \cdots \oplus V_r$. Let $\mathcal{P} = \{(A_i, B_i) : i \in [m]\}$ be a skew Bollob\'as system of subspaces of $V$ such that each $i\in [m]$, $ A_i = \bigoplus_{k=1}^r (A_i \cap V_k)$ and $ B_i = \bigoplus_{k=1}^r (B_i \cap V_k)$. We prove that $$\sum_{i=1}^{m} \prod_{k=1}^{r} \left[ \binom{a_{i,k} + b_{i,k}}{a_{i,k}} (1 + a_{i,k} + b_{i,k})^{-1} \right] \leq 1,$$ where $a_{i,k} = \dim(A_i \cap V_k)$ and $b_{i,k} = \dim(B_i \cap V_k)$. This extends a recent result of Yue from set systems to finite dimensional subspaces. We then consider Tuza's theorem on weak Bollob\'as system for $d$-tuples. We give an alternative proof of the original set version of Tuza, and also establish its vector space analogue. Precisely, let $\mathcal{P} = \{(A_i^{(1)}, \ldots, A_i^{(d)}) : i \in [m]\}$ be a skew Bollob\'as system of $d$-tuples of subspaces of finite dimensional space $V$ with $a^{(\ell)}_i=\dim (A_i^{(\ell)})$. Then, for any positive real numbers $p_1, \ldots, p_d$ satisfying $p_1 + \cdots + p_d = 1$, we prove that $ \sum_{i=1}^{m} \prod_{\ell=1}^{d} p_{\ell}^{a_i^{(\ell)}} \leq 1. $

math.CO

DisenReason: Behavior Disentanglement and Latent Reasoning for Shared-Account Sequential Recommendation

Shared-account usage is common on streaming and e-commerce platforms, where multiple users share one account. Existing shared-account sequential recommendation (SSR) methods often assume a fixed number of latent users per account, limiting their ability to adapt to diverse sharing patterns and reducing recommendation accuracy. Recent latent reasoning technique applied in sequential recommendation (SR) generate intermediate embeddings from the user embedding (e.g, last item embedding) to uncover users' potential interests, which inspires us to treat the problem of inferring the number of latent users as generating a series of intermediate embeddings, shifting from inferring preferences behind user to inferring the users behind account. However, the last item cannot be directly used for reasoning in SSR, as it can only represent the behavior of the most recent latent user, rather than the collective behavior of the entire account. To address this, we propose DisenReason, a two-stage reasoning method tailored to SSR. DisenReason combines behavior disentanglement stage from frequency-domain perspective to create a collective and unified account behavior representation, which serves as a pivot for latent user reasoning stage to infer the number of users behind the account. Experiments on four benchmark datasets show that DisenReason consistently outperforms all state-of-the-art baselines across four benchmark datasets, achieving relative improvements of up to 12.56\% in MRR@5 and 6.06\% in Recall@20.

cs.IR

Ahead of the Spread: Agent-Driven Virtual Propagation for Early Fake News Detection

Early detection of fake news is critical for mitigating its rapid dissemination on social media, which can severely undermine public trust and social stability. Recent advancements show that incorporating propagation dynamics can significantly enhance detection performance compared to previous content-only approaches. However, this remains challenging at early stages due to the absence of observable propagation signals. To address this limitation, we propose AVOID, an \underline{a}gent-driven \underline{v}irtual pr\underline{o}pagat\underline{i}on for early fake news \underline{d}etection. AVOID reformulates early detection as a new paradigm of evidence generation, where propagation signals are actively simulated rather than passively observed. Leveraging LLM-powered agents with differentiated roles and data-driven personas, AVOID realistically constructs early-stage diffusion behaviors without requiring real propagation data. The resulting virtual trajectories provide complementary social evidence that enriches content-based detection, while a denoising-guided fusion strategy aligns simulated propagation with content semantics. Extensive experiments on benchmark datasets demonstrate that AVOID consistently outperforms state-of-the-art baselines, highlighting the effectiveness and practical value of virtual propagation augmentation for early fake news detection. The code and data are available at https://github.com/Ironychen/AVOID.

cs.IR

The product measures of cross $t$-intersecting families

We investigate the product measures of intersection problems in extremal combinatorics. Invoking a recent result of He--Li--Wu--Zhang, we prove that for any $ n \geq t \geq 3$ and $ p_1, p_2 \in (0, \frac{1}{t+1})$, if $ \mathcal{F}_1, \mathcal{F}_2 \subseteq 2^{[n]}$ are cross $ t$-intersecting families, then $\mu_{p_1}(\mathcal{F}_1)\mu_{p_2}(\mathcal{F}_2)\le (p_1p_2)^t$. Secondly, we study the intersection problems for integer sequences by proving that if $\mathcal{H}_1, \mathcal{H}_2 \subseteq [m]^{n}$ are cross $t$-intersecting with $ m > t+1$, then $|\mathcal{H}_1|| \mathcal{H}_2|\leq (m^{n-t})^2$. These results confirm two classical conjectures of Tokushige. As an application, we strengthen a recent theorem of Frankl--Kupavskii, generalizing the well-known IU-Theorem. Finally, we show that if $ p \geq \frac{1}{2}$ and $ \mathcal{F}_1, \mathcal{F}_2 \subseteq 2^{[n]}$ are cross $t$-intersecting families, then $\min \left\{\mu_{p}(\mathcal{F}_1),\mu_{p}(\mathcal{F}_2)\right\} \leq \mu_{p}(\mathcal{K}(n,t))$, where $\mathcal{K}(n,t)$ denotes the Katona family. This recovers an old result of Ahlswede--Katona.

math.CO

Two results on set families: sturdiness and intersection

This paper resolves two open problems in extremal set theory. For a family $\mathcal{F} \subseteq 2^{[n]}$ and $i, j\in [n]$, we denote $\mathcal{F} (i,\bar{j})=\{F\backslash\{i\}: F\in \mathcal{F}, F\cap\{i,j\}=\{i\}\}$. The sturdiness $\beta (\mathcal{F})$ is defined as the minimum $|\mathcal{F} (i,\bar{j})|$ over all $i\neq j$. A family $\mathcal{F}$ is called an IU-family if it satisfies the intersection constraint: $F\cap F'\neq \emptyset $ for all $F,F'\in \mathcal{F}$, as well as the union constraint: $F\cup F' \neq [n]$ for all $F,F'\in \mathcal{F}$. The well-known IU-Theorem states that every IU-family $\mathcal{F}\subseteq 2^{[n]}$ has size at most $ 2^{n-2}$. In this paper, we prove that if $\mathcal{F}\subseteq 2^{[n]}$ is an IU-family, then $\beta (\mathcal{F})\le 2^{n-4}$. This confirms a recent conjecture proposed by Frankl and Wang. As the second result, we establish a tight upper bound on the sum of sizes of cross $t$-intersecting separated families. Our result not only extends a previous theorem of Frankl, Liu, Wang and Yang on separated families, but also provides explicit counterexamples to an open problem proposed by them, thereby settling their problem in the negative.

math.CO

Never compromise with vulnerabilities: a comprehensive survey on AI governance

The rapid advancement of AI has expanded its capabilities across domains, yet introduced critical technical vulnerabilities, such as algorithmic bias and adversarial sensitivity, that pose significant societal risks, including misinformation, inequity, security breaches, physical harm, and eroded public trust. These challenges highlight the urgent need for robust AI governance. We propose a comprehensive framework integrating technical and societal dimensions, structured around three interconnected pillars: Intrinsic Security (system reliability), Derivative Security (real-world harm mitigation), and Social Ethics (value alignment and accountability). Uniquely, our approach unifies technical methods, emerging evaluation benchmarks, and policy insights to promote transparency, accountability, and trust in AI systems. Through a systematic review of over 300 studies, we identify three core challenges: (1) the generalization gap, where defenses fail against evolving threats; (2) inadequate evaluation protocols that overlook real-world risks; and (3) fragmented regulations leading to inconsistent oversight. These shortcomings stem from treating governance as an afterthought, rather than a foundational design principle, resulting in reactive, siloed efforts that fail to address the interdependence of technical integrity and societal trust. To overcome this, we present an integrated research agenda that bridges technical rigor with social responsibility. Our framework offers actionable guidance for researchers, engineers, and policymakers to develop AI systems that are not only robust and secure but also ethically aligned and publicly trustworthy. The accompanying repository is available at https://github.com/Tele-EVOL/AI-Governance.

cs.CR

Development and Commissioning of a Compact Cosmic Ray Muon Imaging Prototype

Due to the muon tomography's capability of imaging high Z materials, some potential applications have been reported on inspecting smuggled nuclear materials in customs. A compact Cosmic Ray Muons (CRM) imaging prototype, Lanzhou University Muon Imaging System (LUMIS), is comprehensively introduced in this paper including the structure design, assembly, data acquisition and analysis, detector performance test, and material imaging commissioning etc. Casted triangular prism plastic scintillators (PS) were coupled with Si-PMs for sensitive detector components in system. LUMIS's experimental results show that the detection efficiency of an individual detector layer is about 98%, the position resolution for vertical incident muons is 2.5 mm and the angle resolution is 8.73 mrad given a separation distance of 40.5 cm. Moreover, the image reconstruction software was developed based on the Point of Closest Approach (PoCA) to detect lead bricks as our target. The reconstructed images indicate that the profile of the lead bricks in the image is highly consistent with the target. Subsequently, the capability of LUMIS to distinguish different materials, such as Pb, Cu, Fe, and Al, was investigated as well. The lower limit of response time for rapidly alarming high-Z materials is also given and discussed. The successful development and commissioning of the LUMIS prototype have provided a new solution option in technology and craftsmanship for developing compact CRM imaging systems that can be used in many applications.

physics.ins-det

Energy-based Periodicity Mining with Deep Features for Action Repetition Counting in Unconstrained Videos

Action repetition counting is to estimate the occurrence times of the repetitive motion in one action, which is a relatively new, important but challenging measurement problem. To solve this problem, we propose a new method superior to the traditional ways in two aspects, without preprocessing and applicable for arbitrary periodicity actions. Without preprocessing, the proposed model makes our method convenient for real applications; processing the arbitrary periodicity action makes our model more suitable for the actual circumstance. In terms of methodology, firstly, we analyze the movement patterns of the repetitive actions based on the spatial and temporal features of actions extracted by deep ConvNets; Secondly, the Principal Component Analysis algorithm is used to generate the intuitive periodic information from the chaotic high-dimensional deep features; Thirdly, the periodicity is mined based on the high-energy rule using Fourier transform; Finally, the inverse Fourier transform with a multi-stage threshold filter is proposed to improve the quality of the mined periodicity, and peak detection is introduced to finish the repetition counting. Our work features two-fold: 1) An important insight that deep features extracted for action recognition can well model the self-similarity periodicity of the repetitive action is presented. 2) A high-energy based periodicity mining rule using deep features is presented, which can process arbitrary actions without preprocessing. Experimental results show that our method achieves comparable results on the public datasets YT Segments and QUVA.

cs.CV

A Circumstantial Evidence for the Possible Production of QGP in the 158A GeV/c Central Pb+Pb Collisions

Hadron and string cascade model (JPCIAE) with the hypothesis without introducing the quark-gluon plasma (QGP), is employed to study the direct photon and $π^0$ transverse momentum distributions for central $^{208}$Pb+$^{208}$Pb collisions at 158A GeV/c . JPCIAE model, is based on LUND model, especially on the envent generator PYTHIA, and can be used to simulate the relativistic nucleus-nucleus collisions where PYTHIA is called to deal with hadron-hadron collisions. In our work, the theoretical results of transverse momentum distribution for both the direct photon and the $π^0$ particle are lower than the data of WA98 experiment. However, JPCIAE model can ever explain successfully the results of WA80 and WA93 experiments of central S+Au collisions at 200A GeV/c where no evidence of direct photon excess. Having considered the results of WA80 and WA93 experiments can be explained but WA98's can't, that might provide a circumstantial evidence for the possible production of QGP in the high-energy central Pb+Pb collisions.

hep-ph