Searcharxiv⌕ Search

arXiv subjects

Jun Yan

Publications and source records attributed to Jun Yan.

At least 55 records · Page 3Linked to original sources

Privacy-Preserving EHR Data Transformation via Geometric Operators: A Human-AI Co-Design Technical Report

Electronic health records (EHRs) and other real-world clinical data are essential for clinical research, medical artificial intelligence, and life science, but their sharing is severely limited by privacy, governance, and interoperability constraints. These barriers create persistent data silos that hinder multi-center studies, large-scale model development, and broader biomedical discovery. Existing privacy-preserving approaches, including multi-party computation and related cryptographic techniques, provide strong protection but often introduce substantial computational overhead, reducing the efficiency of large-scale machine learning and foundation-model training. In addition, many such methods make data usable for restricted computation while leaving them effectively invisible to clinicians and researchers, limiting their value in workflows that still require direct inspection, exploratory analysis, and human interpretation. We propose a real-world-data transformation framework for privacy-preserving sharing of structured clinical records. Instead of converting data into opaque representations, our approach constructs transformed numeric views that preserve medical semantics and major statistical properties while, under a clearly specified threat model, provably breaking direct linkage between those views and protected patient-level attributes. Through collaboration between computer scientists and the AI agent \textbf{SciencePal}, acting as a constrained tool inventor under human guidance, we design three transformation operators that are non-reversible within this threat model, together with an additional mixing strategy for high-risk scenarios, supported by theoretical analysis and empirical evaluation under reconstruction, record linkage, membership inference, and attribute inference attacks.

cs.CR↗

Ordered Ramsey and Turán numbers of alternating paths and their variants

An ordered graph is a graph whose vertex set is equipped with a total order. The ordered complete graph $K_N^<$ is the complete graph with vertex set $[N]$ equipped with the natural ordering of the integers. Given an ordered graph $H$, the ordered Ramsey number $R_<(H)$ is the smallest integer $N$ such that every red/blue edge-colouring of $K_N^<$ contains a monochromatic copy of $H$ with vertices appearing in the same relative order as in $H$. Balko, Cibulka, Král, and Kynčl asked whether, among all ordered paths on $n$ vertices, the ordered Ramsey number is minimised by the alternating path $\mathrm{AP}_n$ -- the ordered path with vertex set $[n]$ such that the vertices encountered along the path are $1, n, 2, n - 1,3, n-2,\dots$. Motivated by this problem, we make progress on establishing the value of $R_<(\mathrm{AP}_n)$ by proving that \[ R_{<}(\mathrm{AP}_n)\leq \left(2+\frac{\sqrt{2}}{2}+o(1)\right)n. \] We then use similar methods to determine the exact ordered Turán number of $\mathrm{AP}_n$, and study the ordered Ramsey and Turán numbers of several related ordered paths.

math.CO↗

More Shattering News

An ordered variant of the well-known set theory concept of shattering was introduced by Anstee, Rónyai, and Sali. In this paper, we prove several new results related to order shattering. Given a family $\mathcal F$ of subsets of $[n]$, we show that $\mathrm{osh}(\mathcal F)$, the family of all sets order shattered by $\mathcal F$, coincides with $T(\mathcal F)$, the family obtained from $\mathcal F$ by the down-shift operation. We then give a full characterization of all sets that can be order shattered by some $\ell$-Sperner family. Finally, we completely determine $\mathrm{osh}\left(\binom{[n]}{a}\cup\binom{[n]}{b}\right)$.

math.CO↗

ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory

With the growing adoption of large language model agents in persistent real-world roles, they naturally encounter continuous streams of tasks. A key limitation, however, is their failure to learn from the accumulated interaction history, forcing them to discard valuable insights and repeat past errors. We propose ReasoningBank, a novel memory framework that distills generalizable reasoning strategies from an agent's self-judged successful and failed experiences. At test time, an agent retrieves relevant memories from ReasoningBank to inform its interaction and then integrates new learnings back, enabling it to become more capable over time. Building on this powerful experience learner, we further introduce memory-aware test-time scaling (MaTTS), which accelerates and diversifies this learning process by scaling up the agent's interaction experience. By allocating more compute to each task, the agent generates abundant, diverse experiences that provide rich contrastive signals for synthesizing higher-quality memory. The better memory in turn guides more effective scaling, establishing a powerful synergy between memory and test-time scaling. Across web browsing and software engineering benchmarks, ReasoningBank consistently outperforms existing memory mechanisms that store raw trajectories or only successful task routines, improving both effectiveness and efficiency; MaTTS further amplifies these gains. These findings establish memory-driven experience scaling as a new scaling dimension, enabling agents to self-evolve with emergent behaviors naturally arise. Our code can be found at https://github.com/google-research/reasoning-bank.

cs.AI↗

Domain-Skewed Federated Learning with Feature Decoupling and Calibration

Federated learning (FL) allows distributed clients to collaboratively train a global model in a privacy-preserving manner. However, one major challenge is domain skew, where clients' data originating from diverse domains may hinder the aggregated global model from learning a consistent representation space, resulting in poor generalizable ability in multiple domains. In this paper, we argue that the domain skew is reflected in the domain-specific biased features of each client, causing the local model's representations to collapse into a narrow low-dimensional subspace. We then propose Federated Feature Decoupling and Calibration ($F^2$DC), which liberates valuable class-relevant information by calibrating the domain-specific biased features, enabling more consistent representations across domains. A novel component, Domain Feature Decoupler (DFD), is first introduced in $F^2$DC to determine the robustness of each feature unit, thereby separating the local features into domain-robust features and domain-related features. A Domain Feature Corrector (DFC) is further proposed to calibrate these domain-related features by explicitly linking discriminative signals, capturing additional class-relevant clues that complement the domain-robust features. Finally, a domain-aware aggregation of the local models is performed to promote consensus among clients. Empirical results on three popular multi-domain datasets demonstrate the effectiveness of the proposed $F^2$DC and the contributions of its two modules. Code is available at https://github.com/mala-lab/F2DC.

cs.LG↗

A Stable, High-Order Time-Stepping Scheme for the Drift-Diffusion Model in Modern Solar Cell Simulation

This paper presents a one-dimensional transient drift--diffusion simulator for advanced solar cells, integrating a structure-preserving finite-volume spatial discretization with Scharfetter--Gummel--type fluxes and a high-order, L-stable implicit Runge--Kutta (Radau IIA) temporal integrator. The scheme ensures local charge conservation, handles sharp material interfaces, and achieves second-order spatial and fifth-order temporal convergence. Its accuracy is verified against the classical depletion approximation in $p$--$n$ junction and validated through excellent agreement with the established simulator for an organic photovoltaic device. The framework's extensibility is demonstrated by incorporating exciton kinetics in organic solar cells, capturing multi-timescale dynamics, and by modeling mobile ions in perovskite solar cells, reproducing characteristic $\tmem{J}$--$\tmem{V}$ hysteresis without empirical parameters. This work provides a robust, high-order numerical foundation for simulating coupled charge, exciton, and ion transport in next-generation photovoltaic devices.

physics.app-ph↗

Diagnostics for Semiparametric Accelerated Failure Time Models with R Package afttest

The semiparametric accelerated failure time (AFT) model offers a direct and interpretable alternative to the Cox proportional hazards model, yet practical diagnostic tools for this framework remain limited. We introduce afttest, an R package that implements martingale-residual-based goodness-of-fit procedures for semiparametric AFT models. In addition to the recently developed multiplier bootstrap diagnostics, the package introduces a new computationally efficient resampling strategy based on an influence-function linear approximation. Unlike the original approach, which requires repeatedly solving estimating equations for each bootstrap replicate, the proposed method avoids iterative optimization and substantially reduces computation time while preserving asymptotic validity. Both the standard multiplier bootstrap and the accelerated linear approximation are implemented, allowing users to balance finite-sample performance and computational scalability. The package supports rank-based and least-squares estimators, provides omnibus, link function, and functional form tests, and includes graphical tools for visualizing residual processes. An application to the Mayo Clinic primary biliary cirrhosis study illustrates the workflow.

stat.CO↗

Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoning

Large Language Models (LLMs) often struggle with problems that require multi-step reasoning. For small-scale open-source models, Reinforcement Learning with Verifiable Rewards (RLVR) fails when correct solutions are rarely sampled even after many attempts, while Supervised Fine-Tuning (SFT) tends to overfit long demonstrations through rigid token-by-token imitation. To address this gap, we propose Supervised Reinforcement Learning (SRL), a framework that reformulates problem solving as generating a sequence of logical "actions". SRL trains the model to generate an internal reasoning monologue before committing to each action. It provides smoother rewards based on the similarity between the model's actions and expert actions extracted from the SFT dataset in a step-wise manner. This supervision offers richer learning signals even when all rollouts are incorrect, while encouraging flexible reasoning guided by expert demonstrations. As a result, SRL enables small models to learn challenging problems previously unlearnable by SFT or RLVR. Moreover, initializing training with SRL before refining with RLVR yields the strongest overall performance. Beyond reasoning benchmarks, SRL generalizes effectively to agentic software engineering tasks, establishing it as a robust and versatile training framework for reasoning-oriented LLMs.

cs.CL↗

Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning

Large Language Models (LLMs) are widely used as judges to evaluate response quality, providing a scalable alternative to human evaluation. However, most LLM judges operate solely on intrinsic text-based reasoning, limiting their ability to verify complex constraints or perform accurate computation. Motivated by the success of tool-integrated reasoning (TIR) in numerous tasks, we propose TIR-Judge, an end-to-end RL framework for training LLM judges that integrates a code executor for precise evaluation. TIR-Judge is built on three principles: (i) diverse training across verifiable and non-verifiable domains, (ii) flexible judgment formats (pointwise, pairwise, listwise), and (iii) iterative RL that bootstraps directly from the initial model without distillation. On seven public benchmarks, TIR-Judge surpasses strong reasoning-based judges by up to 6.4% (pointwise) and 7.7% (pairwise), and achieves listwise performance comparable to Claude-Opus-4 despite having only 8B parameters. Remarkably, TIR-Judge-Zero - trained entirely without distilled judge trajectories, matches the performance of distilled variants, demonstrating that tool-augmented judges can self-evolve through iterative reinforcement learning.

cs.CL↗

On the Validity of Isotropic Covariance Functions for Set-indexed Random Fields

Distances between sets arise naturally when modeling stochastic dependence on collections of spatial supports, including settings with point-referenced and areal observations. However, commonly used constructions of distances on sets, including those derived from the Hausdorff distance, generally fail to be conditionally negative definite, precluding their use in isotropic covariance models. We propose the ball--Hausdorff distance, defined as the Hausdorff distance between the minimum enclosing balls of bounded sets in a metric space. For length spaces, we derive an explicit representation of this distance in terms of the associated centers and radii. We show that the ball--Hausdorff distance is conditionally negative definite whenever the underlying metric is conditionally negative definite. By Schoenberg's theorem, this implies an isometric embedding into a Hilbert space and guarantees the validity of broad classes of isotropic covariance functions, including the Matérn and powered exponential families, for set-indexed random fields. The construction reduces dependence between sets to low-dimensional geometric summaries, leading to substantial simplifications in covariance evaluation.

stat.ME↗

Static class-guided selection of elementary solutions in non-monotone vanishing discount problems

We study a generalized vanishing discount problem for Hamilton--Jacobi equations, removing the standard monotonicity assumption, either in a global sense or when integrated against all Mather measures. Specifically, we consider \[ λa(x)u(x)+H(x,Du(x))-Aλ=c_0, \] with a suitably chosen constant $A>0$. By appropriately changing the signs of the function $a(x)$ on different static classes associated with $H$, we show that the maximal viscosity solution converges uniformly as $λ\to 0^+$ and that all elementary solutions of the stationary equation \[ H(x,Du(x))=c_0 \] can be selected as limits. This provides the first result for selecting multiple viscosity solutions in vanishing discount problems beyond the usual monotonicity and integral assumptions, as long as $a(x)$ is positive on one static class. Our results highlight the crucial role of static classes in controlling the asymptotic behavior of viscosity solutions. Previously, under usual monotonicity assumptions, only a single solution could be selected (as discussed in \cite{GL}), whereas our approach allows controlled selection of multiple solutions via static class-guided discount coefficients.

math.AP↗

R-Matrix Theory for Electron-Ion Collisions in Plasmas

Electron-atom collisions in warm dense plasmas are crucial for astrophysics and controlled fusion research, where calculating short-range scattering matrices under screening plasma potentials is essential. While electron-neutral atom collisions are tractable using the standard Riccati-Bessel wavefunctions in the asymptotic region, electron-ion collisions face challenges due to the extended range of the screened Coulomb potential, which lacks analytical solutions or numerical code packages for asymptotic regular and irregular wavefunctions. We introduce an R-matrix theoretical framework for general screened potentials and develop a numerical method to compute these asymptotic wavefunctions efficiently. Our approach yields short-range scattering phase shifts that remain invariant with respect to the matching point in the asymptotic region. Applying the Debye screening potential as an illustrative example, we calculate elastic and electron-impact excitation collision strengths for H-like ions (He, C, Ne) across varying temperatures and densities. The calculations show that Debye screening systematically modifies resonance structures and progressively lowers excitation thresholds. Nevertheless, the effective collision strengths and rate coefficients exhibit approximate scaling laws. These findings enable convenient access to electron collision data in plasma environments, advancing plasma diagnostics and modeling.

physics.atom-ph↗

SAGE: Steerable Agentic Data Generation for Deep Search with Execution Feedback

Deep search agents, which aim to answer complex questions requiring reasoning across multiple documents, can significantly speed up the information-seeking process. Collecting human annotations for this application is prohibitively expensive due to long and complex exploration trajectories. We propose an agentic pipeline that automatically generates high quality, difficulty-controlled deep search question-answer pairs for a given corpus and a target difficulty level. Our pipeline, SAGE, consists of a data generator which proposes QA pairs and a search agent which attempts to solve the generated question and provide execution feedback for the data generator. The two components interact over multiple rounds to iteratively refine the question-answer pairs until they satisfy the target difficulty level. Our intrinsic evaluation shows SAGE generates questions that require diverse reasoning strategies, while significantly increases the correctness and difficulty of the generated data. Our extrinsic evaluation demonstrates up to 23% relative performance gain on popular deep search benchmarks by training deep search agents with our synthetic data. Additional experiments show that agents trained on our data can adapt from fixed-corpus retrieval to Google Search at inference time, without further training.

cs.AI↗

Distribution of independent sets in perfect $r$-ary trees

Given a graph $G$, the family of all independent sets of size $k$ containing a fixed vertex $v$ is called a star with centre $v$, and is denoted by $\mathcal{I}_G^k(v)$. Motivated by a generalisation of the Erdős-Ko-Rado Theorem to the setting of independent sets in graphs, Hurlbert and Kamat conjectured that for every tree $T$ and every $k$, the maximum of $|\mathcal{I}_T^k(v)|$ can always be attained by a leaf of $T$. While this conjecture turns out to be false in general, it is known to hold for specific families of trees like spiders and caterpillars. In this paper, we prove that this conjecture holds for a new family of trees, the perfect $r$-ary trees, by constructing injections from stars centred at arbitrary vertices to stars centred at leaves. We also show that the analogous property holds for every forest $\mathcal{T}$ that is the disjoint union of perfect trees with possibly varying sizes and arities, and determine the leaf that maximises $|\mathcal{I}_{\mathcal{T}}^k(v)|$.

math.CO↗

When Does the Silhouette Score Work? A Comprehensive Study in Network Clustering

Selecting the number of communities is a fundamental challenge in network clustering. The silhouette score offers an intuitive, model-free criterion that balances within-cluster cohesion and between-cluster separation. Albeit its widespread use in clustering analysis, its performance in network-based community detection remains insufficiently characterized. In this study, we comprehensively evaluate the performance of the silhouette score across unweighted, weighted, and fully connected networks, examining how network size, separation strength, and community size imbalance influence its performance. Simulation studies show that the silhouette score accurately identifies the true number of communities when clusters are well separated and balanced, but it tends to underestimate under strong imbalance or weak separation and to overestimate in sparse networks. Extending the evaluation to a real airline reachability network, we demonstrate that the silhouette-based clustering can recover geographically interpretable and market-oriented clusters. These findings provide empirical guidance for applying the silhouette score in network clustering and clarify the conditions under which its use is most reliable.

cs.SI↗

Co-Evolution of Types and Dependencies: Towards Repository-Level Type Inference for Python Code

Python's dynamic typing mechanism, while promoting flexibility, is a significant source of runtime type errors that plague large-scale software, which inspires the automatic type inference techniques. Existing type inference tools have achieved advances in type inference within isolated code snippets. However, repository-level type inference remains a significant challenge, primarily due to the complex inter-procedural dependencies that are difficult to model and resolve. To fill this gap, we present \methodName, a novel approach based on LLMs that achieves repository-level type inference through the co-evolution of types and dependencies. \methodName~constructs an Entity Dependency Graph (EDG) to model the objects and type dependencies across the repository. During the inference process, it iteratively refines types and dependencies in EDG for accurate type inference. Our key innovations are: (1) an EDG model designed to capture repository-level type dependencies; (2) an iterative type inference approach where types and dependencies co-evolve in each iteration; and (3) a type-checker-in-the-loop strategy that validates and corrects inferences on-the-fly, thereby reducing error propagation. When evaluated on 12 complex Python repositories, \methodName~significantly outperformed prior works, achieving a \textit{TypeSim} score of 0.89 and a \textit{TypeExact} score of 0.84, representing a 27\% and 40\% relative improvement over the strongest baseline. More importantly, \methodName~removed new type errors introduced by the tool by 92.7\%. This demonstrates a significant leap towards automated, reliable type annotation for real-world Python development.

cs.SE↗

A Further Comparison of MPS and TTNS for Nonadiabatic Dynamics of Exciton Dissociation

Tensor networks, such as matrix product states (MPS) and tree tensor network states (TTNS), are powerful ansätze for simulating quantum dynamics. While both ansätze are theoretically exact in the limit of large bond dimensions, [J. Chem. Theory Comput. 2024, 20, 8767-8781] reported a non-negligible discrepancy in its calculations for exciton dissociation. To resolve this inconsistency, we conduct a systematic comparison using Renormalizer, a unified software framework for MPS and TTNS. By revisiting the benchmark P3HT:PCBM heterojunction model, we show that the observed discrepancies arise primarily from insufficient bond dimensions. By increasing bond dimensions, we reduce the relative difference in occupancy for weakly populated electronic states from up to 60% towards the end of the simulation to less than 10% and the absolute difference from 0.05 to 0.005. We also discuss the impact of tensor network structures on accuracy and efficiency, with the difference further reduced by an optimized TTNS structure. Our results confirm that both methods converge to numerically exact solutions when bond dimensions are adequately scaled. This work not only validates the reliability of both methods but also provides high-accuracy benchmark data for future developments in quantum dynamics simulations.

physics.chem-ph↗

Self-Refined Generative Foundation Models for Wireless Traffic Prediction

With a broad range of emerging applications in 6G networks, wireless traffic prediction has become a critical component of network management. However, the dynamically shifting distribution of wireless traffic in non-stationary 6G networks presents significant challenges to achieving accurate and stable predictions. Motivated by recent advancements in Generative AI (GenAI)-enabled 6G networks, this paper proposes a novel self-refined Large Language Model (LLM) for wireless traffic prediction, namely TrafficLLM, through in-context learning without parameter fine-tuning or model training. The proposed TrafficLLM harnesses the powerful few-shot learning abilities of LLMs to enhance the scalability of traffic prediction in dynamically changing wireless environments. Specifically, our proposed TrafficLLM embraces an LLM to iteratively refine its predictions through a three-step process: traffic prediction, feedback generation, and prediction refinement. Initially, the proposed TrafficLLM conducts traffic predictions using task-specific demonstration prompts. Recognizing that LLMs may generate incorrect predictions on the first attempt, this paper designs feedback demonstration prompts to provide multifaceted and valuable feedback related to these initial predictions. The validation scheme is further incorporated to systematically enhance the accuracy of mathematical calculations during the feedback generation process. Following this comprehensive feedback, our proposed TrafficLLM introduces refinement demonstration prompts, enabling the same LLM to further refine its predictions and thereby enhance prediction performance. Evaluations on two realistic datasets demonstrate that the proposed TrafficLLM outperforms LLM-based in-context learning methods, achieving performance improvements of 23.17% and 17.09%, respectively.

eess.SY↗