Searcharxiv⌕ Search

arXiv subjects

Yan Sun

Publications and source records attributed to Yan Sun.

At least 37 records · Page 2Linked to original sources

OneReason Technical Report

Generative recommendation models in the OneRec family have been widely deployed in many real-world services, such as short-video, live-streaming, advertising, and e-commerce. However, these generative models can only benefit from the scaling advantage, while their reasoning ability is hard to activate, since we cannot construct meaningful Chain-of-Thought (CoT) sequences consisting of itemic tokens only. Inspired by the success of the reasoning-style ``think before answer'' paradigm in the LLM field, we conduct preliminary studies (i.e., OneRec-Think, OpenOneRec) to explore reasoning capability in generative recommendation. Nevertheless, we notice an unexpected phenomenon: the thinking mode does not show advantages over the non-thinking mode. Drawing insights from recent findings on CoT robustness in multi-modal language models, we argue that effective reasoning in recommendation rests on two factors: perception, the ability to ground itemic tokens in their underlying language semantics, and cognition, the ability to reorganize a user's behavior sequence into coherent latent interest points. We therefore propose OneReason, which includes: (1) strong itemic token perception in pre-training, (2) a three-level cognition-enhanced CoT format for recommendation tasks in SFT, and (3) a specialize-then-unify training recipe in RL to enhance the thinking ability.

cs.IR↗

Scaling-Aware Adapter for Structure-Grounded LLM Reasoning

Large language models (LLMs) are enabling reasoning over 2D and 3D structures, yet existing methods remain modality-specific and typically compress structural inputs through sequence-based tokenization or fixed-length query connectors. Such architectures either omit the geometric grounding requisite for mitigating structural hallucinations, or impose inflexible modality fusion bottlenecks that concurrently over-compress and suboptimally allocate structural tokens, thereby impeding the realization of generalized all-atom reasoning. We introduce Cuttlefish, a unified multimodal LLM that grounds language reasoning in geometric cues while scaling modality tokens with structural complexity. First, Scaling-Aware Patching leverages an instruction-conditioned gating mechanism to generate variable-size patches over structural graphs, adaptively scaling the query token budget with structural complexity to mitigate fixed-length connector bottlenecks. Second, Geometry Grounding Adapter refines these adaptive tokens via cross-attention to modality embeddings and injects the resulting modality tokens into the LLM, exposing explicit geometric cues to reduce structural hallucination. Experiments across interdisciplinary all-atom benchmarks demonstrate that Cuttlefish achieves superior performance in heterogeneous structure-grounded reasoning. Code: github.com/zihao-jing/Cuttlefish.

cs.AI↗

Convergent Differential Privacy Analysis for General Federated Learning

The powerful cooperation of federated learning (FL) and differential privacy~(DP) provides a promising paradigm for the large-scale private clients. However, existing analyses in FL-DP mostly rely on the composition theorem and cannot tightly quantify the privacy leakage challenges, which is tight for a few communication rounds but yields an arbitrarily loose and divergent bound eventually. This also implies a counterintuitive judgment, suggesting that FL-DP may not provide adequate privacy support during long-term training under constant-level noisy perturbations, yielding discrepancy between the theoretical and experimental results. To further investigate the convergent privacy and reliability of the FL-DP framework, in this paper, we comprehensively evaluate the worst privacy of two classical methods under the non-convex and smooth objectives based on the $f$-DP analysis. With the aid of the shifted interpolation technique, we successfully prove that privacy in {\ttfamily Noisy-FedAvg} has a tight convergent bound. Moreover, with the regularization of the proxy term, privacy in {\ttfamily Noisy-FedProx} has a stable constant lower bound. Our analysis further demonstrates a solid theoretical foundation for the reliability of privacy in FL-DP. Meanwhile, our conclusions can also be losslessly converted to other classical DP analytical frameworks, e.g. $(ε,δ)$-DP and R$\acute{\text{e}}$nyi-DP~(RDP), to provide more fine-grained understandings for the FL-DP frameworks.

cs.LG↗

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs

The rapid scaling of large language models~(LLMs) has made inference efficiency a primary bottleneck in the practical deployment. To address this, semi-structured sparsity offers a promising solution by strategically retaining $N$ elements out of every $M$ weights, thereby enabling hardware-friendly acceleration and reduced memory. However, existing (N:M)-compatible approaches typically fall into two categories: rule-based layerwise greedy search, which suffers from considerable errors, and gradient-driven combinatorial learning, which incurs prohibitive training costs. To tackle these challenges, we propose a novel linear-space probabilistic framework named MaskPro, which aims to learn a prior categorical distribution for every $M$ consecutive weights and subsequently leverages this distribution to generate the (N:M)-sparsity throughout an $N$-way sampling without replacement. Furthermore, to mitigate the training instability induced by the high variance of policy gradients in the super large combinatorial space, we propose a novel update method by introducing a moving average tracker of loss residuals instead of vanilla loss. Finally, we conduct comprehensive theoretical analysis and extensive experiments to validate the superior performance of MaskPro, as well as its excellent scalability in memory efficiency and exceptional robustness to data samples. Our code is available at \href{https://github.com/woodenchild95/Maskpro.git}{\ttfamily https://github.com/woodenchild95/Maskpro.git}.

cs.LG↗

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization

Pretraining large language models (LLMs) with next-token prediction has led to remarkable advances, yet the context-dependent nature of token embeddings in such models results in high intra-class variance and inter-class similarity, thus hindering the efficiency of representation learning. While similarity-based regularization has demonstrated benefit in supervised fine-tuning and classification tasks, its application and efficacy in large-scale LLM pretraining remains underexplored. In this work, we propose the SimReg, an embedding similarity regularization loss that explicitly encourages token representations with the same ground-truth label within each sequence to be more similar, while enforcing separation from different-label tokens via a contrastive loss. Our analysis reveals that this mechanism introduces gains by enlarging multi-classification margins, thereby enabling more efficient classification. Extensive experiments across dense and Mixture-of-Experts (MoE) architectures demonstrate that SimReg consistently accelerates training convergence by over 30% and improves average zero-shot downstream performance by over 1% across standard benchmarks. Further ablation studies and analyses offer practical insights into hyperparameter tuning and loss effectiveness.

cs.CL↗

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR

Reinforcement learning with verifiable rewards (RLVR) has emerged as a central paradigm for improving the reasoning capabilities of large language models. Group-based policy optimization methods, such as GRPO, typically allocate a fixed number of rollouts to every prompt. This uniform allocation can be inefficient: it over-allocates compute to prompts whose sampled groups are already saturated while under-exploring prompts for which additional samples may reveal useful correct trajectories. To address this limitation, we introduce hit utility, the posterior probability that at least one rollout in a proposed additional allocation for a prompt will be correct. Building on this notion, we propose Hit-Utility Optimal Rollout Allocation (HORA), a learning-free rollout allocation policy that maximizes total posterior hit utility within each allocation batch. HORA adaptively reallocates rollout budgets while leaving the downstream reward evaluation and group-based advantage estimator unchanged. Across four mathematical reasoning benchmarks and three model scales, HORA preserves comparable Pass@1 and improves Pass@K over compute-matched GRPO in ten of twelve model--benchmark configurations, with one tie and one saturated exception. It is also drop-in compatible with other group-based estimators such as RLOO. Ablation studies indicate that the uniform prior used by HORA is competitive with five prompt-conditioned learned-prior alternatives.

cs.LG↗

East Asian VLBI Network astrometry toward the star-forming region G040.96+02.48 in the Extreme Outer Galaxy

Accurate astrometric measurements for star-forming regions located on the far side of the Milky Way remain scarce. In this work, we present the astrometric results for a 22\,GHz water maser associated with star-forming region G040.96+02.48 located on the far side of the Milky Way, using the East Asian VLBI Network. The target water maser's proper motion was determined to be ($μ_α\cosδ, μ_δ$) = ($-2.06_{-0.51}^{+0.53}$, $-2.95_{-0.44}^{+0.45}$)~mas~yr$^{-1}$. The derived three-dimensional kinematic distance to the star-forming region is 20.2$\pm$3.2\,kpc, placing it slightly outside the Outer Scutum$-$Centaurus Arm. The corresponding vertical height of 872$\pm$139\,pc indicates a significant warp of the outer Galactic disk, which is in good agreement with the latest precessing warp model. Moreover, the resulting peculiar motions reveal a complex kinematic pattern, characterized by a large outward radial velocity of $-32\pm$18\,km~s$^{-1}$. Our observations substantially expand the valuable sample of star-forming regions with accurate astrometric measurements in the Extreme Outer Galaxy.

astro-ph.GA↗

Bayesian Analysis of Gravitational Wave Microlensing Effects from Galactic Double White Dwarfs

Gravitational waves (GWs) from the galactic double white dwarf (DWD) systems are one of the primary targets for upcoming space-based detectors. Due to their vast abundance and widespread distribution throughout the Galactic disk and bulge, these systems may provide a high-statistical population for probing GW microlensing effects induced by Galactic compact objects. To evaluate the detectability of such effects, in this work we simulate the four-year observation of DWD systems by Taiji, in the form of a second-generation Time Delay Interferometry (TDI) data stream. Within a Bayesian inference framework, we estimate parameters for lensed GWs from DWD systems for different values of the lens parameters, including the lens mass $M_\mathrm{L}\in [10, 10^6]$\,M$_\odot$, the effective velocity $v_\mathrm{eff}\in [50, 500]$\,km/s and the initial separation $L\in [R_\mathrm{E}, 3R_\mathrm{E}]$, and obtain the uncertainties of the corresponding parameters. These results characterize the capability of future Taiji observations to probe such systems. We further employ the Bayesian model selection framework to distinguish between lensed and unlensed scenarios, and investigate the impacts of three key physical parameters of the lens system: $M_\mathrm{L}$, $v_\mathrm{eff}$, and $L$ on distinguishing lensing events. Our results show that when $M_\mathrm{L}$ is below $10^5$\,M$_\odot$ or $L\geq3R_\mathrm{E}$, it is not possible to distinguish between lensed and unlensed models. For $v_\mathrm{eff}$, although the Bayes factor decreases as $v_\mathrm{eff}$ decreases, the lensed and unlensed models can still be distinguished within our parameter range.

astro-ph.GA↗

Rethinking the Personalized Relaxed Initialization in the Federated Learning: Consistency and Generalization

Federated learning (FL) is a distributed paradigm that coordinates massive local clients to collaboratively train a global model via stage-wise local training processes on the heterogeneous dataset. Previous works have implicitly studied that FL suffers from the ``client-drift'' problem, which is caused by the inconsistent optimum across local clients. However, till now it still lacks solid theoretical analysis to explain the impact of this local inconsistency. To alleviate the negative impact of ``client drift'' and explore its substance in FL, in this paper, we first propose an efficient FL algorithm FedInit, which allows employing the personalized relaxed initialization state at the beginning of each local training stage. Specifically, FedInit initializes the local state by moving away from the current global state towards the reverse direction of the latest local state. Moreover, to further understand how inconsistency disrupts performance in FL, we introduce the excess risk analysis and study the divergence term to investigate the test error in FL. Our studies show that optimization error is not sensitive to this local inconsistency, while it mainly affects the generalization error bound. Extensive experiments are conducted to validate its efficiency. The proposed FedInit method could achieve comparable results compared to several advanced benchmarks without any additional training or communication costs. Meanwhile, the stage-wise personalized relaxed initialization could also be incorporated into several current advanced algorithms to achieve higher generalization performance in the FL paradigm.

cs.LG↗

A Comparative Study of TeV Gamma-Ray Sources with Various Objects

We investigate the relationships between LHAASO TeV gamma-ray sources and various kinds of objects, including pulsar wind nebulae (PWNe), supernova remnants (SNRs), HII regions, microquasars, and OB associations. We propose a Randomization-Adjusted Overlap Correlation (RAOC) method to statistically assess association probabilities and evaluate association proportions across catalogs. The results reveal statistically significant overlaps between LHAASO sources and SNRs, PWNe, and microquasars, supporting their role as important contributors to TeV gamma-ray emission. The estimated association proportions of LHAASO sources are 0.19$\pm$0.08 with SNRs, 0.20$\pm$0.04 with PWNe, and 0.027$\pm$0.008 with microquasars. The proportion of the gamma-ray sources associated with the subsample of shell-type SNRs is ~0.1. While HII regions also show potential association, particularly with the KM2A component, their large self-overlap ratio complicates precise estimation. In contrast, OB associations exhibit a high probability of chance coincidence, suggesting their limited contribution to TeV gamma-ray emission. Our analysis of TeV gamma-ray emission capabilities shows that ~60% of PWNe are gamma-ray bright in both the WCDA and KM2A energy ranges. For SNRs and microquasars, the TeV gamma-ray bright fraction is ~10%. The subsample of PWNe associated with molecular clouds (MCs) shows enhanced gamma-ray emission. Furthermore, positional analysis reveals a systematic offset of the gamma-ray sources overlapping with PWNe toward the associated MCs. These findings imply a role for MCs in PWN gamma-ray production. Additionally, self-correlation analysis indicates that about 70% of the WCDA and KM2A gamma-ray components share a common origin. The study also identifies selection effects in existing SNR catalogs and notes clustering among approximately 30% of HII regions within larger star-forming regions.

astro-ph.HE↗

Berry curvature induced giant anomalous and spin texture driven Hall responses in the layered kagome antiferromagnet GdTi3Bi4

In recent years, layered kagome magnets have emerged as promising platforms for Berry-curvature engineering and unconventional transport phenomena. Here, we present the single-crystal growth, magnetization, and electrical transport characterizations of the van der Waals-like layered antiferromagnet GdTi3Bi4. The system exhibits pronounced field-induced first-order phase transitions. Comprehensive frequency, temperature, and field-dependent ac susceptibility measurements, and Hall analysis, reveals the formation of a spin-cluster-like glassy magnetic phase attributed to noncollinear spin textures. Additionally, the system demonstrates a colossal anomalous Hall conductivity σ_xy^{A}~ 8.6(7)10^{3} Ohm-1 cm-1 at 2 K). Detailed scaling analyses reveal the coexistence of skew scattering and intrinsic Berry-curvature contributions to the anomalous Hall effect. First-principles calculations highlight flat-band near the Fermi level, with f-electrons of the Gd ion contributing large intrinsic Hall response. Thus, GdTi3Bi4 emerges as a rare layered kagome magnet, exhibiting Berry curvature-induced giant anomalous and spin texture-driven Hall responses, providing a versatile platform for exploring spin-texture physics and advancing low-dimensional spintronic functionalities.

cond-mat.mtrl-sci↗

Multinoulli Extension: A Lossless Continuous Relaxation for Partition-Constrained Subset Selection

Identifying the most representative subset for a close-to-submodular objective while satisfying the predefined partition constraint is a fundamental task with numerous applications in machine learning. However, the existing distorted local-search methods are often hindered by their prohibitive query complexities and the rigid requirement for prior knowledge of difficult-to-obtain structural parameters. To overcome these limitations, we introduce a novel algorithm titled Multinoulli-SCG, which not only is parameter-free, but also can achieve the same approximation guarantees as the distorted local-search methods with significantly fewer function evaluations. More specifically, when the objective function is monotone $α$-weakly DR-submodular or $(γ,β)$-weakly submodular, our Multinoulli-SCG algorithm can attain a value of $(1-e^{-α})\text{OPT}-ε$ or $(\frac{γ^{2}(1-e^{-(β(1-γ)+γ^2)})}{β(1-γ)+γ^2})\text{OPT}-ε$ with only $O(1/ε^{2})$ function evaluations, where OPT denotes the optimal value. The cornerstone of our Multinoulli-SCG algorithm is an innovative continuous-relaxation framework named Multinoulli Extension(ME), which can effectively convert the discrete subset selection problem subject to partition constraints into a solvable continuous maximization focused on learning the optimal multinoulli priors across the concerned partition. In sharp contrast with the well-established multi-linear extension for submodular subset selection, a notable advantage of our proposed ME is its intrinsic capacity to provide a lossless rounding scheme for any set function. Furthermore, based on our proposed ME, we also present two novel online algorithms, namely, Multinoulli-OSCG and Multinoulli-OSGA, for the unexplored online subset selection problems over partition constraints.

cs.LG↗

Time Tracker: Mixture-of-Experts-Enhanced Foundation Time Series Forecasting Model with Decoupled Training Pipelines

In the past few years, time series foundation models have achieved superior predicting accuracy. However, real-world time series often exhibit significant diversity in their temporal patterns across different time spans and domains, making it challenging for a single model architecture to fit all complex scenarios. In addition, time series data may have multiple variables exhibiting complex correlations between each other. Recent mainstream works have focused on modeling times series in a channel-independent manner in both pretraining and finetuning stages, overlooking the valuable inter-series dependencies. To this end, we propose Time Tracker for better predictions on multivariate time series data. Firstly, we leverage sparse mixture of experts (MoE) within Transformers to handle the modeling of diverse time series patterns, thereby alleviating the learning difficulties of a single model while improving its generalization. Besides, we propose Any-variate Attention, enabling a unified model structure to seamlessly handle both univariate and multivariate time series, thereby supporting channel-independent modeling during pretraining and channel-mixed modeling for finetuning.Furthermore, we design a graph learning module that constructs relations among sequences from frequency-domain features, providing more precise guidance to capture inter-series dependencies in channel-mixed modeling. Based on these advancements, Time Tracker achieves state-of-the-art performance in predicting accuracy, model generalization and adaptability.

cs.LG↗

High-pressure phase stability and superconductivity in La-Zr-H hydrides

Hydrogen-rich ternary hydrides are promising candidates for high-Tc superconductivity at megabar pressures, yet their chemical space is vast and largely unexplored. Combining evolutionary structure searches with first-principles calculations, we comprehensively investigate the La-Zr-H ternary system in the 150-300 GPa pressure range. Zero-point energy-corrected convex hull analysis identifies multiple stable superconducting phases, including R3m-Zr2H17 at 300 GPa and P6/mmm-LaZr2H24 at 200 GPa, both of which are thermodynamically and dynamically stable and exhibit strong electron-phonon coupling. Solution of the Eliashberg equations predicts high superconducting transition temperatures of Tc = 209 K for R3m-Zr2H17 at 300 GPa and Tc = 202 K for P6/mmm-LaZr2H24 at 200 GPa. In addition to these stable phases, we identify a high-symmetry metastable compound, P6m2-LaZrH18, which lies just 0.027 eV/atom above the convex hull yet remains dynamically stable and exhibits a high predicted Tc of 206 K at 300 GPa. We find that, across all phases, the elevated Tc correlates with the high-symmetry structure with dense hydrogen cages, favorable electron counts per hydrogen, and a large hydrogen-derived density of states at the Fermi level. Finally, a random- forest machine learning model, trained on diverse hydrides superconductivity data, reproduces these structure-property trends across predicted structures, enabling to identify potential hydrides with high predicted Tc for targeted follow-up calculations and future high-pressure experiments.

cond-mat.mtrl-sci↗

VLBI astrometry of radio stars to link radio and optical celestial reference frames - II. 11 radio stars

The alignment between the radio-based International Celestial Reference Frame (ICRF) and the optical Gaia Celestial Reference Frame (Gaia-CRF) is critical for multi-waveband astronomy, yet systematic offsets at the optical bright end (G<13) limit their consistency. While radio stars offer a potential link between these frames, their utility has been restricted by the scarcity of precise Very Long Baseline Interferometry (VLBI) astrometry. In this study, we present new VLBI astrometry of 11 radio stars using the Very Long Baseline Array (VLBA), expanding the existing sample with positions, parallaxes, and proper motions measured. All 11 radio stars were detected, for 10 of which parallaxes and proper motions can be estimated, achieving median uncertainties better than 0.1 mas and 0.1 mas/yr, respectively. These new samples greatly contribute to the link between ICRF and Gaia-CRF at the optical bright end.

astro-ph.SR↗

Entropy-Guided Dynamic Tokens for Graph-LLM Alignment in Molecular Understanding

Molecular understanding is central to advancing areas such as scientific discovery, yet Large Language Models (LLMs) struggle to understand molecular graphs effectively. Existing graph-LLM bridges often adapt the Q-Former-style connector with fixed-length static tokens, which is originally designed for vision tasks. These designs overlook stereochemistry and substructural context and typically require costly LLM-backbone fine-tuning, limiting efficiency and generalization. We introduce EDT-Former, an Entropy-guided Dynamic Token Transformer that generates tokens aligned with informative molecular patches, thereby preserving both local and global structural features for molecular graph understanding. Beyond prior approaches, EDT-Former enables alignment between frozen graph encoders and LLMs without tuning the LLM backbone (excluding the embedding layer), resulting in computationally efficient finetuning, and achieves stateof-the-art results on MoleculeQA, Molecule-oriented Mol-Instructions, and property prediction benchmarks (TDC, MoleculeNet), underscoring its effectiveness for scalable and generalizable multimodal molecular understanding

cs.LG↗

Stability and Generalization of Push-Sum Based Decentralized Optimization over Directed Graphs

Push-Sum-based decentralized learning enables optimization over directed communication networks, where information exchange may be asymmetric. While convergence properties of such methods are well understood, their finite-iteration stability and generalization behavior remain unclear due to structural bias induced by column-stochastic mixing and asymmetric error propagation. In this work, we develop a unified uniform-stability framework for the Stochastic Gradient Push (SGP) algorithm that captures the effect of directed topology. A key technical ingredient is an imbalance-aware consistency bound for Push-Sum, which controls consensus deviation through two quantities: the stationary distribution imbalance parameter $δ$ and the spectral gap $(1-λ)$ governing mixing speed. This decomposition enables us to disentangle statistical effects from topology-induced bias. We establish finite-iteration stability and optimization guarantees for both convex objectives and non-convex objectives satisfying the Polyak--Łojasiewicz condition. For convex problems, SGP attains excess generalization error of order $\tilde{\mathcal{O}}\!\left(\frac{1}{\sqrt{mn}}+\fracγ{δ(1-λ)}+γ\right)$ under step-size schedules, and we characterize the corresponding optimal early stopping time that minimizes this bound. For PŁ objectives, we obtain convex-like optimization and generalization rates with dominant dependence proportional to $κ\!\left(1+\frac{1}{δ(1-λ)}\right)$, revealing a multiplicative coupling between problem conditioning and directed communication topology. Our analysis clarifies when Push-Sum correction is necessary compared with standard decentralized SGD and quantifies how imbalance and mixing jointly shape the best attainable learning performance.

cs.LG↗

Foundations of Top-$k$ Decoding For Language Models

Top-$k$ decoding is a widely used method for sampling from LLMs: at each token, only the largest $k$ next-token-probabilities are kept, and the next token is sampled after re-normalizing them to sum to unity. Top-$k$ and other sampling methods are motivated by the intuition that true next-token distributions are sparse, and the noisy LLM probabilities need to be truncated. However, to our knowledge, a precise theoretical motivation for the use of top-$k$ decoding is missing. In this work, we develop a theoretical framework that both explains and generalizes top-$k$ decoding. We view decoding at a fixed token as the recovery of a sparse probability distribution. We consider \emph{Bregman decoders} obtained by minimizing a separable Bregman divergence (for both the \emph{primal} and \emph{dual} cases) with a sparsity-inducing $\ell_0$ regularization. Despite the combinatorial nature of the objective, we show how to optimize it efficiently for a large class of divergences. We show that the optimal decoding strategies are greedy, and further that the loss function is discretely convex in $k$, so that binary search provably and efficiently finds the optimal $k$. We show that top-$k$ decoding arises as a special case for the KL divergence, and identify new decoding strategies that have distinct behaviors (e.g., non-linearly up-weighting larger probabilities after re-normalization).

cs.AI↗