SearcharxivSearch

arXiv subjects

Jun Xu

Publications and source records attributed to Jun Xu.

At least 19 recordsLinked to original sources

Cosmic-ray electron propagation in the peculiar barred spiral galaxy NGC 2442

Face-on spiral galaxies offer a favorable geometry for studying magnetic-field structures and cosmic-ray (CR) propagation because projection effects and structural overlap are reduced. We investigate cosmic-ray electron (CRE) transport in the nearby face-on spiral galaxy NGC 2442 and assess how its environment and magnetic-field structure influence propagation. We combine radio continuum (RC) observations from ASKAP at 943 MHz, MeerKAT at 1.28 and 1.7 GHz, and ATCA at 5 GHz with optical H$\alpha$ and infrared data, and compare them with 2D CRE transport simulations. NGC 2442 has a steep integrated RC spectrum, with $\alpha=-0.96\pm0.04$ for the total emission and $\alpha_{\rm nt}=-1.21\pm0.04$ for the synchrotron emission over 408 MHz-5 GHz. A break near 1 GHz indicates substantial radiative aging. Under equipartition, we derive a mean magnetic-field strength of $10.8\,\mu{\rm G}$. RC-SFR smoothing gives effective CRE propagation lengths of $\sim0.65$-$0.89$ kpc at 943-1700 MHz and $\sim0.44$ kpc at 5 GHz, corresponding to diffusion coefficients of order $10^{28}\,{\rm cm^2\,s^{-1}}$. We identify a steep-spectrum synchrotron ``island'' in the southeast, with $\alpha\sim-1.09$ and no clear H$\alpha$, infrared, FUV, or NUV counterpart, indicating that CREs are unlikely to be injected in situ. Our 2D CRPropa simulations show that anisotropic diffusion along ordered magnetic fields enables CREs to reach the island more efficiently than isotropic diffusion. NGC 2442 therefore shows that environmental disturbances and ordered magnetic fields can strongly regulate CRE propagation in disturbed spiral galaxies.

astro-ph.GA

Key Path Identification for Resolving Knowledge Conflicts via SAE-based Steering

Sparse autoencoder (SAE)-based steering has been widely used to address knowledge conflicts by guiding LLMs to be more faithful to the contextual knowledge. Existing methods usually perform mass steering, which modifies a large batch of SAE features identified via correlation-based methods. However, due to the inaccurate correlation and the neglected feature interactions, mass steering methods fail to precisely identify the features that play the key roles in steering and introduce a large number of redundant ones, which add noise and weaken the steering effects. Our empirical studies reveal that steering only a small subset of the identified features can achieve comparable or even better performance. Motivated by this finding, we propose Key Path Identification (KPI), a novel method that identifies key steering features characterized by strong causal dependencies with both upstream and downstream features. From these features, KPI constructs key paths and steers through less feature modifications. In this way, KPI advances SAE-based steering from quantity-driven to quality-focused, offering a perspective for more precise and interpretable model editing. Experiments in RAG tasks with knowledge conflicts show that our method improves the accuracy by 18% on average compared to the best baseline of mass steering, effectively filtering redundant features, alleviating side effects and demonstrating the core role of key paths in steering.

cs.AI

QPS-ToR: A Parallel Iterative Switching Algorithm for Reconfigurable Optical Datacenter Switching

Reconfigurable optical data center networks (RODCNs) have emerged as a promising solution for scaling DCN capacity, yet their scheduling mechanisms remain a performance bottleneck: traffic-oblivious schemes inherently limit throughput, while the state-of-the-art traffic-aware scheme, NegotiaToR, uses single-iteration iSLIP as its scheduling engine, which limits throughput to around 60% and treats all source-destination pairs with equal priority regardless of queue length. We propose QPS-ToR, which replaces NegotiaToR's scheduling logic with SW-QPS, a sliding-window algorithm originally proposed for crossbar scheduling that achieves around 90% throughput with a single low-complexity iteration. QPS-ToR operates within NegotiaToR's existing workflow, requiring only a revision of the scheduling cycle from three-step Request-Grant-Accept (RGA) to two-step Request-Grant (RG) with a sliding window mechanism. In flow-level simulations with 128 ToRs under realistic datacenter workloads, QPS-ToR achieves up to 36% higher throughput and 82% lower flow completion time (FCT) compared to NegotiaToR, and consistently outperforms RotorNet, a representative traffic-oblivious scheme, on the parallel network topology.

cs.NI

High-purity fluorescence photon bundle emission from two separate emitters

We propose a scheme for generating high-purity fluorescence two-photon bundles from two spatially separate two-level emitters, whose correlated excitation is mediated by a strongly driven auxiliary emitter. Unlike bosonic-mode-based approaches, the finite excitation space of the two target emitters intrinsically excludes higher-excitation manifolds, and thus eliminates impurity photons associated with undesired higher-excitation states. We further analytically characterize the residual population of off-resonant single-excitation states and show that suppressing the corresponding leakage improves photon bundle purity. Our scheme offers a feasible pathway for constructing high-purity quantum light sources with promising applications in quantum information processing.

quant-ph

Synergistic Information Disentanglement for Omni-modal Slide Representation Learning in Computational Pathology

In computational pathology (CPath), developing omni-modal self-supervised learning (SSL) models that integrate histology, genomics, and clinical reports enables transferable representation learning for whole slide images (WSIs). Existing approaches implicitly force heterogeneous modalities into a uniform latent space by contrastive alignment, causing modality collapse where unique, synergistic diagnostic signals (termed as $\mathrm{\Phi}$) are discarded in favor of trivial redundancy. We hypothesize that the strongest task-agnostic SSL training signal stems from distilling the synergistic interactions over merely aligning shared redundancy. To this end, we introduce \textsc{$\mathrm{\Phi}$-Omni}, a synergistic information disentanglement framework grounded in Partial Information Decomposition (PID) theory for slide representation learning. Unlike standard contrastive approaches, \textsc{$\mathrm{\Phi}$-Omni} employs a Synergistic Information Bottleneck (SIB) regulated by the proposed $\mathrm{\Phi}\text{ID}$ objective, which explicitly suppresses marginal redundancy while maximizing irreducible synergy, thereby distilling high-order cross-modal interactions. Following pretraining on breast ($n$=1031) and lung ($n$=919) cohorts, \textsc{$\mathrm{\Phi}$-Omni} demonstrates superior few-shot performance across five independent external datasets spanning eight tasks compared to supervised and SSL baselines. Source code is available here.

cs.CV

Spin transport in intermediate-energy heavy-ion collisions

In this mini-review, we provide a brief status report on investigating spin transport phenomena in intermediate-energy heavy-ion collisions by simulating solutions of the spin- and isospin-dependent Boltzmann-Uehling-Uhlenbeck (SIBUU) transport equation using the test-particle method. We outline the physics foundation and technical approach, summarize our main results and identify key challenges in further studying spin transport within the SIBUU framework. As examples, we show how the global and local spin polarizations of nucleons, the spin splitting of nucleon collective flows as well as flows of light clusters at different spin states can be used to explore interesting new physics associated with the nuclear spin-orbit potential in dense neutron-rich medium. The spin-orbit potential and the rigorous angular momentum conservation are also shown to affect spin-averaged observables. Hopefully the new findings and challenges identified here will stimulate further studies both theoretically and experimentally on spin transport in intermediate-energy heavy-ion collisions.

nucl-th

An improved view of cosmic-ray transport and the galactic outflow in NGC 253

The nearly edge-on starburst galaxy NGC 253 exhibits extended multiwavelength halo emission, making it an ideal laboratory for studying disk-halo transport. We present improved ASKAP 943 MHz and MWA 216 MHz total-intensity images with resolutions of 13 and 45 arcsec and rms noise levels of 16 $\mu$Jy beam$^{-1}$ and 1 mJy beam$^{-1}$, respectively. After subtracting the thermal emission, we fitted the vertical synchrotron emission intensity and spectral-index profiles with one-dimensional advection and diffusion models. The ASKAP image reveals a loop-like structure in the northwestern radio spur extending to $\sim9$ kpc above the disk, while the southeastern spur reaches $\sim8$ kpc. The vertical profiles are best fitted by exponential components in the central region and Gaussian components in the outer regions, indicating advection-dominated CRE transport in the center and diffusion elsewhere. In the central region, the advection speed increases exponentially with height and reaches the estimated escape speed at about 5.5 kpc. The spatial correspondence with star-forming and X-ray-emitting regions indicates that CRE advection traces the bulk motion of the magnetized outflow. Below $\sim5.5$ kpc, the combined thermal, magnetic, cosmic-ray, and ram pressures exceed the estimated gravitational pressure, consistent with acceleration of the galactic wind. These results demonstrate the power of sensitive low-frequency radio observations for probing CRE transport and galactic outflows.

astro-ph.GA

Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning

We introduce Mobius-v0, an architecture that comprises a globally shared Memory (FFN) that stores knowledge vectors and multiple Reasoners (Self-Attn) that iteratively achieve compositional reasoning. Using hidden states as cache and carrier, reasoners repeatedly query memory for required knowledge-vectors, while the knowledge is transmitted back to reasoning operators. Through this knowledge-reasoning-separation architecture, Mobius achieves better knowledge compression and reasoning efficiency. Built upon Mobius-v0 architecture: 1) Our 7B model trained-from-scratch achieves similar downstream score as a 7B Transformer baseline with 62.6% of baseline's training data. 2) Our Intern-S2-Mobius, continually-pretrained from Qwen3.5-35B, achieves similar downstream score while delivering nearly 4x end-to-end inference speedup.

cs.AI

Intern-S2-Preview: Scientific Agentic Foundation Model

Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sustain progress across long task horizons. We present Intern-S2-Preview, a series of scientific agentic foundation models designed to support multimodal scientific understanding, reasoning, generation, and long-horizon tasks. The training pipeline begins with scientific multimodal pre-training over rendered scientific documents, interleaved image-text data, and diverse scientific corpora. Starting from the pretrained checkpoint, we apply a unified post-training pipeline consisting of supervised fine-tuning, scalable multi-task reinforcement learning (RL), black- and white-box agentic RL, and on-policy distillation. This pipeline is supported by practical techniques that improve rollout and training stability and efficiency, including partial rollout with off-policy correction, adaptive length regularization, online speculative decoding, robust multi-task optimization, and trace-aware experience assembly for agentic tasks. At the architecture level, Intern-S2-Preview-397B extends time series modelling from efficient long-sequence understanding to numerical forecasting, while Memory Decoder is studied as a separate memory-augmented path for rapid scientific specialization without modifying the frozen 397B backbone. Evaluations across scientific, multimodal, agentic, and general-purpose benchmarks show that Intern-S2-Preview-397B achieves competitive or leading results in multiple settings. The time series modules improve scientific signal understanding and forecasting on SciTS, while the separate Intern-MemDec-4B extension improves the Biology-Instructions average score from 56.92 to 60.32 without modifying the frozen 397B backbone.

cs.LG

ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation

Text-to-image (T2I) models can produce visually compelling images, yet they remain limited on open-world tasks that require complex semantic understanding, multi-step reasoning, and the integration of external world knowledge. Existing efforts introduce agent capabilities into image generation, but they either prescribe a fixed workflow or place only a subset of the open-world image generation process under agent control. Consequently, reasoning, tool invocation, and image generation are not coordinated by a single policy. We propose ToolArtist, a fully agentic image generation model obtained by post-training a Unified Multimodal Model (UMM). ToolArtist dynamically orchestrates reasoning, external tool use, and native image generation within one unified policy. During Supervised Fine-Tuning (SFT), we equip a teacher agent with search tools alongside an image-generation tool. We then convert the collected trajectories into a UMM compatible format, where the image-generation tool is concealed while the resulting generated images are retained. During Reinforcement Learning (RL), we develop an agentic RL infrastructure for UMMs and introduce Reason-Act-Draw GRPO (RAD-GRPO), which uses complementary intent and quality rewards to jointly optimize the model. Experiments show that placing the entire open-world image-generation process under an agent policy consistently outperforms approaches with fixed pipelines or only partially agent-controlled components. We release the training data and the complete post-training infrastructure.

cs.CV

TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning

Tool-Integrated Reasoning (TIR) enables LLMs to solve complex tasks through iterative tool interactions. However, existing reinforcement learning methods often rely on trajectory-level supervision, limiting fine-grained credit assignment in long-horizon TIR scenarios. On-policy self-distillation offers denser signals through teacher branches with privileged context, but existing approaches typically derive such context from ground-truth answers or retrieved skills, which may not reflect the states actually visited by the agent. Moreover, token-level supervision fails to capture the turn-level structure of tool interactions. To address this, we propose TurnSight, a turn-level hindsight self-distillation framework that derives supervision directly from execution-conditioned hindsight. It then constructs multiple hindsight views with different lookahead horizons and selects reliable supervision through cross-horizon directional agreement. Finally, the selected hindsight signal is normalized across sibling rollouts and used to adaptively modulate RL advantages while preserving their original optimization direction. Extensive experiments on three benchmarks demonstrate the effectiveness of TurnSight. Our codes are available at https://github.com/quchangle1/TurnSight.

cs.CL

KuaiLive-M3: A Multi-Modal, Multi-Domain, and Multi-Feedback Dataset for Live Streaming Recommendation

Existing public live streaming datasets suffer from three major limitations: they provide limited access to temporally evolving multimodal live content, overlook users' cross-domain interactions between short videos and live streams, and contain only implicit behavioral signals without explicit feedback that captures users' perceived content quality and satisfaction. These limitations prevent existing benchmarks from faithfully reflecting real-world live streaming scenarios and hinder comprehensive research on live streaming recommendation. To address these limitations, we introduce KuaiLive-M3, a multi-modal, multi-domain, and multi-feedback dataset for live streaming recommendation, collected from Kuaishou, a leading live streaming and short video platform in China. KuaiLive-M3 covers 21,938 users and contains 35 million live streaming interactions and 111 million short video interactions, with fine-grained timestamps and diverse user behaviors. It further provides approximately 88 million timestamped segment-level multi-modal embeddings that capture the temporal evolution of live streaming content, as well as 25,403 questionnaire-based feedback records that bridge implicit user behaviors and explicit user preferences. Based on these unique signals, we establish benchmarks for cross-domain recommendation, live stream highlight prediction, and questionnaire-enhanced recommendation. Extensive experiments with representative baselines demonstrate that KuaiLive-M3 provides a challenging and realistic benchmark for future live streaming recommendation research. The results further highlight the importance of modeling temporally evolving content, transferring user preferences across domains, and bridging the gap between implicit behaviors and explicit user feedback. The dataset and benchmark code are publicly available at https://imgkkk574.github.io/KuaiLive-M3/.

cs.IR

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks

Group-based policy optimization has been increasingly used to train large language model (LLM) agents from sparse outcome rewards by comparing trajectories or steps within a group. However, on difficult long-horizon tasks, this comparison can suffer from a sampling imbalance: repeated or low-effect actions dominate the high-probability region of the policy while useful state-changing actions remain under-sampled. This imbalance produces many all-failed rollout groups, where outcome rewards provide no direction for correcting the policy. Together, these effects can form a self-reinforcing credit trap: failure-dominated sampling yields no outcome-based correction, allowing repeated low-effect actions to persist. To break this loop, we propose Progress-conditioned Group Policy Optimization (ProGPO), which uses first-visit observation coverage only when all samples in a group receive zero outcome reward. Specifically, within such groups, ProGPO assigns higher relative advantages to trajectories or steps that visit more new states since reaching new observations is a prerequisite for task success. Experiments on two challenging agentic benchmarks, ALFWorld and WebShop with Qwen2.5-1.5/7B-Instruct, show that ProGPO consistently improves over group-based baselines, with particularly large gains on hard tasks.

cs.LG

pyHB: an open-source automatic-differentiation-enhanced semi-analytical solver for nonlinear dynamics

The Harmonic Balance (HB) method is widely used to compute and analyze the periodic responses of nonlinear systems. However, its application to high-dimensional complex systems is limited by the burden of handling the partial derivatives of the nonlinearities. This work presents pyHB, an open-source, automatic-differentiation-enhanced semi-analytical framework that integrates the complete HB workflow for general user-defined nonlinear systems. The proposed formulation exploits localized nonlinearities and applies PyTorch-based automatic differentiation (AD) only to the reduced nonlinear force, thereby avoiding the need for user-supplied derivatives of the nonlinear force and maintaining controllable GPU memory usage. Weighted arc-length continuation, sparse matrix assembly, a blocked solution strategy for the augmented continuation equations, and Floquet-based stability analysis are incorporated within a modular architecture that separates model definition from reusable numerical procedures. Hence, pyHB can provide a complete landscape of the nonlinear system's periodic response based solely on the user-defined dynamical equations. Four examples, including a quasi-zero-stiffness isolator, a nonlinear piezoelectric energy harvester, a 284 degrees of freedom (DOFs) aeroengine model, and a 2000 DOFs Bernoulli beam, demonstrate the ability of pyHB to trace stable and unstable solution branches and capture subharmonic resonance, combination resonance, and mixed-order electromechanical responses. Notably, in the Bernoulli beam example with 202000 HB unknowns, the AD-enhanced solver requires approximately 0.44s per continuation point, achieving several-hundred-fold speedup compared to the Newmark-$\beta$ method and remaining 637.8MB of additional RAM and 243.5MB of GPU memory. The proposed pyHB provides a general, one-stop benchmark platform for HB-based nonlinear dynamics analysis.

cs.MS

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tightly coupled, hindering reproducibility and causing redundant engineering. To address this, we introduce AgentCompass, an open-source, lightweight, and extensible infrastructure for evaluating LLM-based agents. AgentCompass organizes the evaluation process around three independent components, namely Benchmark, Harness, and Environment, thereby enabling flexible configurations without requiring the reimplementation of complex execution logic. Furthermore, it features a fault-tolerant asynchronous runtime and comprehensive trajectory analysis tools to transparently diagnose nuanced failure modes like reward-hacking. Natively supporting over 20 benchmarks across five capability dimensions, AgentCompass provides the community with a scalable and reproducible infrastructure for advancing agent research.

cs.AI

MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation

Large language models (LLMs) are increasingly deployed in online medical consultation, yet existing benchmarks remain poorly aligned with real clinical practice. Many rely on synthetic conversations or patient simulators, omit patient-uploaded medical images, or evaluate open-ended clinical responses using multiple-choice or lexical-overlap metrics that poorly reflect clinical quality. We introduce \textbf{MedRealMM}, a large-scale benchmark for multimodal online medical consultation built from de-identified patient-doctor interactions collected from a nationwide Chinese internet hospital. MedRealMM uses a Multimodal Clinical Challenge Point (MCCP) extraction framework to identify clinically demanding moments in authentic consultation trajectories and converts each into a standardized next-response generation task while preserving the preceding text-image context. Each instance is paired with a case-specific rubric refined by physicians that rewards clinically desirable behaviors and penalizes unsafe, unsupported, or contradictory responses. The current release contains 5,620 real-world multimodal cases spanning 64 clinical departments. We evaluate 19 general-purpose and medical-specialized LLMs, including text-only and multimodal systems. Our results show that image information is critical for reliable clinical performance and that current frontier models remain below the online physician response. Although some frontier models satisfy as many or more positive clinical criteria than physicians, they trigger more negative criteria, indicating that safety-sensitive error avoidance remains a central bottleneck. MedRealMM offers a realistic and reproducible benchmark for evaluating multimodal medical reasoning in real-world online consultation. The dataset will be publicly available on Hugging Face at https://huggingface.co/datasets/jdh-algo/MedRealMM.

cs.AI

Enhanced two-photon sources in a cavity-coupled two-atom system

We propose a component-selective scheme for improving two-photon sources in a cavity-coupled two-atom system, where a single cavity mode interacts with two two-level atoms driven by phase-controlled classical fields of the same frequency. By controlling the atomic detunings and driving phase, the system can be tailored toward optimized cavity-field two-photon blockade or strongly correlated fluorescence photon-pair emission. When the two-cavity-photon component is enhanced, the cavity field exhibits optimized two-photon blockade with simultaneous suppression of unwanted one- and three-photon components at a comparable two-photon population. In another parameter regime, strongly correlated fluorescence photon pairs can also be generated from the two atoms by selecting the double-atomic-excitation component in the same two-excitation manifold. This approach provides a route toward high-quality and versatile two-photon sources, with potential applications in few-photon quantum optics and quantum information processing.

quant-ph

Autonomous Information Seeking: A Roadmap for Agentic Recommender Systems

The rapid integration of large language model-based agents into recommender systems has driven a shift from static, ranking-based pipelines toward autonomous and interactive systems that can reason, plan, and act. This survey provides a comprehensive overview of this emerging landscape by introducing a unified taxonomy grounded in the level of autonomy and three core paradigms of agentic recommender systems: agent-assisted recommendation, agent-as-recommender, and agent-as-user-simulator. The autonomy framework organizes existing methods along increasing capabilities in proactivity, context awareness, interaction flexibility, and adaptivity. Building on this framework, the survey analyzes how each paradigm adopts different agentic architectures and how agents enhance key components such as profiles, memory, tool use, workflows, and optimization mechanisms. We further examine evaluation methodologies for agentic recommendation, covering automated metrics, LLM-based judging, and simulation-based assessment, and discuss their limitations in capturing reasoning quality, user experience, and system behavior. Beyond existing evaluation protocols, we further discuss unresolved issues in evaluating agentic recommender systems, including trajectory-level assessment, agent contribution analysis, and calibration of user simulation. Lastly, the survey outlines open challenges in lifelong user modeling, contextual abstraction, multimodal alignment, controllability, trustworthiness, privacy, scalability, and efficiency. Together, these analyses establish a unified foundation for understanding the current progress of agentic recommender systems and highlight promising opportunities for developing more autonomous, reliable, and human-aligned recommendation agents.

cs.IR