SearcharxivSearch

arXiv subjects

Yang Sun

Publications and source records attributed to Yang Sun.

At least 19 recordsLinked to original sources

Learning Human Health and Diseases from 24-hour Wrist Movement

Much of human health and function unfolds beyond the clinic, through the movements of everyday life. Wrist-worn accelerometers capture these movements continuously, yet their rich signals are often reduced to a small set of predefined behavioural summary measures. Here, we present Sensori, a self-supervised foundation model that learns general-purpose health representations directly from 24 hours of raw tri-axial wrist movement. We developed and evaluated the model across four population-based cohorts from the United Kingdom, China and the United States, comprising 122,640 participants contributing 683,617 person-days of free-living recordings. Sensori condensed each day of movement into a representation that captured diverse movement behaviours, demographic characteristics, health axes and physical function. Evaluation in independent cohorts showed that these representations generalised across populations and measurement settings without retraining. When added to common clinical covariates, Sensori significantly improved prevalent disease classification for 52 of 102 eligible conditions (median delta AUROC, 0.060; range, 0.012-0.242) and incident disease risk prediction for 26 of 87 eligible conditions (median delta Uno's C-index, 0.064; range, 0.025-0.172), with the largest gains for neurological and psychiatric disorders. These findings establish 24-hour wrist movement as a rich and scalable source of health information, with the potential to support passive health monitoring and disease prediction at population scale.

cs.LG

No Silver Bullet: Boosting GaussDB Performance on the 30TB TPC-H Workload

GaussDB is Huawei's premier database system, designed for large-scale deployments and the most demanding workloads. It is a distributed shared-nothing system, capable of handling all types of workloads. This paper outlines a series of modifications to GaussDB aimed at improving its performance on large-scale and complex analytical workloads. After these changes, its performance on the TPC-H workload exceeded the best published result by 40% at 30 TB. The key enhancements to achieve this elite performance include adopting a pipeline execution model, a faster and more scalable inter-node data shuffle, exploiting a unified bus and unified remote memory access. We also expanded the support of cost-based Bloom filter placement and implemented several Bloom filter streaming strategies, enabling their use across nodes.

cs.DB

Quantifying the effects of nickel on Earth's inner-core nucleation

The formation of Earth's solid inner core marks a major transition in the thermal and chemical evolution of the deep Earth, yet its origin remains paradoxical, as initial nucleation appears to require unrealistically large undercooling in the outer core. Here, we use atomistic simulations to quantify how Ni affects this process under inner-core conditions. While Fe-Ni alloys preserve strong thermodynamic competition between the hcp and bcc phases, the bcc phase consistently forms smaller critical nuclei and has lower nucleation barriers than hcp. Increasing Ni content in the melts further lowers the nucleation barrier and shortens the nucleation waiting time. Local chemical fluctuations also strongly affect the macroscopic nucleation rate. Combining these effects, bcc nucleation in Fe80Ni20 reaches about 250 K of undercooling, approaching geophysical constraints. We demonstrate that Ni enrichment, bcc nucleation, and chemical heterogeneity substantially narrow the inner-core nucleation paradox.

physics.geo-ph

High-throughput identification of ferromagnetic Kagome candidates in the AT6X4 and AT6X5 families

We present a systematic high-throughput density-functional theory study of the thermodynamic stability, collinear magnetic ground states, and electronic structures of layered kagome compounds in the AT6X4 and AT6X5 families. Using the experimentally reported structure types as templates, we screened 78 substitutional compositions in each family. Our calculations reproduce the stability and antiferromagnetic character of the known Fe-based Ge compounds and identify six additional stable candidates with robust ferromagnetism. Within collinear spin configurations, we find a clear chemistry-dependent trend: stable Fe-based Ge compounds predominantly adopt AFM2 ground states, whereas stable Mn-based Ge compounds consistently favor ferromagnetic order. Exchange analysis further shows that the magnetic phase space is governed by competing interlayer interactions, consistent with the mechanism established for AT6X6 kagome magnets. Representative ferromagnetic members from the two structural families also retain kagome-derived dispersive band features near K, although the AT6X5 phase exhibits stronger band folding and hybridization. Overall, these results establish AT6X4 and AT6X5 as promising layered kagome families for realizing ferromagnetism and kagome-derived electronic states.

cond-mat.mtrl-sci

VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?

Constructing an interactive 3D open world from a user query is important. However, existing methods are primarily evaluated on idealized, simple queries, making it difficult to systematically analyze and compare how multimodal agents understand user intent, use 3D tools, and reason over textual and visual 3D world information. To this end, we propose VibeWorlding, a unified framework for benchmarking and training vibe worlding agents: a multimodal agent that can autonomously infer user intent, plan scene layout, invoke 3D tools, and reflect on the multimodal feedback in a multi-turn agent-environment interaction process. To achieve this, we first build VWE-BENCH, a benchmark of 2,616 high-quality 3D assets, 323 human-annotated seed 3D worlds, and 6,828 reverse-synthesized multimodal user queries, split into verified queries with ground-truth and unverified queries with carefully designed rubrics. Moreover, we develop VibeWorlding-Gym, a joint multimodal RL post-training framework that integrates (1) a sandbox environment unifying asset retrieval, editing, and image rendering as MCP tools, and (2) a rubric-based verifier that combines physical feasibility and intent fulfillment verification, supporting both fair model evaluation and scalable multimodal RL reward service. Our experiments show that current frontier MLLMs are far from solving the vibe worlding agent task, with even GPT-5.5 and Qwen3.8-Max reaching below 60% success rate, and trace the bottleneck to precise 3D world editing. We further find that RL training can ease this weakness and enable open-source MLLMs to even surpass closed-source frontiers: our VibeWorlder-8B is comparable to frontier MLLMs, while our flagship VibeWorlder-30B-A3B attains the best overall Pass@1 among all evaluated models.

cs.AI

Reconfigurable microwave photonic Fano filters based on optical Kerr microcombs

Microwave photonic (MWP) Fano filters, featuring asymmetric filter shapes that enable steep spectral transitions, are attractive for high bandwidth microwave signal processing such as frequency discrimination. However, achieving both steep spectral transitions and a high degree of reconfigurability remains challenging for conventional methods relying on direct mapping of Fano resonances generated by optical filters. Here, we propose and experimentally demonstrate a new way for realizing MWP Fano filters based on a microcomb-driven transversal filter system. Leveraging the large number of comb lines provided by microcombs as discrete taps, the transversal filter system can synthesize filter response that closely resembles Fano resonances, yielding high rolloff rates and slope rates up to 33.8 dB / GHz and 25.7 dB / GHz in our experiments, respectively. In addition, by simply programming the tap coefficients without changing any hardware, highly reconfigurable filter response can be realized. We experimentally demonstrate independent tuning of all three Fano characteristic parameters, including the asymmetry factor, resonance linewidth, and center frequency. These results verify the effectiveness of our approach for implementing highly reconfigurable MWP Fano filters with steep spectral transitions, offering strong versatility for meeting diverse requirements in practical applications.

physics.optics

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation

On-policy distillation (OPD) trains a student on its own trajectories under dense token-level supervision from a teacher. Reward-extrapolation methods such as ExOPD amplify the teacher-reference log-likelihood ratio to move beyond direct imitation, but apply a single global coefficient $\lambda$ to every token. This can drive the student to fit extreme peaks in the implicit reward, causing reward hacking and unstable training, and the optimal $\lambda$ varies across domains, requiring costly sweeps. We propose REOPD, a reliability-adaptive reward extrapolation framework for OPD. REOPD combines a token-level compatibility weight with a batch-level adaptive budget, yielding a token-wise coefficient $\lambda_{b,t}=1+\gamma_b q_t$ that preserves teacher alignment while selectively extrapolating along reliable teacher-reference directions. It requires no verifier, reward model, value model, or extra rollout beyond standard OPD. REOPD outperforms G-OPD on single-teacher mathematics and on both domains in the multi-teacher setting, while matching G-OPD on single-teacher code, demonstrating effective fine-grained reliability adaptation across domains and teacher configurations.

cs.LG

Projected shell model description of nuclear level density: Collective, pair-breaking, and multiquasiparticle regimes in even-even nuclei

There is overwhelmingly experimental evidence indicating that excited nuclear states are dominated by quasiparticle (qp) excitations, which form many-body configurations with broken nucleon-pairs from different orbitals. By using these multi-qp states as building blocks for a shell-model basis, we propose a novel shell-model method to calculate the nuclear level density (NLD) in deformed nuclei. The shell-model diagonalization with two-body residual interactions yields a large ensemble of eigenstates of angular momentum and parity. We demonstrate that NLD as a statistical quantity depends sensitively on the structure of deformed single-particle states. As the first example to introduce this method, we take a well-deformed rare-earth nucleus, $^{164}$Dy, for which NLD has been studied extensively by the Oslo method. By a quantitative comparison with discrete levels from spectroscopic measurements, we show that while the pronounced stepwise structure in the low-energy NLD curve can be understood as the collective excitation and nucleon-pair breaking, the exponential growth of levels in the higher-energy NLD can be described by the combination of the broken-pair states, subject to the Pauli principle. According to the nature of NLD with increasing excitation, we divide the entire NLD curve into (1) collective regime, (2) pair-breaking regime, and (3) multi-qp regime. We discuss the formation mechanism and characteristic features of NLD for the three regimes. In addition, the parity dependence and angular-momentum dependence in NLD are investigated with a strong emphasis on the structure effect.

nucl-th

Nuclear level density studied in odd-mass nuclei in the framework of the projected shell model

In a recent article [Phys. Rev. C 108, 034309 (2023)], we proposed a projected shell model method for the calculation of nuclear level density (NLD) in deformed even-even nuclei. The current article presents the subsequent study of NLDs in odd-mass nuclei as well as a comparative analysis between our calculated NLDs in adjacent even-even and odd-A systems. Since one nucleon in the odd-mass system remains blocked from participating in the pair formation, resulting in a weakened pairing (assessed by a smaller BCS pairing gap), pronounced differences between the NLDs in an odd-mass (both even-odd and odd-even) nucleus and its immediate even-even neighbour have been found. In general, the structure-dominated variations, which were found to be prominent in the even-even NLD at low energies, are greatly suppressed in the odd-mass systems. Specifically, from excitation energy as low as 2 MeV, the calculated densities of odd-parity and even-parity levels in odd-mass nuclei show an equal division signaling faster attainment of the statistical behavior. Nuclear level-spin distributions of both parities have been seen to adopt a regular Gaussian shape earlier than that found in the even-even system. Moreover, the pleasant property of our shell-model results, that each of our calculated levels is an eigenstate of angular momentum, allows us to extract the values of the energy-dependent dispersion $\sigma$ of Ericson's spin-distribution formula and plot $\rho(E, I, \pi)$, the energy-, spin-, and parity-dependent level density.

nucl-th

Long-Horizon Agent Trajectory Attribution: A Unified Benchmark and Fine-Grained Annotation Framework

Large language model (LLM) agents increasingly operate through long-horizon trajectories involving user instructions, tool use, external observations, and memory. Existing benchmarks primarily evaluate behavioral outcomes but provide limited support for fine-grained attribution analysis. We introduce trajectory attribution and develop a benchmark and annotation framework for this task. The benchmark organizes heterogeneous trajectories under a unified component schema and provides annotations of the primary attribution component, together with attack and execution chains where applicable. Instantiating the benchmark with trajectories from AgentDojo and the Stage and Canary settings of Agent3Sigma yields more than 1,300 annotated trajectories covering task-aligned actions, unsafe actions, and safety refusals. The benchmark defines two evaluation tasks, primary attribution localization and attribution-chain recovery, and provides reference baselines based on incremental trajectory contribution and component-level leave-one-out perturbation. It captures diverse attribution settings, including local and long-range attribution as well as structured attribution chains. Reference baseline results exhibit substantial performance differences across these settings, providing an initial characterization of the benchmark's attribution challenges. Beyond this initial instantiation, we release a reusable annotation skill that enables trajectories generated by new agent models to be standardized, annotated, and evaluated under the same framework. Project resources and future releases are available at https://github.com/chenjing-2024/agent-trajectory-attribution.

cs.AI

How Much, Then Where: Credit-Conserving Action-to-Token Allocation for Multi-Turn Agent Reinforcement Learning

Credit assignment in multi-turn agent reinforcement learning operates at two levels: assigning trajectory-level credit to actions and distributing each action's credit across its tokens. In this paper, we introduce FACTOR, which separates these decisions. FACTOR uses checkpoint-calibrated TD residuals to assign per-action credits that telescope to the trajectory advantage, and feedback-conditioned teacher-student likelihood gaps to allocate each credit across the realized action tokens. Per-action normalization preserves the action-average coefficient and prevents token-level sign flips. We pair this construction with an action-mean reduction, removing the implicit dependence of an action's scalar surrogate weight on its token length. At the behavior policy and before clipping, each action's inner action-mean surrogate equals its TD credit. FACTOR consistently improves over competitive baselines across ALFWorld, WebShop, and ScienceWorld, with every environment-seed comparison favoring FACTOR and the largest gains emerging on the longest-horizon environment. The same hyperparameters transfer without retuning to a larger backbone and to a different model family. Ablations identify TD action credit as the dominant driver of the improvement, with hindsight token allocation contributing complementary gains.

cs.AI

Antiferromagnetic Phases in Zr-Fe-Ge Kagome Systems

A wide variety of chemical substitutions in ferromagnetic Kagome systems can lead to diverse magnetic phases with electronic structures suitable for topological or quantum material properties. Here, we study the electronic structure and magnetic orderings using first-principles calculations for the magnetic Kagome compounds ZrFe6Ge6, ZrFe6Ge4, and ZrFe6Ge5. For ZrFe6Ge6, the obtained ground-state magnetic structure is A-type antiferromagnetic (AFM), in agreement with existing experiments. We predicted that the magnetic ground states of ZrFe6Ge4 and ZrFe6Ge5 are collinear A-type bilayer AFM structures with long-period ordering that involves a mix of FM and AFM interlayer orientations. The formation of such long-range magnetic structures appears to be a general feature and is not tied to specific substitutions. The magnetic moments in these systems are largely local and only weakly dependent on the magnetic configuration, with magnitudes in good agreement with available experimental estimates. Neutron scattering experiments, which could provide direct verification of these predictions, are therefore of particular importance.

cond-mat.mtrl-sci

Determining Total Infrared Luminosities from Submm Measurements of High Redshift Galaxies

Determining total infrared luminosities for very high redshift galaxies is important to estimate the rate of star formation in heavily dust-embedded environments. It is also challenging because the most sensitive far infrared observatory, Herschel, was limited in sensitivity and its deepest measurements are subject to confusion noise. Thus, these determinations largely depend on ALMA, observing in the mm- and/or submm-wavelengths, which sample only the long-wavelength part of the spectral energy distributions (SEDs). Luminosities are conventionally estimated with modified blackbody fits to these measurements, but do not include the emission in the mid-infrared by warmer dust; there is evidence that this mid-IR component may be relatively strong in very high-redshift galaxies compared with local ones. A correction factor must be applied to the modified black body luminosities to derive the total infrared luminosity. We study infrared SEDs using simulations tuned to galactic conditions typical of high-z galaxies. We find that the different behaviors of infrared SEDs are dominated by a single key physical parameter, the luminosity density. This allows us to estimate the corrections for the missing mid-infrared luminosity in a general way. We find that a factor of ~ 1.6 - 1.7 (0.2 dex) is appropriate in most circumstances, with a larger factor of ~ 1.75 - 1.85 (~ 0.25 dex) up to 2 (0.3 dex) necessary for high redshift (z > 4) galaxies at the highest luminosities, > 10^{12} Lsun. These corrections are needed to estimate star formation rates based on total infrared luminosity.

astro-ph.GA

A quasar hatching from a buried red phase at z = 3.7

We present JADES-GS 209777, previously cataloged as CANDELS J033238.02-274626.2, hereafter "the Hatchling," a red quasar at $z=3.711$. While the source has been reported in earlier deep-field catalogs, our multiwavelength analysis reveals a visible active nucleus still embedded in a dense gas- and dust-rich environment. Red quasar continua are often attributed to dust attenuation, including non-standard extinction curves, but the highly comprehensive multiwavelength data for this source provide direct constraints on the material being cleared. Using JWST/NIRSpec, NIRCam, MIRI, HST, MUSE, Chandra, ALMA, and VLA data, we detect broad emission lines and strong X-ray emission, showing that the active nucleus is at least partially exposed. We also detect H$\alpha$ and He I absorption, indicating dense gas close to the nucleus. Kinematically disturbed O I, Mg II, Na D, and [O III] features, together with extended Ly$\alpha$ emission over $\gtrsim 20$ kpc, further show that multiphase gas is being accelerated from the nuclear region into the host-galaxy environment. The ALMA detection reveals strong dust emission, with the inferred infrared luminosity placing the system in the ULIRG regime. The continuum is red and sharply declining toward the rest-frame UV, resembling compact red AGNs, and may reflect extreme dust attenuation, gas reprocessing, possible BAL-like absorption, or a combination of these effects. Regardless of which mechanism dominates the continuum shape, the line diagnostics show that the visible nucleus remains partially obscured by nearby material. The Hatchling therefore represents a unique opportunity to explore a poorly known transition phase in which feedback is likely clearing an enshrouded quasar and allowing it to emerge toward a more unobscured active nucleus.

astro-ph.GA

PhotoIFU: NIRCam as a Photometric Integral Field Unit for Mapping Feedback in Galaxies

We present PhotoIFU, a workflow that uses deep multi-band imaging as a low-resolution photometric integral field unit. Applied to PSF-matched JWST/NIRCam imaging, PhotoIFU treats each spatial pixel as a coarse SED element and fits the pixel SEDs with Prospector to map resolved stellar-population and ISM-related properties. We apply this approach to three galaxies at $z=1.3$--3.7 in JADES: two systems with extended ionized line emission and one post-starburst galaxy with an exceptionally strong neutral outflow. Pixel-by-pixel SED fitting gives maps of stellar-mass surface density, specific star formation rate, dust attenuation, gas-phase metallicity, and recent star-formation history. We find that regions selected from the extended-emission or outflow geometry occupy distinct parts of the resolved SED-property distribution compared with the full host. In the systems with extended ionized emission, these regions are generally less dusty, consistent with ionized emission being observed along dust-poor, low-column-density pathways through the host. In the neutral-outflow system, the selected regions show enhanced recent star formation, suggesting that compact rejuvenation may mark the aftermath of an earlier energetic phase. These results show that galactic outflows and extended emission-line structures can be spatially associated with measurable differences in resolved host-galaxy stellar populations and ISM-related properties. PhotoIFU provides an imaging-based method for resolved SED mapping of feedback-related structures in larger galaxy samples where full spectroscopic integral-field mapping is unavailable.

astro-ph.GA

Alchemical thermodynamic integration for ab initio free-energy calculations in solutions

We develop an alchemical thermodynamic integration scheme that couples ab initio force calculators on the fly during Monte Carlo and molecular dynamics simulations. The implementation is validated against existing hybrid-Hamiltonian approaches. The scheme yields ab initio free energies of high-pressure Fe-Ni and ambient Li-Na liquid solutions that agree with previous calculations and reproduce the experimentally observed Li-Na miscibility gap. The code has an efficiency comparable to standard ab initio molecular dynamics. These results establish this scheme as a practical alchemical-integration framework for first-principles free-energy calculations in solutions.

cond-mat.mtrl-sci

Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning

Egocentric videos of human manipulation provide scalable supervision for embodied intelligence, yet existing resources rarely combine low-cost continuous capture, manipulation-level structured annotations, and reusable tools for robot learning. We present Open-AoE, an open, community-oriented egocentric manipulation dataset and toolchain spanning the full pipeline from smartphone capture to model training. Its first release contains approximately 2,000 hours of manipulation video collected in natural environments by 500+ contributors using 400+ smartphones. The dataset provides text annotations, MANO-based hand poses, camera trajectories, and temporally localized atomic actions. Open-AoE further includes a data processing pipeline that transforms raw recordings into structured samples through temporal action segmentation, semantic annotation, hand reconstruction, and camera trajectory reconstruction. Meanwhile, we provide a separate downstream toolchain supports visualization, cross-embodiment retargeting, model-specific data conversion, and training recipes for VLA policies, WAMs, and World Models. By integrating scalable capture, structured processing, and downstream adaptation, Open-AoE reduces the barriers to both data contribution and reuse, providing practical open infrastructure for embodied model training, human-to-robot transfer, and world modeling.

cs.RO

SemFlowRAG: Directed Semantic Flow from Abstraction to Evidence for Complex Reasoning

Retrieval-Augmented Generation (RAG) enhanced by Knowledge Graphs has shown promise in complex multi-hop reasoning tasks. However, existing graph-based retrieval methods typically rely on flat, undirected topologies. During the retrieval process, the probability flow often gets trapped in high-degree abstract concept nodes which we define as ``probability black holes'', leading to semantic drift and noise accumulation. To address this, we propose SemFlowRAG, a framework that reconstructs the flat retrieval space into a corpus-adaptive semantic gradient graph. This data-driven self-organization enables a hierarchical structure to emerge naturally from the data distribution, capturing the intrinsic semantic granularity of the corpus to suppress structural noise. By quantifying the semantic abstractness of entities through the embedding variance of their associated passages, we transform static undirected edges into directed semantic constraints. Furthermore, we design an abstractness-guided directed PageRank algorithm that forces the retrieval trajectory to follow a ``high-to-low semantic abstractness'' gradient. This mechanism ensures layer-by-layer evidence convergence, smoothly guiding the retrieval process from abstract concepts to specific document evidence. Extensive experiments on complex QA datasets demonstrate that SemFlowRAG effectively mitigates the ``probability black holes'' issue, outperforming existing baselines in both retrieval and downstream reasoning performance.

cs.IR