SearcharxivSearch

arXiv subjects

Shaolong Chen

Publications and source records attributed to Shaolong Chen.

12 recordsLinked to original sources

FedSubMuon: Communication-Efficient Federated LLM Fine-Tuning via Structured Subspace Muon

Federated fine-tuning adapts large language models (LLMs) to decentralized client data, but its scalability in cross-device training is often limited by the high communication cost. Muon is an optimizer that improves optimization performance by orthogonalizing momentum for matrix-valued parameters. Existing federated Muon methods demonstrate the benefit of matrix-aware optimization in federated learning, but still require transmitting full layer-size updates and optimizer state. A natural way to reduce communication is to directly apply Muon to LoRA factors, but this changes the optimized object and weakens Muon's matrix-aware update geometry. We propose FedSubMuon, a communication-efficient federated Muon fine-tuning method that optimizes compact coefficient matrices within shared structured subspaces. This design keeps Muon on a single matrix-valued trainable object, while reducing the client upload to compact coefficient matrices. We further introduce FedSubMuon-GT, an accuracy-oriented extension that uses projected gradients to adapt tracked subspace bases toward task-relevant gradient directions. Experiments on instruction tuning and mathematical reasoning show that FedSubMuon-GT achieves the best overall accuracy on four of five dataset-model pairs, while FedSubMuon performs best under all matched communication budgets. On Dolly-15K, the closest communication baseline requires 5.5 times and 1.4 times more total communication on Llama-1B and Qwen-4B, respectively.

cs.LG

Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies

Can a language model recover the true research idea of a published paper when given only that paper's pre-publication bibliography? We introduce Reconstruction, a blind idea-recovery benchmark that withholds the seed paper and all contemporaneous or future literature, and asks models to propose hypotheses that an independent large language model judge matches against the held-out ground-truth idea. A strict anti-leakage protocol-temporal citation cutoff, anonymous reference IDs, and frozen per-paper bibliographies, which prevents prompt-time leakage of the seed idea. Across six scientific domains and 643 evaluated papers, seven frontier models achieve only modest Match rates (approx. 3-15%). We then evaluate a reference-only multi-agent (top 4) pipeline that combines cross-model review with a Swiss tournament over aligned hypothesis slots, without external web search. Cross-model review plus tournament selection raises Match rates to approx. 23-42% across all six domains, which is an observed approx. 2.4x lift over the best single-model baseline. This draft reports the protocol, anti-leakage design, and current results as an arXiv timestamp.

cs.AI

How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks

AI has long assisted scientific research, but the rapid advance of LLMs and agentic scaffolds is reshaping the landscape; a single system can now carry whole-stage research from an initial hypothesis all the way to final published paper, which is a paradigm now referred to as AutoResearch. Existing evaluations reveal little about how these agents operate or where they break down. Tasks are narrowly-scoped, evaluation measures performance but not process, and failure diagnoses lack systematic coverage or artifact-level visibility. To address this gap, we introduce AutoResearchEval, featuring 100 tasks grounded in published frontier science across 7 scientific domains and the full research lifecycle, including ideation, retrieval, execution, analysis, writing, and review. Evaluating 8 harness-model combinations yields 800 autoresearch agent trajectories, with process-level annotation. We organize these insights into AutoResearch Failure Taxonomy or ARFT, a framework of 45 empirically-grounded failure patterns. To enable scalable fine-grained attribution, we leverage a human-calibrated agent-as-a-judge pipeline to inspect complete trajectories and intermediate artifacts. Failure patterns converge on a single overarching limitation, namely that current agents lack a metacognitive loop, which entails the ability to check what they produced against what they found, revise when it does not hold up, and question whether the path they took was sound. The same patterns recur across all 8 harness-model combinations, including the strongest models tested, locating the deficit at the model level rather than in any particular scaffold; whether orchestration-level interventions can close it is an open question this work does not test. We publicly release AutoResearchEval and ARFT to facilitate continued research and development in autonomous scientific discovery.

cs.CL

AdaDPO: Self-Adaptive Direct Preference Optimization with Balanced Gradient Updates

DPO has become a widely adopted alternative to RLHF for aligning LLMs with human preferences, eliminating the need for a separate reward model or RL loop. Recent theoretical analysis uncovers an asymmetric gradient behavior in DPO: the loss suppresses dispreferred responses substantially faster than it promotes preferred ones, causing the model to learn to avoid bad answers rather than to generate good ones. We propose AdaDPO, a Self-Adaptive variant of the DPO algorithm that introduces per-preference-pair, stop-gradient-based coefficients derived directly from the policy model's generation probabilities, with the reference model's probabilities as an optional component. AdaDPO is constructed to enforce equality of gradient magnitudes between preferred and dispreferred probabilities; the practical implementation balances per-token gradients and applies a numerical clipping bound for stability, while retaining DPO's original hyperparameter structure. On Llama-3-8B-Instruct trained on UltraFeedback under a SimPO similar setup, AdaDPO consistently outperforms DPO on AlpacaEval 2: it achieves higher length-controlled win rates (LC) in 81% of hyperparameter combinations, attains the global best LC (48.3%) and raw win rate (46.1%), and enlarges the LC-over-WR margin in 88% of combinations, indicating effective mitigation of length bias. Additional analyses on KL divergence, reward margin, and reward accuracy confirm that AdaDPO rectifies the gradient imbalance and yields more efficient optimization. Because it operates purely at the loss level, AdaDPO can be dropped into existing preference-based alignment pipelines without changing data collection or model architectures. The method requires only a few lines of code, and the same self-adaptive principle generalizes to a broad family of pairwise contrastive preference losses including SimPO, R-DPO, IPO, CPO, and ORPO.

cs.CL

Coulomb Crystallization of Highly Charged Ni^12+ Ions in a Linear Paul Trap

Optical clocks have garnered widespread attention due to their unparalleled precision in time-frequency standards, geodetic measurements, and fundamental physics research. Among emerging developments, highly charged ion (HCI)-based optical clocks have attracted significant scientific interest owing to their exceptional resilience against electromagnetic perturbations and enhanced sensitivity to variations in the fine-structure constant ($\alpha$). While the recent successful demonstration of an Ar$^{13+}$ optical clock has validated the feasibility of HCI-based systems, Ni$^{12+}$ -- featuring an ultranarrow clock transition linewidth -- stands out as a superior candidate for achieving HCI optical clocks with $10^{-19}$ level uncertainty and stability. In this work, we report the Coulomb crystallization of nickel highly charged ions (Ni-HCIs). Through a precision deceleration and sympathetic cooling protocol in a room-temperature Paul trap, high-energy Ni-HCI bunches were sympathetically cooled from megakelvin to the 100-millikelvin range using laser-cooled Be$^{+}$ ions. This work represents a pivotal step toward the realization of an optical clock based on the Ni$^{12+}$ ion.

physics.atom-ph

Precision Measurement of M1 Optical Clock Transition in Ni12+

Highly charged ions (HCIs) have drawn significant interest in quantum metrology and in search for new physics. Among these, Ni12+ is considered as one of the most promising candidates for the next generation of HCI optical clocks, due to its two E1-forbidden transitions M1 and E2, which occur in the visible spectral range. In this work, we used the Shanghai-Wuhan Electron Beam Ion Trap to perform a high-precision measurement of the M1 transition wavelength. Our approach involved an improved calibration scheme for the spectra, utilizing auxiliary Ar+ lines for calibration and correction. Our final measured result of the M1 transition wavelength demonstrates a five-fold improvement in accuracy compared to our previous findings, reaching the sub-picometer level accuracy. In combination with our rigorous atomic-structure calculations to capture the electron correlations and relativistic effects, the quantum electrodynamic (QED) corrections were extracted. Moreover, comparing with an estimate of the one-electron QED contributions by using the GRASP2018 package, we found that the present experimental accuracy is high enough for testing the higher-order QED corrections for such a complex system with four electrons in the p subshell.

physics.atom-ph

An atrium segmentation network with location guidance and siamese adjustment

The segmentation of atrial scan images is of great significance for the three-dimensional reconstruction of the atrium and the surgical positioning. Most of the existing segmentation networks adopt a 2D structure and only take original images as input, ignoring the context information of 3D images and the role of prior information. In this paper, we propose an atrium segmentation network LGSANet with location guidance and siamese adjustment, which takes adjacent three slices of images as input and adopts an end-to-end approach to achieve coarse-to-fine atrial segmentation. The location guidance(LG) block uses the prior information of the localization map to guide the encoding features of the fine segmentation stage, and the siamese adjustment(SA) block uses the context information to adjust the segmentation edges. On the atrium datasets of ACDC and ASC, sufficient experiments prove that our method can adapt to many classic 2D segmentation networks, so that it can obtain significant performance improvements.

eess.IV

Automatic segmentation of meniscus based on MAE self-supervision and point-line weak supervision paradigm

Medical image segmentation based on deep learning is often faced with the problems of insufficient datasets and long time-consuming labeling. In this paper, we introduce the self-supervised method MAE(Masked Autoencoders) into knee joint images to provide a good initial weight for the segmentation model and improve the adaptability of the model to small datasets. Secondly, we propose a weakly supervised paradigm for meniscus segmentation based on the combination of point and line to reduce the time of labeling. Based on the weak label ,we design a region growing algorithm to generate pseudo-label. Finally we train the segmentation network based on pseudo-labels with weight transfer from self-supervision. Sufficient experimental results show that our proposed method combining self-supervision and weak supervision can almost approach the performance of purely fully supervised models while greatly reducing the required labeling time and dataset size.

cs.CV

Highly charged Nd$^{9+}$ Ion: A potential candidate of $\upmu$Hz linewidth optical clocks for probing fundamental physics

An active optical clock based on highly charged Nd$^{9+}$ ion is proposed for the first time. The clock can offer ultra-narrow linewidth at the $\upmu$Hz-level which is more than two-order of magnitude below the currently recorded laser linewidth. Operating at 605(90) nm superradiation lasing transition between the $5p^2~4f$ ground state and one of the long-lived $5p~4f^2$ excited state, the proposed active clock is inherently immune against the cavity noise which provides high stability. The clock serves as a sensitive probe with high sensitivity to variation of the fine-structure constant with accuracies below 10$^{-19}$ level. Sophisticated relativistic many-body methods are employed to predict related atomic properties that have corroborated the above findings.

physics.atom-ph

A low-energy compact Shanghai-Wuhan electron beam ion trap for extraction of highly charged ions

A low-energy, compact and superconducting electron beam ion trap (the Shanghai-Wuhan EBIT or SW-EBIT) for extraction of highly charged ions is presented. The magnetic field in the central drift tube of the SW-EBIT is approximately 0.21 T produced by a pair of high-temperature superconducting coils. The electron-beam energy of the SW-EBIT is in the range of 30-4000 eV, and the maximum electron-beam current is up to 9 mA. Acting as a source of highly charged ions, the ion-beam optics for extraction is integrated, including an ion extractor and an einzel lens. A Wien filter is then used to measure the charge-state distribution of the extracted ions. In this work, the tungsten ions below the charge state of 15 have been produced, extracted, and analyzed. The charge-state distributions and spectra in the range of 530-580 nm of tungsten ions have been measured simultaneously with the electron-beam energy of 279 eV and 300 eV, which preliminarily indicates that the 549.9 nm line comes from $W^{14+}$.

physics.atom-ph

Investigation of HIV-1 Gag binding with RNAs and Lipids using Atomic Force Microscopy

Atomic Force Microscopy was utilized to study the morphology of Gag, ΨRNA, and their binding complexes with lipids in a solution environment with 0.1Å vertical and 1nm lateral resolution. TARpolyA RNA was used as a RNA control. The lipid used was phospha-tidylinositol-(4,5)-bisphosphate (PI(4,5)P2). The morphology of specific complexes Gag-ΨRNA, Gag-TARpolyA RNA, Gag-PI(4,5)P2 and PI(4,5)P2-ΨRNA-Gag were studied. They were imaged on either positively or negatively charged mica substrates depending on the net charges carried. Gag and its complexes consist of monomers, dimers and tetramers, which was confirmed by gel electrophoresis. The addition of specific ΨRNA to Gag is found to increase Gag multimerization. Non-specific TARpolyA RNA was found not to lead to an increase in Gag multimerization. The addition PI(4,5)P2 to Gag increases Gag multimerization, but to a lesser extent than ΨRNA. When both ΨRNA and PI(4,5)P2 are present Gag undergoes comformational changes and an even higher degree of multimerization.

q-bio.BM

Modeling Multi-wavelength Pulse Profiles of Millisecond Pulsar PSR B1821-24

PSR B1821$-$24 is a solitary millisecond pulsar (MSP) which radiates multi-wavelength pulsed photons. It has complex radio, X-ray and $γ$-ray pulse profiles with distinct peak phase-separations that challenge the traditional caustic emission models. Using the single-pole annular gap model with suitable magnetic inclination angle ($α=40^\circ$) and viewing angle ($ζ=75^\circ$), we managed to reproduce its pulse profiles of three wavebands. It is found that the middle radio peak is originated from the core gap region at high altitudes, and the other two radio peaks are originated from the annular gap region at relatively low altitudes. Two peaks of both X-ray and $γ$-ray wavebands are fundamentally originated from annular gap region, while the $γ$-ray emission generated from the core gap region contributes somewhat to the first $γ$-ray peak. Precisely reproducing the multi-wavelength pulse profiles of PSR B1821$-$24 enables us to understand emission regions of distinct wavebands and justify pulsar emission models.

astro-ph.SR