SearcharxivSearch

arXiv subjects

Xilu Wang

Publications and source records attributed to Xilu Wang.

At least 19 recordsLinked to original sources

Grounded Checklist Partial Credit for Agent Skill Trajectories

Language-model agents increasingly tackle long-horizon tasks in interactive environments, yet their evaluation commonly relies on task-level success rates by reducing an entire execution trajectory to whether the task passes an official verifier. This binary score hides partial progress and is particularly limited for procedural agent skill evaluations, since a skill can alter execution without changing the final outcome. While checklists provide finer-grained evaluation by scoring individual task requirements, costly manual authoring and unreliable automatic generation make trustworthy evaluation difficult to scale. To address these challenges, we introduce Grounded Checklist Partial Credit (GCPC), a human-governed and LLM-instantiated partial-credit evaluation of agent trajectories. Humans define reusable rules once, from which an LLM instantiates a task-specific checklist grounded in the task instruction and official verifier. To keep judgment tied to evidence, a judge scores each item from execution log evidence alone and abstains when evidence is missing. A separate scripted step then applies the official verifier outcome to the score. Across a 4,455-trajectory, deduplicated SkillsBench evaluation population, GCPC better discriminates official PASS and FAIL outcomes than holistic judging on the shared subset (AUC 0.689 vs. 0.619). Human evaluation on 96 trajectories from 12 tasks shows that GCPC aligns more closely with human assessments of progress. Applied to 1,946 matched with/without-skill pairs, GCPC exposes the effects hidden by pass@1: among 879 pairs whose binary outcome does not change, 20.9% improve by more than 0.10 while 18.7% regress by the same margin. The GCPC pipeline also transfers to Terminal-Bench and SWE-bench, demonstrating applicability beyond skill-conditioned evaluation.

cs.SE

From Relevance to Execution Utility: Reward-Aware Dynamic Execution Gating for Skill-Based LLM Agents

Agent skills are increasingly used to equip large language model (LLM) agents with reusable procedural knowledge. Although recent work has substantially improved skill retrieval due to the increasing skill libraries, retrieving a plausible skill bundle does not guarantee that executing it is worthwhile. Since every skill-conditioned rollout is computationally expensive, deciding whether a retrieved bundle should be executed has become an increasingly important challenge. To this end, we introduce the Reward-Aware Dynamic Execution Gate (RADEG), a lightweight, retriever-agnostic decision layer between skill retrieval and agent execution. RADEG learns a low-cost surrogate model that predicts the execution utility of a query--bundle pair before the expensive rollout is launched. To obtain informative supervision while controlling for task difficulty, we locally perturb each retrieved bundle by deleting, adding, or replacing one skill, producing matched same-query rollouts that isolate the effect of bundle composition on verifier reward. During deployment, RADEG updates only a warm-started logistic head as new verifier feedback becomes available, enabling inexpensive adaptation of the execute/skip boundary without retraining either the retriever or the agent. Under a query-level held-out evaluation on 288 collected rollouts, RADEG substantially reduces unnecessary agent executions while preserving a large fraction of the downstream verifier reward. It consistently outperforms relevance-based and random gating across different execution budgets, demonstrating that execution-aware surrogate modeling provides a practical and cost-effective complement to skill retrieval.

cs.AI

Not All Visual Tokens Are Equally Safe to Remove:Consequence-Sensitive Visual Token Compression

Visual token compression for vision--language models (VLMs) has largely relied on criteria such as attention, redundancy, and uncertainty to maximize average accuracy under a fixed compute budget, implicitly assuming that all errors carry equal cost. However, the consequence of an incorrect prediction on downstream tasks is rarely symmetric: misreading an invoice amount can be far more costly than misclassifying a background color. Motivated by this, we introduce consequence-sensitive visual token compression, which allocates visual computation across requests according to their potential error costs. Our method follows a calibrate-then-allocate procedure, estimating consequence-specific error-budget curves offline and applying the calibrated token budgets online using consequence signals available from question or task information. On a controlled within-task benchmark, high- and low-consequence questions are drawn from the same document images, so content alone cannot reveal which questions are costly to get wrong. In this setting, our method reduces high-stakes errors from 0.300 to 0.133 under the same total token budget, whereas a content-driven allocator performs no better than uniform allocation. Measuring how error rates change with token budget across different cost ratios, we derive an allocation frontier: uniform allocation is optimal when errors are equally costly, and token transfer toward high-consequence questions becomes increasingly beneficial as the cost gap grows. This allocation principle generalizes well across three dense vision-language benchmarks, two budget realization mechanisms (token deletion and resolution reallocation), two VLM architectures, and multiple token selection strategies. On a realistic mixed workload, consequence-sensitive allocation reduces cost-weighted error by 38% while achieving approximately 21% lower latency than full-resolution inference.

cs.CV

CRIP: Channel Level Representation Injection for Personalized One-Shot Federated Learning

One-shot federated learning (OSFL) has emerged as a promising collaborative model learning framework with only a single round of communication, offering significant advantages in communication efficiency and privacy preservation. However, OSFL often faces inherent limitations under severe domain heterogeneity across clients due to the lack of iterative knowledge exchange. Most existing OSFL methods require an auxiliary public dataset for knowledge distillation or leverage statistical information for parameter-level aggregation, overlooking feature shift caused by domain heterogeneity. To address these challenges, we propose CRIP, a personalized OSFL framework that operates in the representation space via channel-level feature alignment. To achieve this, each client uploads its feature extractor to the server, which broadcasts all extractors back to every client. Since not all source clients share compatible feature distributions with the target client, indiscriminate fusion of cross-client features would introduce domain-specific noise. Therefore, CRIP effectively measures the channel-wise representational similarity between the target client and each source client on a small local mini-batch, and selectively fuses only the most compatible features. Extensive experiments on domain-heterogeneous benchmarks such as DomainNet, PACS, and Office-Home demonstrate that CRIP consistently outperforms local models and state-of-the-art baselines, validating the effectiveness of representation-space personalization under extreme domain heterogeneity.

cs.LG

BudgetDraft: Acceptance-Aware Multi-View Training for Sparse-KV Speculative Decoding

Speculative decoding speeds up autoregressive decoding by using a drafter to propose multiple tokens that a verifier validates in parallel. In resource-constrained deployments, the drafter uses a sparse KV cache to limit peak GPU memory and end-to-end latency under a fixed KV budget, while the verifier keeps a full KV cache. Mid-to-long context inference (4K--16K context length) is common in real applications. However, naive sparse/full speculative decoding suffers from the sparse/full mismatch as context length grows, causing the acceptance rate to drop quickly. We propose BudgetDraft, a multi-view sparse training method for sparse drafting in mid-to-long inference. The drafter is exposed to multiple sampled KV budgets during training and learns to align each sparse view with one shared full-cache teacher target. BudgetDraft combines an acceptance-aware loss on a full-cache branch with a multi-view loss on a sparse-cache branch, producing a single budget-robust drafter that recovers acceptance across sparsity levels without extra inference-time components. Experimental results on PG-19, LongBench, and LWM show that BudgetDraft achieves up to 6.55x, 4.46x, and 2.10x end-to-end speedup vs AR at 4K, 8K, and 16K context lengths, while keeping the inference pipeline memory-friendly.

cs.LG

Identifying observable MeV lines from the decays of weak and main $r$-process isotopes in mergers

We consider predictions for the MeV gamma-ray spectrum emitted by the $β$ decays of freshly synthesized isotopes from a neutron star merger at timescales of relevance for post-merger (days) and remnant (years) emission. We develop a search algorithm to identify observable spectral peaks and then determine if a specific isotope has a dominant emission line producing the spectral feature. We predict emission spectra using nucleosynthesis calculations which consider nuclear models with distinct masses, $β$-decays, and fission properties as well as variations on main ($A>130$) and weak ($A<130$) $r$-process astrophysical conditions. We tabulate all lines from decaying isotopes that our procedure identifies and provide the predicted range in time over which each line could be visible. We find that Rh-106 presents a unique opportunity to distinguish between main and weak $r$-process emission, as our calculated spectrum above $\sim 1$ MeV for an event dominated by the weak $r$ process is identical to the Rh-106 emission spectrum from $\sim$ 0.2 to $\sim$17 years. We further find emission from species such as Hf-181, Ta-182, Ta-184, and Re-188 offers the potential to be able to distinguish between nuclear models. We investigate whether the 2.6 MeV strong gamma-ray line from Tl-208 is predicted to be robustly observable across calculation variations on both timescale of days and years. We find Tl-208 to consistently shine through on the order of years, though it can face competition from Ga-72 and La-140 at early times ($\sim$ days). We additionally highlight numerous isotopes of interest for observation and nuclear experiment.

astro-ph.HE

ELAS: Efficient Pre-Training of Low-Rank Large Language Models via 2:4 Activation Sparsity

Large Language Models (LLMs) have achieved remarkable capabilities, but their immense computational demands during training remain a critical bottleneck for widespread adoption. Low-rank training has received attention in recent years due to its ability to significantly reduce training memory usage. Meanwhile, applying 2:4 structured sparsity to weights and activations to leverage NVIDIA GPU support for 2:4 structured sparse format has become a promising direction. However, existing low-rank methods often leave activation matrices in full-rank, which dominates memory consumption and limits throughput during large-batch training. Furthermore, directly applying sparsity to weights often leads to non-negligible performance degradation. To achieve efficient pre-training of LLMs, this paper proposes ELAS: Efficient pre-training of Low-rank LLMs via 2:4 Activation Sparsity, a novel framework for low-rank models via 2:4 activation sparsity. ELAS applies squared ReLU activation functions to the feed-forward networks in low-rank models and implements 2:4 structured sparsity on the activations after the squared ReLU operation. We evaluated ELAS through pre-training experiments on LLaMA models ranging from 60M to 1B parameters. The results demonstrate that ELAS maintains performance with minimal degradation after applying 2:4 activation sparsity, while achieving training and inference acceleration. Moreover, ELAS reduces activation memory overhead, particularly with large batch sizes. Code is available at ELAS Repo.

cs.LG

Gardening on the Moon: An Advection-Diffusion Model to Guide the Search for Supernova Debris in the Lunar Regolith

The vertical redistribution of materials in the lunar regolith - ranging from continuously produced space-weathering products to sporadic pulses of supernova- or kilonova-derived isotopes - remains a fundamental problem in planetary science. We present a unified stochastic model of regolith gardening induced by the impact flux. Treating gardening as a competition between impact-driven advection and diffusion predicts the maturity profiles of Apollo cores over more than two orders of magnitude in time ($1.4 \times 10^7$ to $4.5 \times 10^8$ years). This model describes well the depth profiles of live Fe60 in Apollo regolith samples, suggesting that supernova dust capture is independent of native iron abundance, and is consistent with a uniform influx at the latitudes of the Apollo landing sites. We extend our model to predict lunar signals for live r-process species that might originate from supernovae or kilonovae: Pu244 tied to terrestrial detections, and I129, Hf182, and Cm247 based on r-process calculations. The Pu244/Fe60 depth profile can probe the origin of Pu244, motivating searches in Artemis regolith samples down to depths O(100) cm.

astro-ph.EP

High-energy Neutrino Predictions for T Coronae Borealis: Probing Particle Acceleration in Novae

The MAGIC detection of near-TeV gamma rays from the 2021 RS Oph ($2.45$ kpc) outburst has established recurrent novae as TeV particle accelerators. However, the origin of this emission (hadronic vs leptonic) remains unclear due to the lack of coincident neutrinos detected by IceCube. The upcoming outburst of the much closer T Coronae Borealis (T CrB, $\sim0.887$ kpc) offers a unique opportunity to detect these rare nova neutrinos. Here we present the first comparative analysis of the hadronic secondary fluxes expected from the upcoming T CrB outburst and evaluate their detectability across major observatories, considering two proton-acceleration mechanisms: (i) an external shock (ES) at $\sim10^{13}$ cm, and (ii) magnetic reconnection (MR), near the white dwarf surface at $\sim10^{9}$ cm. While the benchmark ES model predicts a gamma-ray flux detectable by current facilities, its corresponding neutrino flux largely remains undetectable. In contrast, the MR scenario generates a robust neutrino flux within the reach of IceCube and KM3NeT. Importantly, as the MR-produced gamma-rays are absorbed, the escaping MR neutrinos will arrive hours before any ES-origin signals. This distinct temporal separation can create a powerful phenomenological signature to disentangle the nova acceleration physics.

astro-ph.HE

Self-Supervised Federated Learning under Data Heterogeneity for Label-Scarce Diatom Classification

Label-scarce visual classification under decentralized and heterogeneous data is a fundamental challenge in pattern recognition, especially when sites exhibit partially overlapping class sets. While self-supervised federated learning (SSFL) offers a promising solution, existing studies commonly assume the same data heterogeneity pattern throughout pre-training and fine-tuning. Moreover, current partitioning schemes often fail to generate pure partially class-disjoint data settings, limiting controllable simulation of real-world label-space heterogeneity. In this work, we introduce SSFL for diatom classification as a representative real-world instance and systematically investigate stage-specific data heterogeneity. We study cross-site variation in unlabeled data volume during pre-training and label-space misalignment during downstream fine-tuning. To study the latter in a controllable setting, we propose PreDi, a partitioning scheme that disentangles label-space heterogeneity into two orthogonal dimensions, namely class Prevalence and class-set size Disparity, enabling separate analysis of their effects. Guided by the resulting insights, we further propose PreP-WFL (Prevalence-based Personalized Weighted Federated Learning) to adaptively strengthen rare-class representations in low-prevalence scenarios. Extensive experiments show that SSFL consistently outperforms local-only training under both homogeneous and heterogeneous settings. The pronounced heterogeneity in unlabeled data volume is associated with improved representation pre-training, whereas under label-space heterogeneity, prevalence dominates performance and disparity has a smaller effect. PreP-WFL effectively mitigates this degradation, with gains increasing as prevalence decreases. These findings provide a mechanistic basis for characterizing label-space heterogeneity in decentralized recognition systems.

cs.CV

TSegAgent: Zero-Shot Tooth Segmentation via Geometry-Aware Vision-Language Agents

Automatic tooth segmentation and identification from intra-oral scanned 3D models are fundamental problems in digital dentistry, yet most existing approaches rely on task-specific 3D neural networks trained with densely annotated datasets, resulting in high annotation cost and limited generalization to scans from unseen sources. Thus, we propose TSegAgent, which addresses these challenges by reformulating dental analysis as a zero-shot geometric reasoning problem rather than a purely data-driven recognition task. The key idea is to combine the representational capacity of general-purpose foundation models with explicit geometric inductive biases derived from dental anatomy. Instead of learning dental-specific features, the proposed framework leverages multi-view visual abstraction and geometry-grounded reasoning to infer tooth instances and identities without task-specific training. By explicitly encoding structural constraints such as dental arch organization and volumetric relationships, the method reduces uncertainty in ambiguous cases and mitigates overfitting to particular shape distributions. Experimental results demonstrate that this reasoning-oriented formulation enables accurate and reliable tooth segmentation and identification with low computational and annotation cost, while exhibiting strong generalization across diverse and previously unseen dental scans.

cs.CV

Long Chain-of-Thought Compression via Fine-Grained Group Policy Optimization

Large Language Models (LLMs) often generate unnecessarily verbose Chain-of-Thought (CoT) reasoning that increases computational costs and latency without proportional performance gains. In this paper, we propose Fine-grained Group policy Optimization (FGO), a Reinforcement Learning (RL) algorithm that refines group responses by subdividing them and assigning appropriate weights based on length and entropy, thereby enabling effective CoT compression. Meanwhile, as an enhanced variant of Group Relative Policy Optimization (GRPO), FGO successfully addresses two major limitations of the GRPO: inefficient data utilization and entropy collapse. We evaluate FGO on multiple reasoning LLMs and benchmarks, including MATH500, AIME24, AMC23, and Minerva. Experimental results show that FGO achieves efficient CoT compression without degrading performance, and simultaneously resolves the key limitations of GRPO. Code: https://github.com/Mr-XcHan/FGO.

cs.LG

Proton-rich production of lanthanides: the $νi$ process

The astrophysical origin of the lanthanides is an open question in nuclear astrophysics. Besides the widely studied $s$, $i$, and $r$ processes in moderately-to-strongly neutron-rich environments, an intriguing alternative site for lanthanide production could in fact be robustly $\textit{proton-rich}$ matter outflows from core-collapse supernovae under specific conditions -- in particular, high-entropy winds with enhanced neutrino luminosity and fast dynamical timescales. In this environment, excess protons present after charged particle reactions have ceased can continue to be converted to neutrons by (anti-)neutrino interactions, producing a neutron capture reaction flow up to A~200. This scenario, christened the $νi$ process in a recent paper, has previously been discussed as a possibility. Here, we examine the prospects for $νi$ process through the lens of stellar abundance patterns, bolometric lightcurves, and galactic chemical evolution models, with a particular focus on hypernovae as candidate sites. We identify specific lanthanide signatures for which the $νi$ process can provide a credible alternative to $r$/$i$ processes.

astro-ph.HE

Observatory Science with eXTP

Scheduled for launch in 2030, the enhanced X-ray Timing and Polarization (eXTP) telescope is a Chinese space-based mission aimed at studying extreme conditions and phenomena in astrophysics. eXTP will feature three main payloads: Spectroscopy Focusing Arrays (SFAs), Polarimetry Focusing Arrays (PFAs), and a Wide-field Camera (W2C). This white paper outlines observatory science, incorporating key scientific advances and instrumental changes since the publication of the previous white paper [1]. We will discuss perspectives of eXTP on the research domains of flare stars, supernova remnants, pulsar wind nebulae, cataclysmic variables, X-ray binaries, ultraluminous X-ray sources, AGN, and pulsar-based positioning and timekeeping.

astro-ph.IM

OWLed: Outlier-weighed Layerwise Pruning for Efficient Autonomous Driving Framework

The integration of Large Language Models (LLMs) into autonomous driving systems offers promising enhancements in environmental understanding and decision-making. However, the substantial computational demands of deploying LLMs locally on vehicles render this approach unfeasible for real-world automotive applications. To address this challenge, we introduce OWLed, the Outlier-Weighed Layerwise Pruning for Efficient Autonomous Driving Framework that leverages outlier-weighted layerwise sparsity for model compression. Our method assigns non-uniform sparsity ratios to different layers based on the distribution of outlier features, significantly reducing the model size without the need for fine-tuning. To ensure the compressed model adapts well to autonomous driving tasks, we incorporate driving environment data into both the calibration and pruning processes. Our empirical studies reveal that the encoder component is more sensitive to pruning than the LLM, highlighting its critical role in the system. Experimental results demonstrate that OWLed outperforms existing methods in perception, action prediction, and language understanding while substantially lowering computational requirements. These findings underscore the potential of combining advanced pruning techniques with LLMs to develop efficient and robust autonomous driving systems capable of handling complex scenarios. Code will be made publicly available.

cs.LG

LOST: Low-rank and Sparse Pre-training for Large Language Models

While large language models (LLMs) have achieved remarkable performance across a wide range of tasks, their massive scale incurs prohibitive computational and memory costs for pre-training from scratch. Recent studies have investigated the use of low-rank parameterization as a means of reducing model size and training cost. In this context, sparsity is often employed as a complementary technique to recover important information lost in low-rank compression by capturing salient features in the residual space. However, existing approaches typically combine low-rank and sparse components in a simplistic or ad hoc manner, often resulting in undesirable performance degradation compared to full-rank training. In this paper, we propose \textbf{LO}w-rank and \textbf{S}parse pre-\textbf{T}raining (\textbf{LOST}) for LLMs, a novel method that ingeniously integrates low-rank and sparse structures to enable effective training of LLMs from scratch under strict efficiency constraints. LOST applies singular value decomposition to weight matrices, preserving the dominant low-rank components, while allocating the remaining singular values to construct channel-wise sparse components to complement the expressiveness of low-rank training. We evaluate LOST on LLM pretraining ranging from 60M to 7B parameters. Our experiments show that LOST achieves competitive or superior performance compared to full-rank models, while significantly reducing both memory and compute overhead. Moreover, Code is available at \href{https://github.com/JiaxiLi1/LOST-Low-rank-and-Sparse-Training-for-Large-Language-Models}{LOST Repo}

cs.LG

CLIP Brings Better Features to Visual Aesthetics Learners

Image Aesthetics Assessment (IAA) is a challenging task due to its subjective nature and expensive manual annotations. Recent large-scale vision-language models, such as Contrastive Language-Image Pre-training (CLIP), have shown their promising representation capability for various downstream tasks. However, the application of CLIP to resource-constrained and low-data IAA tasks remains limited. While few attempts to leverage CLIP in IAA have mainly focused on carefully designed prompts, we extend beyond this by allowing models from different domains and with different model sizes to acquire knowledge from CLIP. To achieve this, we propose a unified and flexible two-phase CLIP-based Semi-supervised Knowledge Distillation (CSKD) paradigm, aiming to learn a lightweight IAA model while leveraging CLIP's strong generalization capability. Specifically, CSKD employs a feature alignment strategy to facilitate the distillation of heterogeneous CLIP teacher and IAA student models, effectively transferring valuable features from pre-trained visual representations to two lightweight IAA models, respectively. To efficiently adapt to downstream IAA tasks in a low-data regime, the two strong visual aesthetics learners then conduct distillation with unlabeled examples for refining and transferring the task-specific knowledge collaboratively. Extensive experiments demonstrate that the proposed CSKD achieves state-of-the-art performance on multiple widely used IAA benchmarks. Furthermore, analysis of attention distance and entropy before and after feature alignment shows the effective transfer of CLIP's feature representation to IAA models, which not only provides valuable guidance for the model initialization of IAA but also enhances the aesthetic feature representation of IAA models. Code will be made publicly available.

cs.CV

Gamma rays as a signature of r-process producing supernovae: remnants and future Galactic explosions

We consider the question of whether core-collapse supernovae (CCSNe) can produce rapid neutron capture process (r-process) elements and how future MeV gamma-ray observations could address this. Rare types of CCSNe characterized by substantial magnetic fields and rotation, known as magnetorotational supernovae (MR-SNe), are theoretically predicted to produce these elements, although direct observational evidence is lacking. We suggest that this critical question be addressed through the study of some of the eleven CCSN remnants located within 10 kpc, as well as through the detection of gamma-ray emission from a future Galactic supernova. We use a two-dimensional MR-SN model to estimate the expected gamma flux stemming from nuclear decays in the range of a few tens of keV to a few MeV. Our results indicate that an observation of Sn-126 (Sb-126) in a remnant stands out as a signature of an r-process-producing supernova. Since the neutron-rich conditions that lead to the production of the r-process could also enhance the production of Fe-60, the detection of substantial Fe-60 (Co-60) would be indicative of favorable conditions for the r-process. In the case of a future supernova explosion, when the evolution of the spectrum is studied over ten days to a few years, a rich picture emerges. At various epochs, second peak r-process isotopes such as Sb-125, I-131, Te-132, I-132 and La-140 produce gamma-ray signals that emerge above the background from explosive burning products and electron-positron annihilation. The weak r-process isotopes Nb-95, Ru-103, Rh-106 also have periods of prominence. While MR-SNe are predicted to have a relatively small main r-process contribution, third peak isotopes like Ir-194 could still be above next-generation MeV gamma instrument sensitivities.

astro-ph.HE