SearcharxivSearch

arXiv subjects

Weining Zhang

Publications and source records attributed to Weining Zhang.

10 recordsLinked to original sources

UpgradeBench: A Decision-Centric Benchmark for Upgrading Fine-Tuned LLM Specialists

Organizations maintain task-specific adapters for open-weight language models, and each new base-model release forces a migration decision: retain existing specialists, port adapters, refresh from preserved behavior, or retrain. Prior transfer work evaluates isolated model pairs, without studying these choices across real model release sequences. We present UpgradeBench, a decision-driven longitudinal benchmark covering four consecutive Qwen releases, one continuation checkpoint, six tasks, and two model sizes, augmented by OLMo checkpoints with known training lineage. The benchmark disentangles three core questions: whether a new checkpoint improves fixed-recipe retrained specialist performance, whether specialization assets transfer across versions, and what recovery resources are usable. We observe upgrade gains differ across task-scale-release episodes: some retrained baselines improve while others stay within training noise, with durability ranging from under one release interval for text-to-SQL to over fourteen months for intent classification. Direct adapter copying depends neither on architecture nor model family: on OLMo, retention drops from 0.88-0.99 at 46B-token continued pretraining to zero at 2.9T tokens; annealing and model souping introduce no extra harm, with portability decaying with continued-pretraining distance. Given preserved input data, teacher relabeling recovers target-base specialists without fresh gold annotations, though compute savings are not guaranteed. Simulating a fixed decision policy over 33 upgrade episodes yields 0.37pp mean quality regret with zero behavioral regressions at one-third the compute and label cost of full retraining. A lightweight CKA probe over 256 prompts predicts cross-version adapter portability (Spearman 0.74 across eight model pairs). We release per-example predictions, cost logs, split manifests, and evaluation code.

cs.AI

No Judgment Without a Reason: Counterfactual Receipts for Versioned AI Evaluators

Evaluators often produce correct labels via flawed reasoning, a critical failure for agentic systems gating actions, routing reviews, or supplying training feedback. Standard evaluation only verifies final label correctness, ignoring whether judgment changes stem from valid evidence, consistent rules, or proper rule applicability. We formalize evaluator reasoning accountability via three core sources: grounds, norms, and authority. Varying these sources yields an eight-cell counterfactual judgment cube to characterize judgment updates. We define judgment receipts as minimal source replacement sets that reproduce revised verdicts to explain judgment transitions. We derive certification cost bounds for black-box evaluators and present ReasonBench, a policy and logical reasoning benchmark with verifiable receipts covering 19,520 cases and 7,200 controls. In frozen evaluations, Qwen3-1.7B reaches 98.41% receipt accuracy, while cube prediction scores 96.99%, a consistent 1.42-point drop validated by Qwen3-0.6B replication. Strong standard accuracy masks severe robustness flaws. Meaning-preserving source permutations reduce valid receipt recovery to 54.8% and 49.2% for direct and cube prediction. Models trained on simple single-source changes retain 93.75% verdict accuracy but recover only 7.16% of receipts for complex multi-source updates. Permutation retraining boosts consistency to 96.6% yet worsens cube prediction deficits. Structured counterfactual supervision fails to guarantee robust reasoning. We show reason-aware evaluation must decouple prediction and certification, reporting transformation consistency alongside standard accuracy for trustworthy evaluator auditing.

cs.AI

Iterate or Widen? When Test-Time Refinement Helps LiDAR Scene Completion: A Controlled Study of Evidence Geometry, Training Coverage, and Compute

Should a completion model spend extra test-time compute by iterating, or spend a similar parameter budget on a wider one-shot predictor? The answer is easily confounded by denoising curricula, corruption augmentation, capacity, and unpaired evaluation. We study this question in LiDAR semantic scene completion by comparing a one-shot predictor, a parameter-matched wider predictor, and a weight-tied multigrid refiner initialized from the same frozen predictor. The protocol separates coherent region removal, independent thinning, range-dependent attenuation, and additive clutter while preserving exact scene-condition pairing. Across five training seeds and 815 SemanticKITTI sequence-08 frames, the full iterative system improves mIoU over the wide control by 0.911 points under contiguous angular removal, with a 95% moving-block bootstrap interval of [0.804, 1.040] that clears a predeclared 0.5-point practical margin. Under independent 75% thinning, iteration adds only 0.300 points [0.166, 0.436], whereas observation-family augmentation adds 5.975 points [5.662, 6.140]. Neither intervention repairs additive clutter. The iterative system also costs 10.74 ms and 0.75 GiB per frame, versus 6.25 ms and 0.23 GiB for the wide control. These results establish a geometry-conditioned empirical boundary rather than a universal advantage: coherent gaps can justify fixed-depth refinement, broadly thinned evidence is addressed more effectively by training coverage, and spurious evidence requires a different robustness mechanism.

cs.CV

Share the Judge, Learn the Deferral: Where Specialization Helps LLM Evaluation

Agentic systems generate outputs faster than human review. We contrast two LLM evaluator specialization strategies: specialized judge weights, or rule-based deferral policies for safe judgment acceptance. On 99,952 rubric-conditioned samples, correct rubrics improve accuracy by 2.11 points, while incorrect rubrics reduce performance by 2.66 points. Splitting training data across eight criterion-specific LoRA experts lowers accuracy by 10.05 points and reduces 5% error-bound coverage from 24.44% to 5.43%. This loss is independent of model size and training settings, with most performance recoverable by warm-starting experts from a unified judge. Sweeping data budgets confirms scratch expert specialization yields no empirical gains. Warm-started splits appear competitive with unified models, yet under limited data, unified training outperforms split training, with specialization beneficial only after unified training plateaus. Results hold on HealthBench, where physician rubrics improve accuracy while flawed rubrics degrade performance. Unlike weight specialization, deferral policies enable efficient evaluation. On RewardBench 2, lightweight deferral heads form a 0.6B-4B-8B reward cascade with no core scoring modification. Across 20 splits, the cascade achieves 89.40% accuracy versus 84.75% for a standalone 8B judge at 41.5% compute, satisfying 95% risk constraints. Margin-based deferral matches accuracy at far higher compute cost. The design generalizes across models, improving Tulu-3-8B and Skywork-8B performance with a lightweight DeBERTa frontend. We derive simple, robust evaluator design rules: unify judgment training or warm-start split models, and use audited deferral cascades for low-cost, reliable LLM evaluation.

cs.AI

The Boundary Effect of QGP Droplet and Self-similarity Effect of Hadrons on QGP-hadron Phase Transition

We investigate the boundary effect of QGP droplet and self-similarity effect of hadrons on QGP-hadron phase transition. In intermediate or low energy collisions, when the transverse momentum is below QCD scale, QGP cannot be produced. However, if the transverse momentum fluctuates to a relatively large value, small scale QGP droplet is produced. The modified MIT bag model with multiple reflection expansion method is employed to study the QGP droplet with the curved boundary effect. It is found that the energy density, entropy density and pressure of QGP with the influence are smaller than those without the influence. In hadron phase, we propose Two-Body Fractal Model (TBFM) to study the self-similarity structure, arising from the resonance, quantum correlation and interaction effects. It is observed that energy density, entropy density and pressure increase due to the self-similarity structure. We calculate the transverse momentum spectra of pions with the self-similarity structure influence, showing a good agreement with the experimental data. Considering both the boundary effect and self-similarity structure influence, our model predicts an increase in the transition temperature compared to scenarios without these two effects in HIAF energy region $2.2\sim 4.5 \,\text{GeV}$.

hep-ph

Reinforcement Learning with Generalizable Gaussian Splatting

An excellent representation is crucial for reinforcement learning (RL) performance, especially in vision-based reinforcement learning tasks. The quality of the environment representation directly influences the achievement of the learning task. Previous vision-based RL typically uses explicit or implicit ways to represent environments, such as images, points, voxels, and neural radiance fields. However, these representations contain several drawbacks. They cannot either describe complex local geometries or generalize well to unseen scenes, or require precise foreground masks. Moreover, these implicit neural representations are akin to a ``black box", significantly hindering interpretability. 3D Gaussian Splatting (3DGS), with its explicit scene representation and differentiable rendering nature, is considered a revolutionary change for reconstruction and representation methods. In this paper, we propose a novel Generalizable Gaussian Splatting framework to be the representation of RL tasks, called GSRL. Through validation in the RoboMimic environment, our method achieves better results than other baselines in multiple tasks, improving the performance by 10%, 44%, and 15% compared with baselines on the hardest task. This work is the first attempt to leverage generalizable 3DGS as a representation for RL.

cs.CV

Whole-body Humanoid Robot Locomotion with Human Reference

Recently, humanoid robots have made significant advances in their ability to perform challenging tasks due to the deployment of Reinforcement Learning (RL), however, the inherent complexity of humanoid robots, including the difficulty of designing complicated reward functions and training entire sophisticated systems, still poses a notable challenge. To conquer these challenges, after many iterations and in-depth investigations, we have meticulously developed a full-size humanoid robot, "Adam", whose innovative structural design greatly improves the efficiency and effectiveness of the imitation learning process. In addition, we have developed a novel imitation learning framework based on an adversarial motion prior, which applies not only to Adam but also to humanoid robots in general. Using the framework, Adam can exhibit unprecedented human-like characteristics in locomotion tasks. Our experimental results demonstrate that the proposed framework enables Adam to achieve human-comparable performance in complex locomotion tasks, marking the first time that human locomotion data has been used for imitation learning in a full-size humanoid robot.

cs.RO

Analysis of the image of pion-emitting sources in source center of mass frame

In this paper, we try a method to extract the image of pion-emitting source function in the center-of-mass frame of source (CMFS). We choose the identical pion pairs according to the difference of their energy and use these pion pairs to build the correlation function. The purpose is to reduce the effect of $\triangle E \triangle t$, thus the corresponding imaging result can tend to the real source function. We examine the effect of this method by comparing its results with real source functions extracted from models directly.

nucl-th

Thermodynamics of Multiple Two-body Systems with Long-range Correlation

We aim to study thermodynamics of multiple two-body systems with long-range correlation using non-extensive statistics. Long-range correlation will cause multiple systems in anomalous diffusion. We consider the influence of long-range correlation as a background noise effect on a two-body system. We solve probability and entropy equations of a two-body system to obtain the temperature and distance dependence of the non-extensive parameter. The result shows the long-range correlation changes the system's entropy and energy. The more strongly is the system bounded, the less its energy is affected by the long-range correlation. Moreover, the anomalous diffusion approaches Brown motion with increasing temperature. This will help to understand how nonlinear field affects thermodynamics of a system.

cond-mat.stat-mech

Resonance Conversion as a Catalyser of Nuclear Reactions

It is shown that resonance interal conversion offers a feasible tool for mastering nuclear processes with laser or synchrotron radiation. Physics of the process is discussed in detail in historical aspect. Possible way of experimental applicaytion is shown in the case of the $M1$ 70.6-keV transition in nuclei of $^{169}$Yb. Nuclear transition rate in hydrogenlike ions of this nuclide can be enhanced by up to four orders of magnitude.

nucl-th