SearcharxivSearch

arXiv subjects

Xinyang Liu

Publications and source records attributed to Xinyang Liu.

At least 19 recordsLinked to original sources

Do Uncertainty Signals Help? A Systematic Study of Uncertainty-Aware Decoding with Rollback Mechanisms

Prediction uncertainty is a widely adopted metric for quantifying model confidence, with downstream applications spanning model explanation, data selection, and prediction rollback. Despite its demonstrated utility, the potential of uncertainty quantification to enhance code generation in large language models (LLMs) remains largely underexplored, raising a critical question: to what extent can uncertainty serve as an effective signal for improving LLM-based code generation? To answer this question, we study uncertainty-aware rollback decoding, an inference-time strategy that uses uncertainty signals to identify unreliable generation regions and roll back to earlier valid prefixes without retraining the model. We evaluate this framework on seven code LLMs, five code generation benchmarks, and eight token-level uncertainty signals under a unified decoding setup. Our results show that the complete rollback framework improves over equal-budget restart across the evaluated benchmarks and model settings, with gains of up to 0.26 in pass@1 and 0.35 in AvgTestPassRate on functional code generation benchmarks, and an absolute improvement of up to 6.4\% in Patch-Aligned Safe Rate on Dsec-Python. Among the evaluated signals, information-theoretic measures such as token entropy and negative log-likelihood show the most favorable overall trend, frequently achieving the best or near-best results on standard benchmarks. A component-controlled ablation further shows that feedback-guided rollback provides the main improvement, while uncertainty localization provides an additional gain when checking, budget, rollback, and branch decay are held fixed.

cs.LG

FFR: Forward-Forward Learning for Regression

The Forward-Forward (FF) algorithm offers a computationally efficient and biologically plausible alternative to backpropagation (BP) by training neural networks through purely local, layer-wise optimization. However, FF is inherently designed for classification via contrastive positive-negative sample pairs, and extending it to regression poses fundamental challenges: continuous target space lack natural "opposites" for contrastive learning, and the standard goodness function carries no information about target magnitude or ordering. We propose FFR (Forward-Forward for Regression), to our knowledge, the first framework to extend FF to real-world regression and demonstrate competitive performance across diverse real-world datasets. FFR introduces three key innovations: (1) an ordinal competitive goodness function that replaces contrastive pairs with competitive learning between partitioned neuron groups under distance-aware ordinal supervision; (2) a stratified ladder architecture where shallow layers learn coarse ordinal discrimination and deeper layers refine into fine-grained regression, with multi-scale feature aggregation for inter-layer collaboration; and (3) hierarchical prediction with uncertainty estimation, where multi-scale predictors jointly provide robust predictions and prediction confidence as a free-lunch. Extensive experimental results show FFR recovers on average 98.6% of BP's accuracy across five real-world regression benchmarks while reducing peak training memory to only 27% of BP's at depth 8 and 8% at depth 32, with per-iteration time around 72% of BP's, and substantially outperforms all BP-free competitors.

cs.LG

Noise-like pulse laser source with ultrabroadband tunability and coherence-limited sub-structure

High brightness and low coherence laser sources with wideband tunability are essential for many full-field imaging applications aiming for high contrast and speckle free performance. However, this combination of parameters is challenging to achieve. The current solutions focus on decreasing spatial coherence or generation of time-varying speckle patterns, while suppression of temporal coherence typically compromises brightness. Here we demonstrate a wideband pulsed laser source with low temporal coherence and the absence of phase correlation between pulses as an alternative approach with simultaneous time and frequency diversity. The full gain spectrum of a Tm doped fiber laser (1650 nm 2000 nm) is operated in a tunable noise like pulse regime, which by nature is composed of countless structured elementary events with uncorrelated phases randomly varying from bunch to bunch. The measured spectral widths range from 13.8 nm to 18.8 nm, while the average output power varies between 63.3 mW and 213 mW. Numerical simulations reveal that temporal coherence decreases significantly with increasing optical gain, dropping from near unity at low gain to approximately 0.2 at high gain. The startup dynamics of the noise like pulse laser are experimentally studied using the dispersive Fourier transformation (DFT) method. Based on single shot spectra and frequency resolved optical gating traces, the coherence properties of the laser are further analyzed by calculating the mutual coherence function and cross-spectral density. The noise like pulse laser exhibits a coherence time of approximately 100 fs and an average pulse burst duration of about 40 ps in the high-gain regime.

physics.optics

Damping dynamics of the centroid oscillation of a relativistic laser pulse in a plasma channel

The centroid oscillation of an offset laser pulse propagating in a preformed plasma channel is investigated through theoretical analysis and three-dimensional particle-in-cell simulations. For non-relativistic laser pulses, the mode leakage of a finite channel and the temporal walk-off between the fundamental and high order modes of a finite-duration laser induce a decay in the laser centroid oscillation. An analytical model characterizing these decay mechanisms is derived and validated by simulations. For relativistic laser pulses, the slice-based centroid oscillation frequency develops an axial chirp due to relativistic channel modification and photon deceleration. This chirp leads to phase mixing across different axial slices of the pulse, resulting in a rapid damping of the overall centroid oscillation. Understanding this oscillation damping is crucial for mitigating electron beam pointing jitter and maintaining beam quality in high-energy, channel-guided laser wakefield accelerators.

physics.acc-ph

AutoSOTA: An End-to-End Automated Research System for State-of-the-Art AI Model Discovery

Artificial intelligence research increasingly depends on prolonged cycles of reproduction, debugging, and iterative refinement to achieve State-Of-The-Art (SOTA) performance, creating a growing need for systems that can accelerate the full pipeline of empirical model optimization. In this work, we introduce AutoSOTA, an end-to-end automated research system that advances the latest SOTA models published in top-tier AI papers to reproducible and empirically improved new SOTA models. We formulate this problem through three tightly coupled stages: resource preparation and goal setting; experiment evaluation; and reflection and ideation. To tackle this problem, AutoSOTA adopts a multi-agent architecture with eight specialized agents that collaboratively ground papers to code and dependencies, initialize and repair execution environments, track long-horizon experiments, generate and schedule optimization ideas, and supervise validity to avoid spurious gains. We evaluate AutoSOTA on recent research papers collected from eight top-tier AI conferences under filters for code availability and execution cost. Across these papers, AutoSOTA achieves strong end-to-end performance in both automated replication and subsequent optimization. Specifically, it successfully discovers 105 new SOTA models that surpass the original reported methods, averaging approximately five hours per paper. Case studies spanning LLM, NLP, computer vision, time series, and optimization further show that the system can move beyond routine hyperparameter tuning to identify architectural innovation, algorithmic redesigns, and workflow-level improvements. These results suggest that end-to-end research automation can serve not only as a performance optimizer, but also as a new form of research infrastructure that reduces repetitive experimental burden and helps redirect human attention toward higher-level scientific creativity.

cs.CL

Giant Magnetocaloric Effect in a High-Spin Shastry-Sutherland Dipolar Magnet

The Shastry-Sutherland lattice is a prototypical frustrated quantum magnet. It is notable for its exactly solvable dimer-singlet ground state and hosts a wealth of magnetic phenomena under external fields. Here, this work investigates the high-spin (S = 7/2) Eu-based magnet Eu2MgSi2O7 (EMSO) using low-temperature magnetothermal measurements and Monte Carlo simulations, revealing a giant magnetocaloric effect (MCE) in this Shastry-Sutherland compound. The entropy change peak value is found to be 55.0 J kg-1 K-1 under a field change of B = 0-4 T, approximately 1.5 times larger than the commercial Gd3Ga5O12 (GGG). Adiabatic demagnetization refrigeration achieves a lowest temperature of 151 mK, deeply into the sub-Kelvin regime. Furthermore, a distinctive cooling effect persists below about 1 T, a characteristic absent for conventional magnetic coolants. A dipolar Shastry-Sutherland model is introduced as a minimal model to describe this system; in particular, the experimentally revealed 1/3 magnetization pseudo-plateau can be ascribed to the presence of dipolar couplings between Eu2+ ions, further stabilized by the thermal fluctuations, explaining the persistent cooling effect. This work establishes EMSO as a novel platform for exploring the dipolar Shastry-Sutherland system and for sub-Kelvin adiabatic demagnetization refrigeration.

cond-mat.mtrl-sci

Ising Supercriticality and Universal Magnetocalorics in Spiral Antiferromagnet Nd$_3$BWO$_9$

The celebrated analogy between the pressure-temperature phase diagram of a liquid-gas system and the field-temperature phase diagram of a ferromagnet has long been a cornerstone for understanding universality of phase transitions and critical phenomena. Here we extend this analogy to a highly frustrated antiferromagnet, the spiral Ising compound Nd$_3$BWO$_9$ with kagome layers. In its phase diagram, we identify a metamagnetic transition line with a critical endpoint (CEP) located at $\mu_0H_{\mathrm{c}} \simeq 1.04$ T and $T_{\mathrm{c}} \simeq 0.3$ K. Above the CEP, an Ising supercritical regime emerges with crossover lines that follow a universal scaling law, as evidenced by the specific heat, magnetic susceptibility, and magnetocaloric measurements. Remarkably, we observe highly sensitive magnetic cooling near the emergent CEP, characterized by a divergent magnetic Gr\"uneisen ratio $\Gamma_H \propto 1/t^{\beta+\gamma-1}$, with $\beta + \gamma \simeq 1.563$ the sum of critical exponents of the 3D Ising universality class and $t \equiv (T-T_{\rm c})/T_{\rm c}$ the reduced temperature. Adiabatic demagnetization from 2 K and 4 T reaches a minimum temperature of 195 mK, via a self-cascading process that combines supercritical and topological cooling. Our findings open a new avenue for studying supercritical phenomena and magnetic refrigeration with the frustrated rare-earth compounds RE$_3$BWO$_9$ and, more broadly, in Ising-anisotropic antiferromagnets such as spin ices.

cond-mat.str-el

GDEPO: Group Dual-dynamic and Equal-right Advantage Policy Optimization with Enhanced Training Data Utilization for Sample-Constrained Reinforcement Learning

Automated Theorem Proving (ATP) represents a fundamental challenge in Artificial Intelligence (AI), requiring the construction of machine-verifiable proofs in formal languages such as Lean to evaluate AI reasoning capabilities. Reinforcement learning (RL), particularly the high-performance Group Relative Policy Optimization (GRPO) algorithm, has emerged as a mainstream approach for this task. However, in ATP scenarios, GRPO faces two critical issues: when composite rewards are used, its relative advantage estimation may conflict with the binary feedback from the formal verifier; meanwhile, its static sampling strategy may discard entire batches of data if no valid proof is found, resulting in zero contribution to model updates and significant data waste. To address these limitations, we propose Group Dual-dynamic and Equal-right-advantage Policy Optimization (GDEPO), a method incorporating three core mechanisms: 1) dynamic additional sampling, which resamples invalid batches until a valid proof is discovered; 2) equal-right advantage, decoupling the sign of the advantage function (based on correctness) from its magnitude (modulated by auxiliary rewards) to ensure stable and correct policy updates; and 3) dynamic additional iterations, applying extra gradient steps to initially failed but eventually successful samples to accelerate learning on challenging cases. Experiments conducted on three datasets of varying difficulty (MinF2F-test, MathOlympiadBench, PutnamBench) confirm the effectiveness of GDEPO, while ablation studies validate the necessity of its synergistic components. The proposed method enhances data utilization and optimization efficiency, offering a novel training paradigm for ATP.

cs.AI

Canted ferromagnetic order in a distorted triangular-lattice magnet Na$_2$SrCo(VO$_4$)$_2$

Triangular-lattice cobaltates with glaserite-type $X_2Y$Co($T$O$_4)_2$ structure provide an ideal platform to investigate intriguing quantum magnetism. Here we report a comprehensive study of the structural and magnetic properties of a triangular-lattice cobalt vanadate $\rm Na_2SrCo(VO_4)_2$. Room-temperature x-ray and neutron powder diffraction confirm that $\rm Na_2SrCo(VO_4)_2$ crystallizes in the monoclinic $P2_1/c$ space group with slightly distorted triangular layers of $\rm Co^{2+}$ ions. Magnetization measurements reveal a ferromagnetic transition at $T\rm_C \approx 3.4~{\rm K}$, where a sharp $\lambda$-type anomaly is observed in the specific heat. The magnetic entropy recovered up to 55 K approaches 90$\%$ of $R{\rm ln}2$, supporting an effective spin-1/2 state of Co$^{2+}$ ions at low temperature. Neutron diffraction at 2.3 K (below $T_{\rm C}$) further confirms a long-range canted ferromagnetic order with the Co$^{2+}$ moments aligned in the $ac$ plane and the ordered moment size of $\sim$ 2.6 $\mu\rm_{B}$. Comparing with its sister compounds with a trigonal symmetry, $\rm Na_2BaCo(VO_4)_2$ with a collinear ferromagnetic structure and the recently discovered spin supersolid candidate $\rm Na_2BaCo(PO_4)_2$ with a distinct Y-like antiferromagnetic ground state, this study indicates the decisive role of the $T{\rm O_4}$ tetrahedra in tuning exchange interactions and contrasting magnetic behaviors of these glaserite-structure compounds.

cond-mat.str-el

OmniScientist: Toward a Co-evolving Ecosystem of Human and AI Scientists

With the rapid development of Large Language Models (LLMs), AI agents have demonstrated increasing proficiency in scientific tasks, ranging from hypothesis generation and experimental design to manuscript writing. Such agent systems are commonly referred to as "AI Scientists." However, existing AI Scientists predominantly formulate scientific discovery as a standalone search or optimization problem, overlooking the fact that scientific research is inherently a social and collaborative endeavor. Real-world science relies on a complex scientific infrastructure composed of collaborative mechanisms, contribution attribution, peer review, and structured scientific knowledge networks. Due to the lack of modeling for these critical dimensions, current systems struggle to establish a genuine research ecosystem or interact deeply with the human scientific community. To bridge this gap, we introduce OmniScientist, a framework that explicitly encodes the underlying mechanisms of human research into the AI scientific workflow. OmniScientist not only achieves end-to-end automation across data foundation, literature review, research ideation, experiment automation, scientific writing, and peer review, but also provides comprehensive infrastructural support by simulating the human scientific system, comprising: (1) a structured knowledge system built upon citation networks and conceptual correlations; (2) a collaborative research protocol (OSP), which enables seamless multi-agent collaboration and human researcher participation; and (3) an open evaluation platform (ScienceArena) based on blind pairwise user voting and Elo rankings. This infrastructure empowers agents to not only comprehend and leverage human knowledge systems but also to collaborate and co-evolve, fostering a sustainable and scalable innovation ecosystem.

cs.CY

Quantum fluctuations associated with first-order magnetic transition in a frustrated kagome lattice antiferromagnet

Intense quantum fluctuations arising from geometrical frustrations in kagome-lattice magnets provide a feasible approach to exotic quantum states. Here, we document an unexpected isosymmetric first-order magnetic transition in the recently synthesized frustrated kagome-lattice antiferromagnet Nd3ScBi5, which is characterized by significant latent heat and a pronounced magnetocaloric effect, as well as discontinuous Raman shifts and negligible hysteresis. Employing the magnetocaloric effect as a detection method, in conjunction with systematical field-dependent physical properties, we uncover a distinctive 1/2 magnetization plateau phase with significant quantum fluctuations. Our study unveils Nd3ScBi5 as a prototypical model with an emerging phase of enhanced quantum fluctuations triggered by first-order magnetic transitions.

cond-mat.str-el

Route Experts by Sequence, not by Token

Mixture-of-Experts (MoE) architectures scale large language models (LLMs) by activating only a subset of experts per token, but the standard TopK routing assigns the same fixed number of experts to all tokens, ignoring their varying complexity. Prior adaptive routing methods introduce additional modules and hyperparameters, often requiring costly retraining from scratch. We propose Sequence-level TopK (SeqTopK), a minimal modification that shifts the expert budget from the token level to the sequence level. By selecting the top $T \cdot K$ experts across all $T$ tokens, SeqTopK enables end-to-end learned dynamic allocation -- assigning more experts to difficult tokens and fewer to easy ones -- while preserving the same overall budget. SeqTopK requires only a few lines of code, adds less than 1% overhead, and remains fully compatible with pretrained MoE models. Experiments across math, coding, law, and writing show consistent improvements over TopK and prior parameter-free adaptive methods, with gains that become substantially larger under higher sparsity (up to 16.9%). These results highlight SeqTopK as a simple, efficient, and scalable routing strategy, particularly well-suited for the extreme sparsity regimes of next-generation LLMs. Code is available at https://github.com/Y-Research-SBU/SeqTopK.

cs.LG

Vision Transformer for Robust Occluded Person Reidentification in Complex Surveillance Scenes

Person re-identification (ReID) in surveillance is challenged by occlusion, viewpoint distortion, and poor image quality. Most existing methods rely on complex modules or perform well only on clear frontal images. We propose Sh-ViT (Shuffling Vision Transformer), a lightweight and robust model for occluded person ReID. Built on ViT-Base, Sh-ViT introduces three components: First, a Shuffle module in the final Transformer layer to break spatial correlations and enhance robustness to occlusion and blur; Second, scenario-adapted augmentation (geometric transforms, erasing, blur, and color adjustment) to simulate surveillance conditions; Third, DeiT-based knowledge distillation to improve learning with limited labels.To support real-world evaluation, we construct the MyTT dataset, containing over 10,000 pedestrians and 30,000+ images from base station inspections, with frequent equipment occlusion and camera variations. Experiments show that Sh-ViT achieves 83.2% Rank-1 and 80.1% mAP on MyTT, outperforming CNN and ViT baselines, and 94.6% Rank-1 and 87.5% mAP on Market1501, surpassing state-of-the-art methods.In summary, Sh-ViT improves robustness to occlusion and blur without external modules, offering a practical solution for surveillance-based personnel monitoring.

cs.CV

Metrics and evaluations for computational and sustainable AI efficiency

The rapid advancement of Artificial Intelligence (AI) has created unprecedented demands for computational power, yet methods for evaluating the performance, efficiency, and environmental impact of deployed models remain fragmented. Current approaches often fail to provide a holistic view, making it difficult to compare and optimise systems across heterogeneous hardware, software stacks, and numeric precisions. To address this gap, we propose a unified and reproducible methodology for AI model inference that integrates computational and environmental metrics under realistic serving conditions. Our framework provides a pragmatic, carbon-aware evaluation by systematically measuring latency and throughput distributions, energy consumption, and location-adjusted carbon emissions, all while maintaining matched accuracy constraints for valid comparisons. We apply this methodology to multi-precision models across diverse hardware platforms, from data-centre accelerators like the GH200 to consumer-level GPUs such as the RTX 4090, running on mainstream software stacks including PyTorch, TensorRT, and ONNX Runtime. By systematically categorising these factors, our work establishes a rigorous benchmarking framework that produces decision-ready Pareto frontiers, clarifying the trade-offs between accuracy, latency, energy, and carbon. The accompanying open-source code enables independent verification and facilitates adoption, empowering researchers and practitioners to make evidence-based decisions for sustainable AI deployment.

cs.PF

Ultra-Fast Language Generation via Discrete Diffusion Divergence Instruct

Fast and high-quality language generation is the holy grail that people pursue in the age of AI. In this work, we introduce Discrete Diffusion Divergence Instruct (DiDi-Instruct), a training-based method that initializes from a pre-trained diffusion large language model (dLLM) and distills a few-step student for fast generation. The model distilled with DiDi-Instruct matches or surpasses its dLLM teacher and the GPT-2 baseline while providing up to 64$\times$ acceleration. The theoretical foundation of DiDi-Instruct is a novel framework based on integral KL-divergence minimization, which leads to a practical training algorithm. We further introduce grouped reward normalization, intermediate-state matching, and the reward-guided ancestral sampler to improve training stability, model coverage, and inference quality. On the OpenWebText benchmark, DiDi-Instruct achieves perplexity ranging from 62.2 (8 NFEs) to 18.4 (128 NFEs), outperforming prior accelerated dLLMs and the GPT-2 baseline. These gains incur a negligible entropy loss (around $1$%) and reduce additional training wall-clock time by more than $20\times$ compared to competing dLLM distillation methods. We further validate the robustness and effectiveness of DiDi-Instruct through extensive ablation studies, model scaling, downstream task evaluations, and unconditional protein sequence generation. In conclusion, DiDi-Instruct enables efficient and effective distillation for language generation in the blink of an eye.

cs.CL

Realization of large magnetocaloric effect in the Kagome antiferromagnet Gd3BWO9 for Sub-Kelvin cryogenic refrigeration

Rare-earth (RE) based frustrated magnets have attracted great attention as excellent candidates for magnetic refrigeration at sub-Kelvin temperatures, while the experimental identification on systems exhibiting both large volumetric cooling capacity and reduced working temperatures far below 1 K remain to be a challenge. Here, through the ultra-low temperature magnetism and thermodynamic characterizations, we unveil the large magnetocaloric effect (MCE) realized at sub-Kelvin temperatures in the frustrated Kagome antiferromagnet Gd3BWO9 with TN~1.0 K. The isothermal magnetization curves indicate the existence of field (B) induced anisotropic magnetic phase diagrams, where four distinct magnetic phases for B // c-axis and five magnetic phases for B // ab-plane are identified at T< TN. The analysis of magnetic entropy S(B, T) data and direct adiabatic demagnetization tests reveal a remarkable cooling performance at sub-Kelvin temperatures featured by a large volumetric entropy density 502.2 mJ/K/cm3 and a low attainable minimal temperature Tmin~168 mK from the initial cooling condition of 2 K and 6 T, surpassing most of Gd-based refrigerants previously documented in temperature ranges of 0.25-4 K. The realized Tmin~168 mK far below TN ~ 1.0 K in Gd3BWO9 is related to the combined effects of magnetic frustration and criticality-enhanced MCE, which together leave a substantial magnetic entropy at reduced temperatures by enhancing spin fluctuations.

cond-mat.str-el

Antiferromagnetic ordering and critical behavior induced giant magnetocaloric effect in distorted kagome lattice Gd$_3$BWO$_9$

We synthesize the high-quality Gd$_3$BWO$_9$ single crystal and investigate its lowtemperature magnetic and thermodynamic properties. Below $T\rm_{N}$ = 1.08 K, the anisotropic behavior of magnetic susceptibilities reveals that the Gd$^{3+}$ moments exhibit the dominant antiferromagnetic coupling along the $c$-axis, while displaying a ferromagnetic arrangement in kagome plane. With pronounced magnetic frustration, in adiabatic demagnetization refrigeration experiments starting from initial conditions of 9 T and 2 K, Gd$_3$BWO$_9$ polycrystal reaches a minimum temperature of 0.151 K, significantly lower than its $T\rm_{N}$. Due to the high density of Gd$^{3+}$ ions ($S$=7/2), the maximum magnetic entropy change reaches over 50 J kg$^{-1}$ K$^{-1}$ under fields up to 7 T in Gd$_3$BWO$_9$, nearly 1.5 times as large as commercial sub-Kelvin magnetic coolant Gd$_3$Ga$_5$O$_{12}$(GGG). The H-T phase diagram of Gd$_3$BWO$_9$ under $H$//$c$ exhibits field-induced critical behavior near the phase boundaries. This observation aligns with the theoretical scenario in which a quantum critical point acts as the endpoint of a line of classical second-order phase transitions. Such behavior suggests the importance of further investigations into the divergence of magnetic Gr\"uneisen parameter in the vicinity of critical field at ultralow temperatures.

cond-mat.str-el

Conformal Sets in Multiple-Choice Question Answering under Black-Box Settings with Provable Coverage Guarantees

Large Language Models (LLMs) have shown remarkable progress in multiple-choice question answering (MCQA), but their inherent unreliability, such as hallucination and overconfidence, limits their application in high-risk domains. To address this, we propose a frequency-based uncertainty quantification method under black-box settings, leveraging conformal prediction (CP) to ensure provable coverage guarantees. Our approach involves multiple independent samplings of the model's output distribution for each input, with the most frequent sample serving as a reference to calculate predictive entropy (PE). Experimental evaluations across six LLMs and four datasets (MedMCQA, MedQA, MMLU, MMLU-Pro) demonstrate that frequency-based PE outperforms logit-based PE in distinguishing between correct and incorrect predictions, as measured by AUROC. Furthermore, the method effectively controls the empirical miscoverage rate under user-specified risk levels, validating that sampling frequency can serve as a viable substitute for logit-based probabilities in black-box scenarios. This work provides a distribution-free model-agnostic framework for reliable uncertainty quantification in MCQA with guaranteed coverage, enhancing the trustworthiness of LLMs in practical applications.

cs.CL