SearcharxivSearch

arXiv subjects

Xing Wang

Publications and source records attributed to Xing Wang.

At least 19 recordsLinked to original sources

From Language to Behavior: Scaling Sequence Transformers for Industrial Recommendation Ranking with Rec-Native Designs

Scaling Transformers has driven large gains in language modeling, but transplanting this to behavior-sequence modeling in production ranking is challenging: recommendation differs in signal quality, where behavior sequences are noisy, temporally irregular, and sparsely supervised, and in computation asymmetry, where each request scores many candidates against one shared user history under tight latency budgets. We propose ReST, a recommendation-native Transformer scaling framework. For signal quality, it introduces a sequence encoder with dual-gated attention, rotary positional and temporal embedding, stabilized residual normalization, and training-only auxiliary objectives. For computation asymmetry, it factorizes ranking into a heavy reusable encoder and a lightweight cross decoder with projection-free KV attention and token-specific parameterization, coupling user-level shared-prefix training with shared-prefix serving for compute-once, decode-many-times ranking. Across industrial and public benchmarks, ReST achieves higher accuracy and scales more consistently along sequence length, depth, and width, where LLM-style Transformer blocks saturate. A one-week online A/B test on a production advertising platform improves online AUC by 1.31% and lifts a core revenue metric by 11.93% within a 50 ms P99 budget; ReST has since been fully deployed in production, showing that behavior-sequence scaling remains a promising, under-exploited axis for production ranking.

cs.IR

Auditing Harness Tampering in Self-Improving Agents

Self-improving agents iteratively modify their own harness to push the frontier of their performance. However, such modifications can produce illusory performance gains or compromise integrity constraints such as authorization, provenance, and completeness without genuinely improving capability. We term this phenomenon as harness tampering, which extends the concept from reward and measurement tampering to the full self-improvement lifecycle. To systematically study this problem, we propose a two-axis taxonomy that categorizes each misaligned edit by the harness functional role in which it occurs and the obligation it violates. Then we build an annotated corpus by seeding tampered-benign edit pairs into the real trajectories of self-improving agents. We adapt and benchmark diverse audit methods on tampering classification and localization tasks. Finally we systematically audit real trajectories of self-improving agents. The results demonstrate that harness tampering consistently occurs in real runs from different agents, often persists in the lineage of the best agent, and forms distinct system-specific profiles across the taxonomy.

cs.CL

TrustDABench: Benchmarking Reliability and Robustness of LLMs for Structured Data Analysis

LLMs are increasingly used to analyze spreadsheets, CSV files, and other structured data, but producing a correct-looking answer is not the same as producing a trustworthy analysis. A trustworthy result should be supported by a valid path from the user question to the relevant data evidence. This requirement creates two diagnostic questions: whether an LLM can refuse to answer or ask for clarification when such a path does not exist, and whether it can preserve the correct analysis when the same evidence is expressed in different table forms. We introduce TrustDABench, a benchmark that operationalizes these questions as reliability and robustness. Starting from the evidence-path view, we derive 19 perturbation operators and instantiate them through an Agentic-LLM-based generation framework. TrustDABench contains 2,340 human-verified perturbed instances, and we evaluate eight representative LLMs. The results show substantial headroom: the best reliability result is only 24.21% average MRS, achieved by GPT-5.5, while the best robustness result still has 9.10% average ASR, achieved by Claude-Sonnet-5. The failures are systematic: models rarely detect conflicting evidence, often continue along executable but unsupported analysis paths, and remain sensitive to perturbations that change observation boundaries or cross-table relations. These findings suggest that stronger evidence-boundary recognition and representation-invariant reasoning are still needed for reliable structured-data analysis.

cs.CL

Intersection matrices associated to geometric-ordered bases of Feynman integrals

In integration-by-parts reduction of Feynman integrals, the order relation in the Laporta algorithm determines a set of master integrals. In this paper we investigate the intersection matrices of the integrands of the master integrals that are obtained from a geometric order relation. With an appropriate definition of integrands and their duals, we find that the intersection matrices are simpler than expected: For a filtration-compatible basis, the entries of the intersection matrix are Laurent polynomials in the dimensional regularisation parameter $\varepsilon$. For an $\varepsilon$-factorised basis, the entries are instead integers, up to an overall power of $\varepsilon$, if the boundary values for the auxiliary functions of the rotation are chosen appropriately. This has practical consequences: We can systematically eliminate certain auxiliary transcendental functions, introduced in going from a filtration-compatible basis to an $\varepsilon$-factorised basis. We provide an algorithm that performs this elimination while minimising the number of required calculations.

hep-th

Testing for Smooth Structural Change in Cointegrated Systems

This paper develops an econometric framework for analysing smooth structural change in cointegrated systems following a known intervention time. We consider a vector error-correction model in which the cointegration rank and the pre-intervention cointegrating structure are identified from a stable pre-intervention subsample. After the intervention, both the adjustment coefficients and the cointegrating vectors are allowed to evolve smoothly as functions of rescaled time, which are estimated using kernel-weighted local reduced-rank methods. The analysis is formulated directly in a cointegrated VAR/VECM system, which preserves the treatment of long-run relations and short-run error-correction dynamics. By working with the decomposition $\Pi(\delta)=\alpha(\delta)\beta(\delta)'$, the method separates changes in the equilibrium relation from those in the speed of adjustment. We also provide two tests for the parameter consistency and the post-intervention parameter smoothness respectively. An empirical application to energy market, foreign-exchange, and gold-market index around the 24 February 2022 Russia's invasion of Ukraine illustrates how the proposed approach distinguishes between a discrete regime shift and smooth post-intervention evolution. The results suggest that cointegrating relation among the price of Brent crude oil, the spot exchange rate (USD/EUR), and the Credit Suisse NASDAQ Gold Price Index has smoothly changed after the outbreak of war, instead of a constant long-run conintegration system in the pre-intervention period.

econ.EM

Analytic result of a three-loop integral family in the Higgs decay to four massive bottom quarks

When calculating the decay width of $H\to b\bar{b}b\bar{b}$ induced by the effective Higgs-gluon-gluon interaction using the optical theorem, the three-loop Feynman diagrams are encountered. Among them, there is an interesting integral family, which involves not only the sectors evaluated to multiple polylogarithms but also those related to an elliptic curve and a K3 surface. We demonstrate how to construct and solve the canonical differential equation for the entire family, which includes mixing between different sectors, by using the recently proposed algorithm in~\cite{e-collaboration:2025frv, Bree:2025tug}. Analytical results for all master integrals in this family up to the first two orders of $\varepsilon$ are given.

hep-ph

Observation of stopping power reduction at strong ion-plasma coupling

Ion stopping in dense plasma is crucial for stellar evolution and fusion ignition. However, its behavior in the strong ion-plasma coupling regime beyond the linear limit has long remained elusive, due to formidable experimental challenges. Here we report the first experimental investigation of ion stopping at an unprecedented coupling parameter exceeding unity, achieved by sending laser-accelerated short-pulse and intense quasi-monoenergetic carbon ions ($\sim$583 keV/u, C$^{5+}$) into a uniform, long-lived, well-characterized dense plasma target ($T_e$ $\approx$ 17 eV, $n_e$ $\approx$ 4$\times$10$^{20}$ cm$^{-3}$). By simultaneously measuring ion energy loss and charge-state evolution, we eliminated key experimental ambiguities arising from charge-state determination. Our results clearly show a reduction in stopping power compared with predictions from standard linear dielectric response or binary collision models, and they agree well with the hybrid calculation of molecular dynamics with quantum corrections. The importance of nonlinear screening effects arising from many-body interactions and quantum effects due to the wave nature of electrons was demonstrated at strong coupling. This work establishes a definitive high-fidelity experimental benchmark for collisional dynamics in the strong-coupling regime. It offers critical insight for accurate modeling of energy transport in inertial confinement fusion and astrophysical plasmas.

physics.plasm-ph

Precision physics at the muon collider: $m_W$ and CKM matrix elements

We examine the potential for a 10~TeV lepton collider to carry out precision measurements of the W boson mass and W boson couplings strength, i.e. the CKM matrix elements. We consider the several W boson production mechanisms and focus on the most copious at 10~TeV, that is effective $\gamma W \to W$, a process viable at both opposite sign and same-sign leptonic colliders. We find that the leptonic W decay channel can hardly be competitive with present determinations, due to lack of rate. The hadronic channel has potential to improve over the current $\simeq$10~MeV from measurements at hadron colliders, motivating detector developments towards high-precision hadronic energy measurements. We find that the precision understanding of the detector response to hadrons can also lead to a determination of the CKM matrix elements. We expect determination of CKM matrix elements surpassing by far the present precision for couplings involving heavy quarks, notably $V_{cb}$, avoiding the present bottle-necks due to poor knowledge of hadronic matrix elements needed in low energy extractions of CKM matrix elements. Our findings motivate detector developments towards high-precision hadronic energy measurements and flavor tagging.

hep-ph

Motion-Aware Caching for Efficient Autoregressive Video Generation

Autoregressive video generation paradigms offer theoretical promise for long video synthesis, yet their practical deployment is hindered by the computational burden of sequential iterative denoising. While cache reuse strategies can accelerate generation by skipping redundant denoising steps, existing methods rely on coarse-grained chunk-level skipping that fails to capture fine-grained pixel dynamics. This oversight is critical: pixels with high motion require more denoising steps to prevent error accumulation, while static pixels tolerate aggressive skipping. We formalize this insight theoretically by linking cache errors to residual instability, and propose MotionCache, a motion-aware cache framework that exploits inter-frame differences as a lightweight proxy for pixel-level motion characteristics. MotionCache employs a coarse-to-fine strategy: an initial warm-up phase establishes semantic coherence, followed by motion-weighted cache reuse that dynamically adjusts update frequencies per token. Extensive experiments on state-of-the-art models like SkyReels-V2 and MAGI-1 demonstrate that MotionCache achieves significant speedups of $\textbf{6.28}\times$ and $\textbf{1.64}\times$ respectively, while effectively preserving generation quality (VBench: $1\%\downarrow$ and $0.01\%\downarrow$ respectively). The code is available at https://github.com/ywlq/MotionCache.

cs.CV

Multi-output Extreme Spatial Model for Complex Aircraft Production Systems

Problem definition: Data-driven models in machine learning have enabled efficient management of production systems. However, a majority of machine learning models are devoted to modeling the mean response or average pattern, which is inappropriate for studying abnormal extreme events that are often of primary interest in aircraft manufacturing. Since extreme events from heavy-tailed distributions give rise to prohibitive expenditures in system management, sophisticated extreme models are urgently needed to analyze complex extreme risks. Engineering applications of extreme models usually focus on individual extreme events, which is insufficient for complex systems with correlations. Methodology/results: We introduce an extreme spatial model for multi-output response control systems that efficiently captures the dynamics using a bilinear function on two spatial domains for control variables and measurement locations. Marginal parameter modeling and extremal dependence have been investigated. In addition, an efficient graph-assisted composite likelihood estimation and corresponding computational algorithms are developed to cope with high-dimensional outputs. The application to composite aircraft production shows that the proposed model enables comprehensive analyses with superior predictive performance on extreme events compared to canonical methods. Managerial implications: Our method shows how to use an extreme spatial model for predicting extreme events and managing extreme risks in complex production systems such as aircraft. This can help achieve better quality management and operation safety in aircraft production systems and beyond.

stat.AP

Higgs boson decay to massive bottom quarks at order $\alpha_s^4$ induced by top-quark Yukawa couplings

The Higgs boson decay to massive bottom quarks has the largest branching ratio. The decay is mainly induced by the bottom-quark Yukawa coupling with the decay rate calculated up to $O(\alpha_s^4)$ assuming the massless final-state bottom quark. The top-quark Yukawa coupling induced contribution starts at $O(\alpha_s^2)$, and exhibits logarithmic and power enhancements, making the perturbative expansion converge slowly, which is a feature not present in the hadronic Higgs boson decay. We present a calculation of such contributions at $O(\alpha_s^4)$ to the decay into massive bottom quarks in which the squared amplitudes contain two top-quark Yukawa couplings and the final state must include at least a bottom quark pair. We find that they increase the decay width, relative to the result up to $O(\alpha_s^3)$, by $0.4\%$, larger than the experimental precision at future lepton colliders, and reduce the scale dependence significantly down to $0.4\%$.

hep-ph

optimade-maker: Automated generation of interoperable materials APIs from static datasets

Atomistic structural data are central to materials science, condensed matter physics, and chemistry, and are increasingly digitised across diverse repositories and databases. Interoperable access to these heterogeneous data sources enables reusable clients and tools, and is essential for cross-database analyses and data-driven materials discovery. Toward this aim, the OPTIMADE (Open Databases Integration for Materials Design) specification defines a standard REST API for atomistic structures and related properties. However, deploying and maintaining compliant services remains technically demanding and poses a significant barrier for many data providers. Here, we present optimade-maker, a lightweight toolkit for the automated generation of OPTIMADE-compliant APIs directly from raw atomistic structure and property data. The toolkit supports a wide range of raw datasets, enables conversion to a standardised OPTIMADE data representation, and allows for rapid deployment of APIs in both local and production environments. We further demonstrate it through an automated service on the Materials Cloud Archive, which automatically creates and publishes OPTIMADE APIs for contributed datasets, enabling immediate discoverability and interoperability. In addition, we implement data transformation pipelines for the Cambridge Structural Database (CSD) and the Inorganic Crystal Structure Database (ICSD), enabling unified access to these curated resources through the OPTIMADE framework. By lowering the technical barriers to interoperable data publication, optimade-maker represents an important step toward a scalable, FAIR materials data ecosystem integrating both community-contributed and curated databases.

cs.DB

An algorithm towards $\varepsilon$-factorising Feynman Integrals

In this talk, we use several examples to elaborate on how a recently proposed algorithm can turn non-trivial Feynman integrals into an $\varepsilon $-factorised manner, regardless of their hidden geometric essence. In particular, some extra details about three-loop banana integrals with unequal-mass configuration are provided.

hep-th

TAP: A Token-Adaptive Predictor Framework for Training-Free Diffusion Acceleration

Diffusion models achieve strong generative performance but remain slow at inference due to the need for repeated full-model denoising passes. We present Token-Adaptive Predictor (TAP), a training-free, probe-driven framework that adaptively selects a predictor for each token at every sampling step. TAP uses a single full evaluation of the model's first layer as a low-cost probe to compute proxy losses for a compact family of candidate predictors (instantiated primarily with Taylor expansions of varying order and horizon), then assigns each token the predictor with the smallest proxy error. This per-token "probe-then-select" strategy exploits heterogeneous temporal dynamics, requires no additional training, and is compatible with various predictor designs. TAP incurs negligible overhead while enabling large speedups with little or no perceptual quality loss. Extensive experiments across multiple diffusion architectures and generation tasks show that TAP substantially improves the accuracy-efficiency frontier compared to fixed global predictors and caching-only baselines.

cs.CV

S2O: Early Stopping for Sparse Attention via Online Permutation

Attention scales quadratically with sequence length, fundamentally limiting long-context inference. Existing block-granularity sparsification can reduce latency, but coarse blocks impose an intrinsic sparsity ceiling, making further improvements difficult even with carefully engineered designs. We present S2O, which performs early stopping for sparse attention via online permutation. Inspired by virtual-to-physical address mapping in memory systems, S2O revisits and factorizes FlashAttention execution, enabling inference to load non-contiguous tokens rather than a contiguous span in the original order. Motivated by fine-grained structures in attention heatmaps, we transform explicit permutation into an online, index-guided, discrete loading policy; with extremely lightweight preprocessing and index-remapping overhead, it concentrates importance on a small set of high-priority blocks. Building on this importance-guided online permutation for loading, S2O further introduces an early-stopping rule: computation proceeds from high to low importance; once the current block score falls below a threshold, S2O terminates early and skips the remaining low-contribution blocks, thereby increasing effective sparsity and reducing computation under a controlled error budget. As a result, S2O substantially raises the practical sparsity ceiling. On Llama-3.1-8B under a 128K context, S2O reduces single-operator MSE by 3.82$\times$ at matched sparsity, and reduces prefill compute density by 3.31$\times$ at matched MSE; meanwhile, it preserves end-to-end accuracy and achieves 7.51$\times$ attention and 3.81$\times$ end-to-end speedups.

cs.LG

A Real-World Grasping-in-Clutter Performance Evaluation Benchmark for Robotic Food Waste Sorting

Food waste management is critical for sustainability, yet inorganic contaminants hinder recycling potential. Robotic automation accelerates sorting through automated contaminant removal. Nevertheless, the diverse and unpredictable nature of contaminants introduces major challenges for reliable robotic grasping. Grasp performance benchmarking provides a rigorous methodology for evaluating these challenges in underexplored field contexts like food waste sorting. However, existing approaches suffer from limited simulation datasets, over-reliance on simplistic metrics like success rate, inability to account for object-related pre-grasp conditions, and lack of comprehensive failure analysis. To address these gaps, this work introduces GRAB, a real-world grasping-in-clutter (GIC) performance benchmark incorporating: (1) diverse deformable object datasets, (2) advanced 6D grasp pose estimation, and (3) explicit evaluation of pre-grasp conditions through graspability metrics. The benchmark compares industrial grasping across three gripper modalities through 1,750 grasp attempts across four randomized clutter levels. Results reveal a clear hierarchy among graspability parameters, with object quality emerging as the dominant factor governing grasp performance across modalities. Failure mode analysis shows that physical interaction constraints, rather than perception or control limitations, constitute the primary source of grasp failures in cluttered environments. By enabling identification of dominant factors influencing grasp performance, GRAB provides a principled foundation for designing robust, adaptive grasping systems for complex, cluttered food waste sorting.

cs.RO

Train Short, Inference Long: Training-free Horizon Extension for Autoregressive Video Generation

Autoregressive video diffusion models have emerged as a scalable paradigm for long video generation. However, they often suffer from severe extrapolation failure, where rapid error accumulation leads to significant temporal degradation when extending beyond training horizons. We identify that this failure primarily stems from the spectral bias of 3D positional embeddings and the lack of dynamic priors in noise sampling. To address these issues, we propose FLEX (Frequency-aware Length EXtension), a training-free inference-time framework that bridges the gap between short-term training and long-term inference. FLEX introduces Frequency-aware RoPE Modulation to adaptively interpolate under-trained low-frequency components while extrapolating high-frequency ones to preserve multi-scale temporal discriminability. This is integrated with Antiphase Noise Sampling (ANS) to inject high-frequency dynamic priors and Inference-only Attention Sink to anchor global structure. Extensive evaluations on VBench demonstrate that FLEX significantly outperforms state-of-the-art models at 6x extrapolation (30s duration) and matches the performance of long-video fine-tuned baselines at 12x scale (60s duration). As a plug-and-play augmentation, FLEX seamlessly integrates into existing inference pipelines for horizon extension. It effectively pushes the generation limits of models such as LongLive, supporting consistent and dynamic video synthesis at a 4-minute scale. Project page is available at https://ga-lee.github.io/FLEX_demo.

cs.CV

Improving integration-by-parts and differential equations

In this talk, we discuss how ideas from geometry help to improve Feynman integral reduction and the construction of $\varepsilon$-factorised differential equations. In particular, we outline a systematic procedure to obtain an $\varepsilon$-factorised differential equation for any Feynman integral.

hep-th