SearcharxivSearch

arXiv subjects

Yaobo Zhang

Publications and source records attributed to Yaobo Zhang.

13 recordsLinked to original sources

Algebraic versus physical uniqueness of MHV gravity numerators

We study whether a tree-level MHV gravity numerator is determined by its degree and by vanishing on $\langle ij\rangle=[ij]=0$ for every pair. A flag-variety standard-monomial basis and an $S_n$-resolved restriction map reduce the problem to exact finite-dimensional calculations. At seven points we find $W_{7,\mathbb{Q}}\simeq S^{(2,1^5)}\oplus S^{(1^7)}$. The Hodges numerator spans the sign summand, while the six-dimensional hook gives additional algebraic solutions. The pair-ideal conditions therefore do not determine a unique algebraic solution, but Bose symmetry selects the Hodges line. At eight points, pair-ideal conditions and Bose symmetry leave a two-dimensional alternating space. Same-helicity BCFW scaling, normalized collinear factorization, and the leading soft coefficient impose the same linear condition and select the Hodges line. We also prove that, at arbitrary multiplicity, an alternating fixed-degree numerator is determined by its full value on one collinear boundary with the marked legs and their spinor ratio fixed. Together with standard factorization, this determines the numerator up to normalization within the fixed-common-denominator ansatz. All rank and ideal-membership calculations use exact integer or rational arithmetic, and their finite-dimensional consequences are checked separately in Lean.

hep-th

An adaptive inverse-problem framework for one-loop five-gluon BCJ numerators

An inverse problem comprises a matrix equation together with its unknown space, physical data, equivalence relation, and validation tests. We formulate Bern--Carrasco--Johansson (BCJ) numerator construction as an exact adaptive inverse problem. A fixed scientific specification determines the theory, graph conventions, coefficient field, locality and power counting, cut data, observable equivalence, and independent checks. Each finite working specification compiles to $\mathcal{P}_\sigma=(A,b;\mathcal{S};\mathcal{T})$, where $Ax=b$ reconstructs numerator coefficients, $\mathcal{S}$ classifies the solution fiber, and $\mathcal{T}$ tests it on held-out information. Left-null obstructions identify candidate numerator-basis directions needed for consistency, while the action of candidate measurements on the right kernel identifies informative new cut equations. We illustrate these steps by hand at four points and apply them to one-loop five-gluon pure Yang--Mills theory. After kinematic and graph-symmetry reduction, the candidate numerator basis contains 1127 independent coordinates. The combined maximal, box, triple, and double cuts have rank 920, giving a 207-dimensional affine solution fiber. Exact reconstruction determines a particular solution and the complete ordered kernel. The specified $R_{12345}$ color-ring readout $\mathcal{S}$ annihilates every kernel direction, so the full fiber represents one observable class. A published forward-limit numerator lies in this fiber, and fresh cuts, independent integral reductions, and helicity-amplitude benchmarks validate the result. Explicit search rules and agent interfaces can propose revisions. Deterministic compilation and exact evaluation assess each proposal.

hep-th

PJ-RoPE: A Fourier-Jet-Affine Position Space for Relative Attention

We organize relative-position mechanisms in attention as a learnable Fourier-Jet-Affine position space. The starting point is lag-shift dynamics: a relative-position kernel is a response function of the lag \(d=i-j\), and the one-step shift \((Ef)(d)=f(d+1)\) gives a compact classification of finite structured responses through constant-coefficient difference modules. In this view, RoPE supplies simple Fourier roots, Jordan-RoPE thickens these roots into finite Fourier jets, and ALiBi supplies the repeated unit-root affine direction. NTK-aware RoPE scaling fits the same structure as a spectral flow of simple Fourier roots: moving the frequency grid generates first Fourier-jet tangent directions, while higher Taylor directions generate higher jets. PJ-RoPE makes these jet directions explicit and learnable, and uses the resulting space to measure task-level sector selection. The framework separates scalar PJ-bias kernels from exact PJ-rotary feature transforms, introduces sector-gate, effective-mass, functional-energy, and leave-one-order-out diagnostics, and stabilizes high-order coordinates with LC/rapidity compactification. Controlled probes recover designed sectors; synthetic teachers show trainable use; small byte-level language runs favor NTK-aware RoPE plus affine recency; symbolic music-token streams keep LC/affine variants strong with measurable high-order corrections; and LC diagnostics quantify the stability-resolution tradeoff.

cs.LG

Jordan-RoPE: Non-Semisimple Relative Positional Encoding via Complex Jordan Blocks

Relative positional encodings determine which functions of query-key lag can enter the primitive attention logit. RoPE supplies a rotary phase, while ALiBi supplies an additive distance bias. Motivated by group-theoretic views of linear translation-invariant positional encodings, we study a non-semisimple case in which a complex rotary eigenvalue and a nilpotent response live in the same defective Jordan block. The resulting relative operator generates oscillatory-polynomial features such as $e^{-\gamma d}\cos(\omega d)$, $e^{-\gamma d}\sin(\omega d)$, $d e^{-\gamma d}\cos(\omega d)$, and $d e^{-\gamma d}\sin(\omega d)$, for causal lag $d=i-j\geq 0$. Thus the construction realizes a distance-modulated phase basis $d e^{i\omega d}$, rather than merely adding a separate distance channel to RoPE. We formulate Exact Jordan-RoPE as a non-semisimple one-parameter representation, give its real block form, and specify the contragredient query action required by non-orthogonal positional maps. We also distinguish this exact representation from stabilized variants whose bounded shear improves numerical behavior but breaks the exact group law. Kernel-level diagnostics and a Jordan-friendly synthetic language-model task show that the coupled Jordan basis is useful when the target contains distance-modulated phase interactions. On a small WikiText-103 byte language model, a scaled-exact variant improves over RoPE and direct-sum baselines within the Jordan family, while RoPE+ALiBi remains strongest overall. The evidence is structural rather than a broad performance claim.

cs.LG

Exact CHY Integrand Construction Using Combinatorial Neural Networks and Discrete Optimization

Constructing a rational CHY integrand that realizes prescribed physical pole constraints is a discrete inverse problem whose combinatorial complexity grows with multiplicity. We encode the pole hierarchy through generalized pole degrees $K(A)$ (channels $s_A$), defined as signed internal-edge counts associated with particle subsets in a colored integrand graph. Additivity under integrand multiplication together with the elementary face recursion on the subset lattice expresses all higher-channel $K(A)$ as linear functions of the two-particle data $\{K(s_{ij})\}$ and reduces the inverse step to a mixed-integer linear feasibility problem. The subset lattice provides a fixed dependency graph for deterministic message passing with forward evaluation and backward residual propagation; this computation is parameter-free and involves no training. In factorial-rescaled variables $\widetilde K(A)=(|A|-2)!\,K(A)$, every local update is integral, so propagation is exact in the rescaled recursion variables and does not rely on numerical reconstruction. We further organize generalized integrand graphs by an $n$-regular grading under multiplication, where degree-zero (0-regular) factors act as M\"obius-invariant insertions that can be decomposed into four-point cross ratios. We illustrate the construction at six and eight points, including pick-pole selection and higher-order pole reduction.

hep-th

General One-loop Generating Function by IBP relations

In this paper we have studied the most general generating function of reduction for one loop integrals with arbitrary tensor structure in numerator and arbitrary power distribution of propagators in denominator. Using IBP relations, we have established the partial differential equations for these generating functions and solved them analytically. These results provide useful guidance for applying generating function method to reductions of higher loop integrals.

hep-ph

Reduction of General One-loop Integrals Using Auxiliary Vector

As a key method to deal with loop integrals, Integration-By-Parts (IBP) method can be used to do reduction as well as establish the differential equations for master integrals. However, when talking about tensor reduction, the Passarino-Veltman (PV) reduction method is also widely used for one-loop integrals. Recently, we have proposed an improved PV reduction method, i.e., the PV reduction method with auxiliary vector $R$, which can easily give analytical reduction results for any tensor rank. However, our results are only for integrals with propagators with power one. In this paper, we generalize our method to one-loop integrals with general tensor structures and propagators with general powers. Our ideas are simple. We solve the generalised reduction problem by combining differentiation over masses and proper limit of reduction with power-one propagators. Finally, we demonstrate our method with several examples. With the result in this paper, we have shown that our improved PV-reduction method with auxiliary vector is a self-completed reduction method for one-loop integrals.

hep-th

Note on solutions of scattering equations

In the CHY-frame for the amplitudes, there are two kinds of singularities we need to deal with. The first one is the pole singularities when the kinematics is not general, such that some of $S_A\to 0$. The second one is the collapse of locations of points after solving scattering equations (i.e., the singular solutions). These two types of singularities are tightly related to each other, but the exact mapping is not well understood. In this paper, we have initiated the systematic study of the mapping. We have demonstrated the different mapping patterns using three typical situations, i.e., the factorization limit, the soft limit and the forward limit.

hep-th

Rademacher complexity of noisy quantum circuits

Noise in quantum systems is a major obstacle to implementing many quantum algorithms on large quantum circuits. In this work, we study the effects of noise on the Rademacher complexity of quantum circuits, which is a measure of statistical complexity that quantifies the richness of classes of functions generated by these circuits. We consider noise models that are represented by convex combinations of unitary channels and provide both upper and lower bounds for the Rademacher complexities of quantum circuits characterized by these noise models. In particular, we find a lower bound for the Rademacher complexity of noisy quantum circuits that depends on the Rademacher complexity of the corresponding noiseless quantum circuit as well as the free robustness of the circuit. Our results show that the Rademacher complexity of quantum circuits decreases with the increase in noise.

quant-ph

Effects of quantum resources on the statistical complexity of quantum circuits

We investigate how the addition of quantum resources changes the statistical complexity of quantum circuits by utilizing the framework of quantum resource theories. Measures of statistical complexity that we consider include the Rademacher complexity and the Gaussian complexity, which are well-known measures in computational learning theory that quantify the richness of classes of real-valued functions. We derive bounds for the statistical complexities of quantum circuits that have limited access to certain resources and apply our results to two special cases: (1) stabilizer circuits that are supplemented with a limited number of T gates and (2) instantaneous quantum polynomial-time Clifford circuits that are supplemented with a limited number of CCZ gates. We show that the increase in the statistical complexity of a quantum circuit when an additional quantum channel is added to it is upper bounded by the free robustness of the added channel. Finally, we derive bounds for the generalization error associated with learning from training data arising from quantum circuits.

quant-ph

On the statistical complexity of quantum circuits

In theoretical machine learning, the statistical complexity is a notion that measures the richness of a hypothesis space. In this work, we apply a particular measure of statistical complexity, namely the Rademacher complexity, to the quantum circuit model in quantum computation and study how the statistical complexity depends on various quantum circuit parameters. In particular, we investigate the dependence of the statistical complexity on the resources, depth, width, and the number of input and output registers of a quantum circuit. To study how the statistical complexity scales with resources in the circuit, we introduce a resource measure of magic based on the $(p,q)$ group norm, which quantifies the amount of magic in the quantum channels associated with the circuit. These dependencies are investigated in the following two settings: (i) where the entire quantum circuit is treated as a single quantum channel, and (ii) where each layer of the quantum circuit is treated as a separate quantum channel. The bounds we obtain can be used to constrain the capacity of quantum neural networks in terms of their depths and widths as well as the resources in the network.

quant-ph

Depth-Width Trade-offs for Neural Networks via Topological Entropy

One of the central problems in the study of deep learning theory is to understand how the structure properties, such as depth, width and the number of nodes, affect the expressivity of deep neural networks. In this work, we show a new connection between the expressivity of deep neural networks and topological entropy from dynamical system, which can be used to characterize depth-width trade-offs of neural networks. We provide an upper bound on the topological entropy of neural networks with continuous semi-algebraic units by the structure parameters. Specifically, the topological entropy of ReLU network with $l$ layers and $m$ nodes per layer is upper bounded by $O(l\log m)$. Besides, if the neural network is a good approximation of some function $f$, then the size of the neural network has an exponential lower bound with respect to the topological entropy of $f$. Moreover, we discuss the relationship between topological entropy, the number of oscillations, periods and Lipschitz constant.

cs.LG

Note on the Labelled tree graphs

In the CHY-frame for the tree-level amplitudes, the bi-adjoint scalar theory has played a fundamental role because it gives the on-shell Feynman diagrams for all other theories. Recently, an interesting generalization of the bi-adjoint scalar theory has been given in arXiv:1708.08701 by the "Labelled tree graphs", which carries a lot of similarity comparing to the bi-adjoint scalar theory. In this note, we have investigated the Labelled tree graphs from two different angels. In the first part of the note, we have shown that we can organize all cubic Feynman diagrams produces by the Labelled tree graphs to the "effective Feynman diagrams". In the new picture, the pole structure of the whole theory is more manifest. In the second part, we have generalized the action of "picking pole" in the bi-adjoint scalar theory to general CHY-integrands which produce only simple poles.

hep-th