Searcharxiv⌕ Search

arXiv subjects

Jianxiong Wang

Publications and source records attributed to Jianxiong Wang.

12 recordsLinked to original sources

GSM: Efficient Language Modeling with Shared Global State

Efficient language models must reduce not only the cost of individual accesses to past context but also the overhead of repeatedly selecting and processing historical information across layers. We introduce the Global State Model (GSM), a causal encoder--decoder architecture that concentrates the selection and aggregation of long-range information in the encoding stage. Through multiple stages of history retrieval, the encoder progressively incorporates long-range information into representations at recent positions, forming a shared state with a fixed window size. Each decoder layer accesses this same state using queries updated from the preceding layer, preserving computational depth while avoiding repeated construction of historical key--value (KV) representations and long-range indexing. As a result, neither the decoder's per-step attention cost nor its KV cache size grows with the history length. Experiments show that GSM improves computational efficiency and reduces cache overhead while maintaining model performance and the ability to use long-range information, offering a shared-state architecture for efficient language modeling.

cs.CL↗

Lngram v2: Latent N-Gram Memory with Interpretable Discrete Representations

Transformers lack a native lookup mechanism, requiring repeated dense computation to recognize and reuse local static patterns. Lngram v1 introduces tokenizer-independent conditional memory through discrete latent n-gram addressing, but its memory capacity is coupled with the backbone width, limiting scalability due to high parameter and activation costs. We propose Lngram v2, which decouples the number of routes, memory dimension, and backbone width, and introduces a context-aware grouped-query attention readout to scale memory capacity independently. A zero-value Sink and counterfactual surrogate gradients further improve readout selectivity and routing trainability while preserving hard discrete addressing. Experiments across vision--language models (VLMs) of different scales show consistent improvements, including successful scaling to a 30B-parameter model. Compared with Lngram v1, Lngram v2 substantially reduces both total and activated memory parameters while maintaining or improving language modeling performance. Further analysis shows that its discrete IDs preserve substantial semantic structure of continuous hidden states, enabling semantic recovery from IDs alone and stable ID--semantic associations across datasets. These results establish Lngram v2 as an efficient and scalable latent conditional memory mechanism whose discrete addresses also provide a structured interface for analyzing internal model representations.

cs.CL↗

Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents

Open-world video understanding often requires a model to locate sparse visual evidence and acquire external knowledge that is absent from the video and its parametric memory. While Thinking-with-Videos enables active temporal perception and Deep Research supports multi-step information seeking, the two capabilities are typically developed in isolation. We introduce VideoRover, a unified Video Deep Research framework that iteratively coordinates video cropping, multimodal search, and webpage browsing. Given a video-question pair, VideoRover uses each tool result to select the next action, so localized video clips guide external retrieval and retrieved evidence triggers further video inspection and verification. To develop this capability, we construct an automated data curation pipeline, producing 26K verified SFT trajectories and 3K challenging RL instances. We also introduce VideoRover-Bench, a benchmark stratified by video duration and research difficulty. Experiments on VideoDR and VideoRover-Bench show that our VideoRover-8B-RL achieves performance comparable to proprietary models in the direct-answer setting without tool use while outperforming larger open-source models equipped with the same tool suite. Ablation studies and training dynamics further validate the complementary roles of active video grounding, external retrieval, and long-horizon reinforcement learning.

cs.CV↗

Symmetry of Solutions to Fractional Semilinear Equations on Hyperbolic Spaces

We study a semilinear equation involving the fractional Laplacian on the hyperbolic space $\mathbb{H}^n$. Unlike in conformally compact Einstein manifolds, the fractional Laplacian on $\mathbb{H}^n$ does not enjoy conformal covariance. By employing Helgason-Fourier analysis, we explicitly derive the Green's function of the fractional Laplacian on $\mathbb{H}^n$ as well as its asymptotic behaviors. We then apply a direct method of moving planes to the integral form of the equation, and show that nonnegative weak solutions are symmetric. In addition, we extend several maximum principles to hyperbolic space.

math.AP↗

Conformally invariant equations with negative critical exponents on the three dimensional hyperbolic space

We establish a symmetry result for positive entire solutions with a prescribed growth rate to the following fourth order equation on the 3-dimensional hyperbolic space $\mathbb{H}^3$: \[ P_2 u = - u^{-7}, \] where $P_2$ denotes the fourth-order Paneitz operator. We prove that any positive solution $u$ on $\mathbb{H}^3$ exhibiting exponential growth at infinity must, up to hyperbolic isometries, be radial and strictly decreasing with respect to some point $P \in \mathbb{H}^3$. Fourth order equations with negative critical growth on 3-dimensional Euclidean space $\mathbb{R}^3$ has been studied by Choi and Xu in \cite{CX09 }, and subsequently by McKenna and Reichel \cite{MR03} and Xu \cite{Xu05}. Unlike the Euclidean case, the behavior of the Green's function of $P_2$ is substantially different, which prevents us from using the moving plane (sphere) method directly.

math.AP↗

Interpretable Logical Anomaly Classification via Constraint Decomposition and Instruction Fine-Tuning

Logical anomalies are violations of predefined constraints on object quantity, spatial layout, and compositional relationships in industrial images. While prior work largely treats anomaly detection as a binary decision, such formulations cannot indicate which logical rule is broken and therefore offer limited value for quality assurance. We introduce Logical Anomaly Classification (LAC), a task that unifies anomaly detection and fine-grained violation classification in a single inference step. To tackle LAC, we propose LogiCls, a vision-language framework that decomposes complex logical constraints into a sequence of verifiable subqueries. We further present a data-centric instruction synthesis pipeline that generates chain-of-thought (CoT) supervision for these subqueries, coupling precise grounding annotations with diverse image-text augmentations to adapt vision language models (VLMs) to logic-sensitive reasoning. Training is stabilized by a difficulty-aware resampling strategy that emphasizes challenging subqueries and long tail constraint types. Extensive experiments demonstrate that LogiCls delivers robust, interpretable, and accurate industrial logical anomaly classification, providing both the predicted violation categories and their evidence trails.

cs.CV↗

Search potential for direct slepton pair production at the CEPC with $\sqrt{s}$ = 360 GeV

The Circular Electron Positron Collider (CEPC) is designed to operate at the key center-of-mass energies: 91.2 GeV as a Z factory for precision Z boson studies,$\approx$ 160 GeV at the threshold for W boson pair production, and 240 GeV as a Higgs factory for copious Higgs boson production. It can be upgraded to 360 GeV (CEPC-360GeV) for enabling top quark-antiquark ($t\bar{t}$) pair production. Beyond enabling high-precision measurements of the Standard Model (SM), CEPC-360GeV is uniquely position to perform searches for new physics beyond the SM (BSM) physics, serving as a valuable complement to hadron colliders. This paper presents a sensitivity study on the direct pair production of staus and smuons at the CEPC with $\sqrt{s}$ = 360 GeV, conducted via full Monte Carlo (MC) simulation. Under the assumptions of integrated luminosity 1.0 ab^{-1} and a flat 5% systematic uncertainty, CEPC-360GeV could potentially discover the combined production of left-handed and right-handed staus up to a mass of 170 GeV (if they exist), or up to 169 GeV for pure left-handed staus and 162 GeV for pure right-handed staus. For direct smuon production, the discovery potential reaches up to 178 GeV under the same conditions.

hep-ex↗

The Fourth Monocular Depth Estimation Challenge

This paper presents the results of the fourth edition of the Monocular Depth Estimation Challenge (MDEC), which focuses on zero-shot generalization to the SYNS-Patches benchmark, a dataset featuring challenging environments in both natural and indoor settings. In this edition, we revised the evaluation protocol to use least-squares alignment with two degrees of freedom to support disparity and affine-invariant predictions. We also revised the baselines and included popular off-the-shelf methods: Depth Anything v2 and Marigold. The challenge received a total of 24 submissions that outperformed the baselines on the test set; 10 of these included a report describing their approach, with most leading methods relying on affine-invariant predictions. The challenge winners improved the 3D F-Score over the previous edition's best result, raising it from 22.58% to 23.05%.

cs.CV↗

The Method of Moving Spheres on the Hyperbolic Space and the Classification of Solutions and the prescribed Q-curvature problem

The classification of solutions to semilinear partial differential equations, as well as the classification of critical points of the corresponding functionals, have wide applications in the study of partial differential equations and differential geometry. The classical moving plane method and the method of moving sphere on the Euclidean space $\mathbb{R}^n$ provide an effective approach to capture the symmetry of solutions. As far as we know, the moving sphere method has yet to be developed on the hyperbolic space $\mathbb{H}^n$. In the present paper, we focus on the following equation \begin{equation*} P_k u = f(u) \end{equation*} on hyperbolic spaces $\mathbb{H}^n$, where $P_k$ denotes the GJMS operators on $\mathbb{H}^n$ and $f : \mathbb{R} \to \mathbb{R}$ satisfies certain growth conditions. We develop a moving sphere approach on $\mathbb{H}^n$ to obtain the symmetry propertyas well as the classification of positive solutions to the above equation. Our methods also rely on the Helgason-Fourier analysis and Hardy-Littlewood-Sobolev inequalities on hyperbolic space together with a Kelvin transform we introduce on the hyperbolic space in this paper. We also present applications to the higher order prescribed $Q$-curvature problem on the hyperbolic space.

math.AP↗

Symmetry of solutions to higher and fractional order semilinear equations on hyperbolic spaces

We show that nontrivial solutions to higher and fractional order equations with certain nonlinearity are radially symmetric and nonincreasing on geodesic balls in the hyperbolic space $\mathbb{H}^n$ as well as on the entire space $\mathbb{H}^n$. Applying the Helgason-Fourier analysis techniques on $\mathbb{H}^n$, we develop a moving plane approach for integral equations on $\mathbb{H}^n$. We also establish the symmetry to solutions of certain equations with singular terms on Euclidean spaces. Moreover, we obtain symmetry to solutions of some semilinear equations involving fractional order derivatives.

math.AP↗

Constructing backbone network by using tinker algorithm

Revealing how a biological network is organized to realize its function is one of the main topics in systems biology. The functional backbone network, defined as the primary structure of the biological network, is of great importance in maintaining the main function of the biological network. We propose a new algorithm, the tinker algorithm, to determine this core structure and apply it in the cell-cycle system. With this algorithm, the backbone network of the cell-cycle network can be determined accurately and efficiently in various models such as the Boolean model, stochastic model, and ordinary differential equation model. Results show that our algorithm is more efficient than that used in the previous research. We hope this method can be put into practical use in relevant future studies.

nlin.AO↗

Next-to-leading-order QCD corrections to the yields and polarisations of J/Psi and Upsilon directly produced in association with a Z boson at the LHC

We update the study of the production of direct J/Psi in association with a Z boson at the Next-to-Leading Order (NLO) in alpha_s by evaluating both the yield differential in P_T and the J/Psi polarisation in the QCD-based Colour-Singlet Model (CSM). Contrary to an earlier claim, QCD corrections at small and mid P_T are small if one assumes that the factorisation and the renormalisation scales are commensurate with the Z boson mass. As it can be anticipated, the t-channel gluon-exchange (t-CGE) topologies start to be dominant only for P_T > mZ/2. The polarisation pattern is not altered by the QCD corrections. This is thus far the first quarkonium-production process where this is observed in the CSM. Along the same lines, our predictions for direct Upsilon+Z are also given.

hep-ph↗