SearcharxivSearch

arXiv subjects

Xingzhong Xu

Publications and source records attributed to Xingzhong Xu.

17 recordsLinked to original sources

FinixDoc: Rethinking Financial Document Parsing Beyond Saturated Benchmarks

Financial document parsing requires accuracy, structural consistency, and verifiability that current benchmarks often fail to reflect. We present FinixDoc, an end-to-end agentic parsing system for real-world financial documents, with FinixDoc-VL, a 4B-scale vision-language model built on Qwen3-VL-4B, as its core parser. To characterize the gap between benchmark and deployment performance, we introduce a Document Parsing Capability Matrix organized along two practical axes: visual quality and document scale. Guided by this matrix, FinixDoc-VL is trained with a domain-adapted recipe combining homoglyph-aware contrastive learning and multi-stage reinforcement learning with composite domain-specific rewards. To better leverage our accumulated advantage in low-quality financial-document data and support large-scale, high-quality data production, we further build a human-in-the-loop Data Factory pipeline with confidence-aware expert review. For evaluation, we construct FinixDocBench, a financial-domain evaluation suite covering digital-native, camera-captured, ultra-large-page, and internal-workflow scenarios, with a compliance-reviewed subset released alongside this technical report. On its main subsets, FinixDoc-VL achieves the highest overall score (81.43) among evaluated baselines, outperforming the next-best open-source model by 5.13 points, with the largest gains on internal financial workflows (FinixInner: 84.08 vs. 78.73).

cs.AI

SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration

Recent advances in Tool-Integrated Large Language Models have made web search a core capability of information-seeking agents. However, as interaction histories grow, agents increasingly struggle to track task progress. When search attempts fail to yield useful evidence, current single- and multi-agent systems can become trapped in repetitive loops, wasting search budgets and ultimately compromising the quality and completeness of the final output. We introduce SearchOS, a system-level multi-agent framework that turns fragile, implicit search progress into explicit, persistent, and shared state. First, we formulate open-domain information seeking as relational schema completion with grounded citations, where agents discover entities, populate attributes across linked tables, and anchor each value to source evidence. Then we design Search-Oriented Context Management (SOCM), which externalizes the evolving state into Frontier Task, an Evidence Graph, a Coverage Map, and Failure Memory. Built on SOCM, SearchOS applies a pipeline-parallel scheduling mechanism that overlaps the execution of sub-agents and continuously refills freed slots with tasks targeting unresolved coverage gaps to improve utilization and throughput. To schedule and control the execution of search agents, SearchOS introduces a Search Tool Middleware Harness that intercepts model and tool interactions to record grounded evidence and react to stalls or budget exhaustion, and provides a reusable hierarchical skill system comprising strategy and access skills to augment the agents' search process and avoid repeating failed search patterns across runs. On WideSearch and GISA, SearchOS leads all metrics among the evaluated single- and multi-agent baselines, paving the way toward robust information-seeking collaboration.

cs.AI

UCPO: Uncertainty-Aware Policy Optimization

The key to building trustworthy large language models (LLMs) lies in endowing them with inherent uncertainty expression capabilities, thereby mitigating overconfident errors in high-stakes applications. However, existing RL paradigms such as GRPO often suffer from Advantage Bias due to binary decision spaces and static uncertainty rewards, inducing either excessive conservatism or overconfidence. To tackle this challenge, this paper unveils the root causes of reward hacking and overconfidence in current RL paradigms incorporating uncertainty-based rewards, based on which we propose the UnCertainty-Aware Policy Optimization (UCPO) framework. UCPO employs Ternary Advantage Decoupling to separate and independently normalize deterministic and uncertain rollouts, thereby eliminating advantage bias. Furthermore, a Dynamic Uncertainty Reward Adjustment mechanism adapts uncertainty weights in real-time according to model evolution and instance difficulty. Experimental results in mathematical reasoning and general tasks demonstrate that UCPO effectively resolves the reward imbalance, significantly improving the reliability of the model beyond their knowledge boundaries.

cs.AI

NGRPO: Negative-enhanced Group Relative Policy Optimization

RLVR has enhanced the reasoning capabilities of Large Language Models (LLMs) across various tasks. However, GRPO, a representative RLVR algorithm, suffers from a critical limitation: when all responses within a group are either entirely correct or entirely incorrect, the model fails to learn from these homogeneous responses. This is particularly problematic for homogeneously incorrect groups, where GRPO's advantage function yields a value of zero, leading to null gradients and the loss of valuable learning signals. To overcome this issue, we propose NGRPO (Negative-enhanced Group Relative Policy Optimization), an algorithm designed to convert homogeneous errors into robust learning signals. First, NGRPO introduces Advantage Calibration. This mechanism hypothesizes the existence of a virtual maximum-reward sample during advantage calculation, thereby altering the mean and variance of rewards within a group and ensuring that the advantages for homogeneously incorrect samples are no longer zero. Second, NGRPO employs Asymmetric Clipping, which relaxes the update magnitude for positive samples while imposing stricter constraints on that of negative samples. This serves to stabilize the exploration pressure introduced by the advantage calibration. Our experiments on Qwen2.5-Math-7B demonstrate that NGRPO significantly outperforms baselines such as PPO, GRPO, DAPO, and PSR-NSR on mathematical benchmarks including MATH500, AMC23, and AIME2025. These results validate NGRPO's ability to learn from homogeneous errors, leading to stable and substantial improvements in mathematical reasoning. Our code is available at https://github.com/nangongrui-ngr/NGRPO.

cs.LG

LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL

Enhancing reasoning in Large Multimodal Models (LMMs) faces unique challenges from the complex interplay between visual perception and logical reasoning, particularly in compact 3B-parameter architectures where architectural constraints limit reasoning capacity and modality alignment. While rule-based reinforcement learning (RL) excels in text-only domains, its multimodal extension confronts two critical barriers: (1) data limitations due to ambiguous answers and scarce complex reasoning examples, and (2) degraded foundational reasoning induced by multimodal pretraining. To address these challenges, we propose \textbf{LMM-R1}, a two-stage framework adapting rule-based RL for multimodal reasoning through \textbf{Foundational Reasoning Enhancement (FRE)} followed by \textbf{Multimodal Generalization Training (MGT)}. The FRE stage first strengthens reasoning abilities using text-only data with rule-based RL, then the MGT stage generalizes these reasoning capabilities to multimodal domains. Experiments on Qwen2.5-VL-Instruct-3B demonstrate that LMM-R1 achieves 4.83\% and 4.5\% average improvements over baselines in multimodal and text-only benchmarks, respectively, with a 3.63\% gain in complex Football Game tasks. These results validate that text-based reasoning enhancement enables effective multimodal generalization, offering a data-efficient paradigm that bypasses costly high-quality multimodal training data.

cs.CL

Transitive fusion systems over a class of finite p-groups

Let $p$ be an odd prime and $S$ a nonabelian finite $p$-group. In [9, 10], they proposed the following conjecture: if $\mathcal{F}$ be a transitive fusion system over a finite $p$-group $S$, then $S$ is either extraspecial of order $p^{3}$ or elementary abelian. In this note, we use an easy method to prove that this conjecture holds when the $p$-rank of $S$ is 2.

math.GR

Some results on a question of M. Newman on isomorphic subgroups of solvable groups

In this paper, we focus on a question of M. Newman on isomorphic subgroups of solvable groups. We get a reduction theorem of this question: for each prime q, assume that this question holds for every characteristic q-groups, then this question holds for every finite solvable groups. Using this reduction theorem, we get some partial answers about this question.

math.GR

Elementary abelian subgroups in some special p-groups

Let $P$ be a finite $p$-group and $p$ be an odd prime. Let $\mathcal{A}_p(P)_{\geq2}$ be a poset consisting of elementary abelian subgroups of rank at least 2. If the derived subgroup $P'\cong C_p\times C_p$, then the spheres occurring in $\mathcal{A}_p(P)_{\geq2}$ all have the same dimension.

math.GR

A note on Oliver's p-group conjecture

Let $S$ be a $p$-group for an odd prime $p$, Oliver proposed the conjecture that the Thompson subgroup $J(S)$ is always contained in the Oliver subgroup $\mathfrak{X}(S)$. That means he conjectured that $|J(S)\mathfrak{X}(S):\mathfrak{X}(S)|=1$. Let $\mathfrak{X}_1(S)$ be a subgroup of $S$ such that $\mathfrak{X}_1(S)/\mathfrak{X}(S)$ is the center of $S/\mathfrak{X}(S)$. In this short note, we prove that $J(S)\leq \mathfrak{X}(S)$ if and only if $J(S)\leq \mathfrak{X}_1(S)$. As an easy application, we prove that $|J(S)\mathfrak{X}(S):\mathfrak{X}(S)|\neq p$.

math.GR

A relation between $m_{G, N}$ and the Euler characteristic of the nerve space of some class poset of $G$

Let $G$ be a finite group and $N\unlhd G$ with $|G: N|=p$ for some prime $p$. In this note, to compute $m_{G,N}$ directly, we construct a class poset $\mathfrak{T}_{C}(G)$ of $G$ for some cyclic subgroup $C$. And we find a relation between $m_{G,N}$ and the Euler characteristic of the nerve space $|N(\mathfrak{T}_{C}(G))|$ (see the Theorem 1.3). As an application, we compute $m_{S_5, A_5}=0$ directly, and get $S_5$ is a $B$-group.

math.GR

Bouc's conjecture on $B$-groups

Bouc proposed the following conjecture: a finite group $G$ is nilpotent if and only if its largest quotient $B$-group $β(G)$ is nilpotent. And he has prove that this conjecture holds when $G$ is solvable. In this paper, we consider the case when $G$ is not solvable. Let $S$ be a nonabelian simple group except the Chevalley groups $A_{n}(q)$, $D_{n}(q)$, $E_{6}(q)$, and $^2A_{n}(q)$, if there exists only one factor of $G$ which is isomorphic to $S$, then $β(G)$ is not solvable, of course, is not nilpotent. That means we prove the conjecture in these cases.

math.GR