SearcharxivSearch

arXiv subjects

Guoqing Wang

Publications and source records attributed to Guoqing Wang.

At least 19 recordsLinked to original sources

GeoAAC: Geometry-Based Adaptive Action Chunking from Denoising Trajectories in VLA Policies

Action chunking is widely used for action generation and execution in Vision-Language-Action (VLA) policies, yet existing approaches commonly use a fixed action horizon. During a rollout, different task stages may require different levels of action continuity, control precision, and closed-loop feedback, making a fixed horizon unable to accommodate changing control requirements. We propose \textbf{GeoAAC}, a geometry-based adaptive action chunking method for flow-based VLA policies that adjusts the action horizon according to the reliability of the current action prediction. We show that the geometry of Flow Matching denoising trajectories provides process-level information for characterizing prediction reliability, with geometric variation across action prefixes remaining positively correlated with predictive uncertainty. GeoAAC uses this prefix-wise geometry to construct a horizon-wise geometric profile and adaptively determine the action horizon from a single generation without additional training. Experiments with GR00T N1.5 and π0.5 on LIBERO, LIBERO-Pro, RoboCasa365, and real-world manipulation tasks show consistent improvements over fixed-action-horizon baselines and existing adaptive methods, including up to 8.7 percentage points in simulation and an increase in average real-world success rate from 53.3\% to 74.4\%.

cs.RO

On a classical zero-sum invariant II: Disproof of a long-standing conjecture

For a nontrivial finite abelian group $G$, let $ν(G)$ be the smallest integer $\ell$ such that every zero-sum free sequence $T$ over $G$ of length at least $\ell$ has the following property: all nonzero elements of $G$ that do not occur as a subsequence sum of $T$ lie in a proper coset of some subgroup of $G$. It is easy to check that $\mathsf d (G)-1 \le ν(G) \le \mathsf d (G)$, where $\mathsf d (G)$ is the small Davenport constant of $G$. A conjecture by Gao from the year 2000 stated that equality should always hold at the lower bound. This conjecture has since been confirmed for many families of groups (including all p-groups and groups of rank at most two). In the current note, we disprove the conjecture.

math.CO

Control incompatibility in multiparameter quantum metrology

In practical applications like quantum sensing and quantum imaging, there is often a necessity to estimate multiple parameters simultaneously. Although the ultimate precision limits for single-parameter estimation are well established, the precision limit of multi-parameter estimation is much less understood. This is primarily due to the inherent incompatibility of the optimal strategies for the estimation of different parameters, particularly those pertaining to optimal control.In this study, we tackle the critical issue of control incompatibility in multi-parameter estimation by presenting explicit cases that expose this challenge. Our research not only pioneers the exploration of control incompatibility but also highlights its pivotal role in the field. Furthermore, our work offers valuable insights into how to minimize trade-offs induced by control incompatibility and enhance precision. This paves the way for future investigations into control strategies that enable optimal estimation of multiple parameters that are incompatible.

quant-ph

The Gao-Zhuang conjecture for the Heisenberg group over $\mathbb{F}_p$

Let $G$ be a finite nonabelian group. The small Davenport constant $\mathsf d(G)$ of $G$ is the largest integer $\ell$ such that there exists a product-one free sequence over $G$ of length $\ell$, while the Gao constant $E(G)$ of $G$ is the least integer $\ell$ such that every sequence over $G$ of length at least $\ell$ contains a product-one subsequence of length exactly $|G|$. A long-standing conjecture of Gao and Zhuang \cite{ZG2005} asserts that $E(G)=\mathsf d(G)+|G|$ for every finite nonabelian group $G$. Let $p$ be an odd prime and let $H_{p^3}=\operatorname{UT}_3(\mathbb F_p)$ be the finite Heisenberg group over $\mathbb F_p$. Godara and Sarkar proved the Gao-Zhuang equality for $H_{27}=\operatorname{UT}_3(\mathbb F_3)$ and asked whether the same equality holds for $H_{p^3}$ for every odd prime $p$. Recently, Volkmann proved that $\mathsf d(H_{p^3})=3p-3$. In this paper, we determine the Gao constant of $H_{p^3}$ and prove that $E(H_{p^3})=\mathsf d(H_{p^3})+|H_{p^3}|=p^3+3p-3$. Together with the known abelian and cyclic-index cases, this completes the verification of the Gao-Zhuang equality for all groups of order $p^3$, for every prime $p$.

math.CO

Spin squeezing by geometric focusing in vacuum Rabi oscillations

We show that vacuum Rabi oscillations can directly generate spin squeezing through geometric focusing on the Bloch sphere. Starting from a coherent spin state resonantly coupled to a cavity initially in the vacuum state, quantum fluctuations are focused by the curvature of the Bloch sphere as the collective spin approaches the atomic ground state, producing squeezing transverse to the direction of motion. The squeezing timescale is set by the collective Rabi frequency $t_s\sim 1/(g\sqrt{N})$. The optimal Wineland squeezing parameter scales as $ξ_{\rm opt}^2\propto N^{-1/3}$, which is an outcome of the competition between the geometric focusing effects and the cavity-field vacuum fluctuations. The squeezing remains robust against realistic dissipation. In the end, an application example of $^{171}$Yb is briefly discussed to show the feasibility of our protocol.

quant-ph

Group-Shared Low-Rank Approximation for Mobile-Efficient Pointwise Convolutions in Large-Kernel CNNs

Large-kernel Convolutional Neural Networks (CNNs) deliver remarkable performance in vision tasks by significantly expanding receptive fields, yet their quadratic parameter growth critically impedes storage-efficient edge deployment. While existing efficient architectures adopt parameter-efficient depthwise separable convolution backbones that leverage techniques like low-rank approximation and weight sharing to compress depthwise convolutions, we identify a critical oversight: pointwise convolutions dominate parameter volume (>87% in models like RepLKNet-31B) and constitute the primary deployment bottleneck on resource-constrained edge devices. This results in prohibitive storage costs and severe memory-loading constraints on resource-limited devices (e.g., smartphones with 4-12 GB Random Access Memory (RAM)). To overcome this, we propose Channel Group-Shared (CGS) low-rank approximation, a novel Singular Value Decomposition (SVD)-based parameter-sharing strategy. CGS constructs a structured low-rank paradigm isomorphic to SVD decomposition, comprising shared (high-parameter-cost) down/up-projection matrices across channel groups within a layer and channel-group-specific (low-parameter-cost) scalable diagonal matrices. This group-sharing design achieves significant parameter reduction. Extensive experiments demonstrate that large-kernel CNNs (RepLKNet, ConvNeXt, SLaK) enhanced with CGS strike an empirically favorable balance between competitive performance and substantially reduced storage costs. Crucially, by alleviating storage constraints, reducing memory bandwidth pressure during loading, and minimizing model loading latency, CGS enables the feasible deployment of pre-trained large-kernel CNN models on edge devices, thereby bridging the gap between high-performance vision models and practical edge deployment.

cs.LG

The equality between the Erdős-Ginzburg-Ziv constant and the short product-one constant for finite nonabelian groups

Let $G$ be a finite group, and let $\exp(G)$ denote its exponent. The Erdős-Ginzburg-Ziv constant $s(G)$ is the least integer forcing a product-one subsequence of length $\exp(G)$, while the short product-one constant $η(G)$ is the least integer forcing a nonempty product-one subsequence of length at most $\exp(G)$. The natural nonabelian extension of a conjecture [W. Gao, \emph{On zero-sum subsequences of restricted size II}, Discrete Math. 2003] on the Erdős-Ginzburg-Ziv constant in finite abelian groups predicts that $s(G)=η(G)+\exp(G)-1.$ We confirm this equality for every finite nonabelian group $G$ having a cyclic subgroup of index $p$, where $p$ is the smallest prime divisor of $|G|$. As further consequences, we determine all generalized Erdős-Ginzburg-Ziv constants $s_{m\exp(G)}(G)$ for this family of groups.

math.CO

High-efficiency loading of 2,400 Ytterbium atoms in optical tweezer arrays

Neutral atom arrays have emerged as a powerful platform for quantum computation, simulation, and metrology.Among them, alkaline-earth-like atoms exhibit distinct advantages, including long coherence time, high-fidelity Rydberg gates, and erasure correction for efficient quantum error correction. However, their scalability has lagged behind that of the alkali atoms. Here, we report 2400 ytterbium-174 atoms trapped in an optical tweezer array with enhanced loading efficiency of 83.5(1)\% via blue-detuned light-assisted collisions. We develop a quantitative model of the collision dynamics and find good agreement between the calculated inelastic collision rates and the experimentally measured loading efficiencies.Notably, the loading efficiency is largely maintained for array sizes ranging from dozens to thousands, exhibiting excellent scalability. We further demonstrate that the enhancement exists robustly across a range of interatomic potentials, suggesting its utility for other atomic species. To establish the capability of the $^{174}$Yb arrays toward universal quantum computation, we propose to encode the qubit in the ground-clock state manifold and estimate a 99.9\% two-qubit gate fidelity with experimentally feasible parameters. Our work advances the prospects for realizing large-scale quantum computers using alkaline-earth-like atoms.

quant-ph

MENTOR: A Metacognition-Driven Self-Evolution Framework for Uncovering and Mitigating Implicit Domain Risks in LLMs

Ensuring the safety of Large Language Models (LLMs) is critical for real-world deployment. However, current safety measures often fail to address implicit, domain-specific risks. To investigate this gap, we introduce a dataset of 3,000 annotated queries spanning education, finance, and management. Evaluations across 14 leading LLMs reveal a concerning vulnerability: an average jailbreak success rate of 57.8\%. In response, we propose MENTOR, a metacognition-driven self-evolution framework. MENTOR performs metacognitive self-assessment, using strategies such as perspective-taking and consequential reasoning to uncover latent model misalignments. MENTOR couples single-pass rule-guided inference for routine requests with a selectively invoked metacognitive evolution cycle that revises residual unsafe responses, distills successful corrections into a dynamic rule graph, and compiles validated rules into activation-level steering signals for future inference. Experiments demonstrate that MENTOR substantially reduces attack success rates across all tested domains and outperforms existing safety alignment methods. The code and dataset for MENTOR are available at: https://anonymous.4open.science/r/MENTOR-Evo.

cs.AI

Hypernetwork-Parameterized Spatially Adaptive Neural Operators for PDE Learning

Spatially heterogeneous partial differential equations (PDEs) exhibit location-dependent dynamics arising from variations in geometry and physical coefficients. Existing neural operators improve localized modeling through multiscale features, attention mechanisms, or domain decomposition, yet their update rules often remain spatially shared. Hypernetwork-based methods adapt parameters across PDE instances but typically generate only one global parameterization per instance. Consequently, shared operators may underfit boundaries and high-gradient regions, with these localized errors accumulating during autoregressive rollout. We propose a spatially adaptive neural operator (SANO), which replaces this spatially shared parameterization with a spatially continuous field of location-dependent operator parameters. SANO uses Fourier-encoded coordinates and a coordinate-conditioned hypernetwork to generate spatial operator-conditioning codes at sampling points. A Hyper-Neural Element (HNE) mechanism interpolates these codes within local subregions, coupling neighboring operators while allowing their update rules to vary across space, and partition-of-unity weights assemble the overlapping local predictions. Experiments on one-, two-, and three-dimensional PDEs and two perforated-domain elliptic benchmarks show that SANO consistently outperforms competitive neural-operator, hypernetwork-based, and physics-informed baselines.

cs.LG

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding

Large vision-language models (LVLMs) have achieved significant progress in video understanding, yet understanding long videos remains challenging due to the large number of visual tokens and limited context windows. Visual sampling provides a practical solution by selecting an informative subset of frames. However, existing methods typically either rely on relevance-aware sampling, leading to redundant frame selection and insufficient temporal coverage, or adopt a fixed sampling strategy regardless of query type. In this paper, we propose VisualRouter, a training-free and plug-and-play framework for query-grounded visual sampling. VisualRouter first classifies each query as either global or local and then applies the corresponding sampling strategy. For global queries, it employs a relevance-coverage hybrid strategy that preserves temporal coverage while retaining query-relevant visual evidence. For local queries, it adopts an event-aware frame selection strategy that performs event partitioning, segment-level frame allocation, and intra-event frame selection, jointly balancing relevance, coverage, and diversity with a limited number of input frames. Experiments show that VisualRouter consistently improves multiple LVLMs over uniform sampling, achieving gains of 5.2%, 7.7%, and 11.6% on Video-MME, LongVideoBench, and MLVU with Qwen2.5-VL-7B, and outperforming existing training-free visual sampling methods under the same setting.

cs.CV

PixelPilot: Scalable Vision-Language-Action Models for End-to-End Autonomous Driving

Vision-Language-Action Models (VLAs), which leverage the advanced reasoning capabilities of Vision-Language Models (VLMs), show promising generalization in complex autonomous driving scenarios. Existing VLAs typically predict and optimize 3D trajectories from 2D images. While intuitive, this 2D-to-3D prediction is inherently entangled with camera parameters, leading to limited data scalability across heterogeneous driving datasets. Moreover, directly optimizing in 3D space induces severe convergence to trivial solutions, where VLAs rely on ego-status rather than visual scene understanding. To address these issues, we propose PixelPilot, a novel VLA featuring a decoupled planning and lifting paradigm. In the planning phase, PixelPilot reformulates scene understanding and trajectory prediction as sensor-agnostic 2D-to-2D tasks in the image plane, thereby facilitating scalable training across diverse datasets. The planned 2D trajectories are then deterministically lifted to 3D only during inference, ensuring the full exploitation of visual cues and generalization across different vehicles. To realize this paradigm, we propose a knowledge-instilled policy learning strategy that applies dense, intermediate rewards via Group Relative Policy Optimization (GRPO) to enforce a rigorous causal chain from visual perception to spatial planning. Extensive experiments demonstrate that PixelPilot achieves state-of-the-art performance in both open-loop and closed-loop settings, validating its superior scalability and visual reasoning capabilities.

cs.CV

Targeted Structure Completion for Sparse-View 3D Reconstruction in Autonomous Driving

Reconstructing 3D scene structures from sparse, low-overlap observations remains a fundamental challenge in autonomous driving. Recent state-of-the-art frameworks achieve promising results by incorporating voxel-based Gaussians, but incur substantial computational redundancy due to a uniform volumetric processing strategy. To bridge the gap between the efficiency of pixel-based Gaussian methods and the structural completeness of voxel-based Gaussian approaches, we propose FocusGS, a simple yet effective framework that shifts the paradigm from global densification to targeted structural completion. Our central insight is that structural completion should be decoupled from deterministic regions, with computation concentrated exclusively on areas exhibiting geometric ambiguity. Specifically, FocusGS addresses the localization challenge by deriving a 3D Geometric Ambiguity Manifold to accurately isolate localized areas prone to occlusion and high geometric uncertainty. To overcome the subsequent manifold completion challenge, we design a lightweight targeted structure completion module that selectively instantiates and optimizes continuous Gaussian queries strictly within this unstructured, sparse topological subspace. Extensive experiments demonstrate that FocusGS achieves a superior efficiency-quality trade-off, advancing state-of-the-art performance on driving-centric benchmarks while naturally reducing the total number of Gaussians by ~74% and decreasing rendering time by ~34%.

cs.CV

SparseOcc++: Geometry-Aware Sparse Latent Representation for Semantic Occupancy Prediction

Vision-based 3D semantic occupancy prediction is essential for autonomous driving, yet dense voxel representations waste computation on largely empty space, while BEV and TPV projections compromise fine-grained 3D structure. Fully sparse representations offer an attractive alternative, but existing methods, including SparseOcc, entangle scene completion with semantic prediction by indiscriminately propagating high-dimensional features into empty regions and applying voxel-wise classification. This creates excessive activations, computational overhead, and geometric ambiguity. We present SparseOcc++, a geometry-aware sparse framework that explicitly decouples scene completion from semantic segmentation. SparseOcc++ reformulates completion as signed-distance regression on sparse anchor voxels through a scene completion field (SCF). To model complex outdoor geometry robustly, it combines orthogonal decomposition with discretized distance learning. A geometry-guided propagation module then converts the SCF into a complete volumetric scene and restricts semantic segmentation to geometrically verified regions. Experiments establish new state of the art: SparseOcc++ improves IoU by 2.3 points and is 3.9x faster than SparseOcc on nuScenes, while achieving a 5.9x speedup over OccFormer on SemanticKITTI.

cs.CV

WPG-MoE: Weak-Prior-Guided Dense Mixture-of-Experts for User-Level Social Media Depression Detection

Online social media posts provide scalable signals for early depression screening, and recent studies mainly improve pre-classification evidence through risk-post selection, symptom grounding, and clinically informed feature construction. However, these screening-stage designs often leave final decisions to a single detector, overlooking how users heterogeneously express depressive risk after screening. A monolithic classifier must average across heterogeneous users, which may dilute localized evidence and cause misclassification, especially for non-self-disclosing users. To address this issue, we propose WPG-MoE, a weak-prior-guided dense mixture-of-experts framework built on a shared large language model (LLM) backbone. WPG-MoE derives user-level weak semantic priors to softly route users to experts matched to different evidence layouts. We formulate this process as learning using privileged information (LUPI): rich LLM-extracted structured evidence guides training-time routing, while inference retains only Patient Health Questionnaire-9 (PHQ-9) template screening and the deployable backbone. Experiments on Chinese and English datasets show that WPG-MoE outperforms strong baselines with interpretable routing behavior.

cs.CL

The universal zero-sum invariant and weighted zero-sum for infinite abelian groups II

Let $G$ be an abelian group, and let $\mathcal F (G)$ be the free commutative monoid with basis $G$, and $\mathcal A (G)$ the set consisting of all minimal zero-sum subsequences over $G$. For any subset $Ω\subset \mathcal F (G)$, we define the universal zero-sum invariant ${\mathsf d}_Ω(G)$ as the minimal positive integer $\ell$ such that every sequence $T$ over $G$ of length $\ell$ contains a subsequence lying in $Ω$. The classical Davenport constant ${\rm D}(G)$ for $G$ can also be written as ${\mathsf d}_{\mathcal A (G)}(G)$. We give a complete classification of all finite abelian groups for which $\mathcal A(G)$ is a minimal set to represent the Davenport constant. We also investigate the weighted Davenport constant over abelian groups (which may be infinite). Let $F$ and $G$ be abelian groups, and let $Ψ\subseteq \mathrm{Hom}(F,G)$ denote a weight set. We reinterpret the weighted Davenport constant $D_Ψ(G)$ in terms of coverings of Cartesian powers $F^n$ by kernels of induced homomorphisms arising from tuples in $Ψ^n$; these homomorphisms are naturally linked to coproducts in the category of abelian groups. This motivates the notion of kernel-cover compactness, a property characterizing when such kernel coverings admit finite subcovers. We establish a correspondence between weighted zero-sum invariants and kernel-cover structures, where the bound $D_Ψ(G)\le n$ is equivalent to a canonical kernel-cover property on $F^n$. We further study finite reduction phenomena for infinite weight sets and provide sufficient conditions ensuring uniform kernel-cover compactness. The present work constitutes a follow-up to [G. Wang, Comm. Algebra, 2025].

math.CO

Deep Reinforcement Learning for Individual Atomic Control and Cooling

Real-time feedback control of quantum systems is often limited by partial observations, nonlinear dynamics and measurement noise, which make accurate model-based controllers difficult to design. Here we show that deep reinforcement learning can cool the motion of a single neutral atom coupled to a high-finesse optical cavity using only the continuously monitored cavity transmission. We first train the controller in simulation and then transfer it to the experiment, where online fine-tuning adapts it to unmodeled experimental dynamics. The learned policy damps the atom's motion in real time and achieves a cooling time constant of 388 +/- 14 microseconds, corresponding to only two motional periods in the trap. It also outperforms a standard linear differentiator controller in cooling speed while maintaining comparable atom retention over a broad range of operating conditions. These results establish reinforcement learning as a practical strategy for feedback control in quantum-limited experiments where compact analytical models are incomplete.

quant-ph

Multimodal Mathematical Reasoning with Diverse Solving Perspective

Recent progress in large-scale reinforcement learning (RL) has notably enhanced the reasoning capabilities of large language models (LLMs), especially in mathematical domains. However, current multimodal LLMs (MLLMs) for mathematical reasoning often rely on one-to-one image-text pairs and single-solution supervision, overlooking the diversity of valid reasoning perspectives and internal reflections. In this work, we introduce MathV-DP, a novel dataset that captures multiple diverse solution trajectories for each image-question pair, fostering richer reasoning supervision. We further propose Qwen-VL-DP, a model built upon Qwen-VL, fine-tuned with supervised learning and enhanced via group relative policy optimization (GRPO), a rule-based RL approach that integrates correctness discrimination and diversity-aware reward functions. Our method emphasizes learning from varied reasoning perspectives and distinguishing between correct yet distinct solutions. Extensive experiments on the MathVista's minitest and Math-V benchmarks demonstrate that Qwen-VL-DP significantly outperforms prior base MLLMs in both accuracy and generative diversity, highlighting the importance of incorporating diverse perspectives and reflective reasoning in multimodal mathematical reasoning.

cs.CL