SearcharxivSearch

arXiv subjects

Shilin Ma

Publications and source records attributed to Shilin Ma.

15 recordsLinked to original sources

Uncertainty-Aware World Model for Aerial Image-Goal Navigation

Aerial image-goal navigation requires an unmanned aerial vehicle (UAV) to reach a target location specified by a goal image. Existing world-model-based methods rank candidate trajectories using predicted futures, but typically rely on only one or a few point predictions, which is inadequate for large-scale outdoor environments with substantial future-state uncertainty. To address this limitation, we propose the Uncertainty-Aware Navigation World Model (UA-NWM), an efficient latent world model for aerial image-goal navigation, which formulates trajectory scoring as conditional out-of-distribution detection. UA-NWM represents plausible futures with an uncertainty subspace and decomposes the prediction--goal discrepancy into uncertainty-explainable and unexplainable components. Only the unexplainable residual is used for scoring, enabling robust selection without multiple future samples. Extensive experiments demonstrate that UA-NWM consistently outperforms existing navigation world models while maintaining low inference latency. Real-world UAV experiments further validate its practical applicability. Project page: https://duryi.github.io/UA-NWM-Project-Page

cs.CV

AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning

Chain-of-thought (CoT) reasoning can improve performance on difficult video questions but often wastes decoding tokens on simple ones. We study whether a video multimodal large language model can adapt its reasoning effort to each question. We propose AdaThinkV, an adaptive framework for video reasoning that learns whether to reason explicitly without offline difficulty labels, manually tuned confidence thresholds, or an external router. During reinforcement learning, AdaThinkV samples matched rollouts in explicit reasoning and direct answering modes for each prompt. ThinkGain estimates the prompt-level utility of explicit reasoning by balancing its accuracy gain against additional response length, providing supervision for both conditional response generation and autonomous mode selection. For difficult prompts, limited rollout exploration can yield groups in which every response is unsuccessful and accuracy rewards show little variation, providing insufficient signal for learning. We therefore introduce Variance Recovery Policy Optimization (VRPO), which retains and progressively expands these groups to recover informative signals from prompts that are difficult yet solvable. At inference, AdaThinkV selects a response mode and generates the response in a single autoregressive sequence. Across a unified suite of video reasoning evaluations, AdaThinkV achieves a mean accuracy of 40.79 with an average of 257.20 output tokens, outperforming the strongest evaluated adaptive baseline by 2.98 points while using 22.7% fewer tokens. Project page: https://trilarflagz.github.io/AdaThinkV/

cs.CV

MorphoQuant: Modality-Aware Quantization for Omni-modal Large Language Models

Conventional Post-Training Quantization (PTQ) methods struggle with 4-bit Omni-modal Large Language Models (OLLMs) due to the extreme distribution heterogeneity and disparate outlier patterns across modalities. To address this, we propose MorphoQuant, a modality-aware PTQ framework engineered to preserve cross-modal morphology and mitigate outlier loss. Specifically, we introduce Distribution-Aware Bias Compensation (DABC), which selectively absorbs long-tailed outliers into channel-wise biases. This mechanism safeguards outlier magnitudes while maintaining high-precision discretization for dense inliers, thereby preserving accurate discretization across diverse modal distribution. Complementing this, we propose Morphology-Directed Quantization Function Optimization (MDQFO) to co-optimize the quantization grid with the bias mask, ensuring fine-grained alignment across modalities. Extensive evaluations on Qwen2.5-Omni across benchmarks like MMMU and Video-MME demonstrate our approach's superiority. Notably, our W4A4 model achieves 76.63% on ScienceQA, significantly outperforming SOTA W4A4 methods and surprisingly surpassing the W4A16 baseline, which fully demonstrates the exceptional accuracy-efficiency trade-off of our framework.

cs.CV

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation

Real-time inference of vision-language-action (VLA) models is essential for robotic control. While visual token pruning has shown strong potential for accelerating inference, most existing methods mainly base pruning decisions on shallow-layer cues and risk discarding visual information required by deep layers. To address this issue, we propose SAFE-Pruner, a plug-and-play pruning framework that incorporates attention cues of future layers into pruning decisions. Specifically, we identify semantic attention consistency, the tendency that VLA models concentrate their attention probability mass on the same semantic entity across control timesteps. Based on this observation, we design a forward-looking strategy to forecast the token saliency in deep layers, which prevents the premature removal of critical tokens and leads to more stable acceleration. We further introduce a reference timestep refresh strategy that triggers updates upon attention shifts, thereby improving forecasting accuracy and pruning reliability. Extensive experiments across diverse evaluation settings demonstrate that our method achieves up to 1.89x speedup with a minimal degradation in success rate of less than 1.5%, while outperforming state-of-the-art methods by up to 1.9%.

cs.CV

VARestorer: One-Step VAR Distillation for Real-World Image Super-Resolution

Recent advancements in visual autoregressive models (VAR) have demonstrated their effectiveness in image generation, highlighting their potential for real-world image super-resolution (Real-ISR). However, adapting VAR for ISR presents critical challenges. The next-scale prediction mechanism, constrained by causal attention, fails to fully exploit global low-quality (LQ) context, resulting in blurry and inconsistent high-quality (HQ) outputs. Additionally, error accumulation in the iterative prediction severely degrades coherence in ISR task. To address these issues, we propose VARestorer, a simple yet effective distillation framework that transforms a pre-trained text-to-image VAR model into a one-step ISR model. By leveraging distribution matching, our method eliminates the need for iterative refinement, significantly reducing error propagation and inference time. Furthermore, we introduce pyramid image conditioning with cross-scale attention, which enables bidirectional scale-wise interactions and fully utilizes the input image information while adapting to the autoregressive mechanism. This prevents later LQ tokens from being overlooked in the transformer. By fine-tuning only 1.2\% of the model parameters through parameter-efficient adapters, our method maintains the expressive power of the original VAR model while significantly enhancing efficiency. Extensive experiments show that VARestorer achieves state-of-the-art performance with 72.32 MUSIQ and 0.7669 CLIPIQA on DIV2K dataset, while accelerating inference by 10 times compared to conventional VAR inference.

cs.CV

The Stability Limit of Prepotentials for Hurwitz-Frobenius Manifolds: An Infinite-Dimensional Approach

The stability of prepotential derivatives for Frobenius manifolds associated with A_N and D_N singularities has been utilized to construct (2+1)-dimensional dispersionless integrable hierarchies. Although the generalization of this construction to genus-zero Hurwitz-Frobenius manifolds was shown to yield the genus-zero Whitham hierarchy, a direct geometric explanation of this correspondence has been lacking. In this note, we provide a direct proof of this identification within the framework of infinite-dimensional Frobenius manifolds. We demonstrate that the stability of prepotentials is an intrinsic property of the tau-structure of the Whitham hierarchy. Furthermore, we extend this identification to the hierarchies arising from the stability of solutions to the open WDVV equations with the extensions of the Whitham hierarchy.

math-ph

Solutions to Open WDVV Equations for the Universal Whitham Hierarchy

In this paper, we construct a pair of solutions to the open WDVV equations associated with the infinite-dimensional Frobenius manifolds that underlie the genus-zero universal Whitham hierarchy, and for the resulting flat F-manifolds, we explicitly construct their principal hierarchies. We further demonstrate that this construction is compatible with finite-dimensional reductions, yielding solutions for Frobenius manifolds associated with general rational superpotentials and those subject to a $\mathbb{Z}_{2}$-symmetry reduction. In particular, the polynomial solutions derived by Basalaev and Buryak via open Saito theory for A- and D-type singularities are recovered as special cases.

math-ph

FADE: Frequency-Aware Diffusion Model Factorization for Video Editing

Recent advancements in diffusion frameworks have significantly enhanced video editing, achieving high fidelity and strong alignment with textual prompts. However, conventional approaches using image diffusion models fall short in handling video dynamics, particularly for challenging temporal edits like motion adjustments. While current video diffusion models produce high-quality results, adapting them for efficient editing remains difficult due to the heavy computational demands that prevent the direct application of previous image editing techniques. To overcome these limitations, we introduce FADE, a training-free yet highly effective video editing approach that fully leverages the inherent priors from pre-trained video diffusion models via frequency-aware factorization. Rather than simply using these models, we first analyze the attention patterns within the video model to reveal how video priors are distributed across different components. Building on these insights, we propose a factorization strategy to optimize each component's specialized role. Furthermore, we devise spectrum-guided modulation to refine the sampling trajectory with frequency domain cues, preventing information leakage and supporting efficient, versatile edits while preserving the basic spatial and temporal structure. Extensive experiments on real-world videos demonstrate that our method consistently delivers high-quality, realistic and temporally coherent editing results both qualitatively and quantitatively. Code is available at https://github.com/EternalEvan/FADE .

cs.CV

Penalty Adversarial Network (PAN): A neural network-based method to solve PDE-constrained optimal control problems

In this work, we introduce a novel strategy for tackling constrained optimization problems through a modified penalty method. Conventional penalty methods convert constrained problems into unconstrained ones by incorporating constraints into the loss function via a penalty term. However, selecting an optimal penalty parameter remains challenging; an improper choice, whether excessively high or low, can significantly impede the discovery of the true solution. This challenge is particularly evident when training neural networks for constrained optimization, where tuning parameters can become an extensive and laborious task. To overcome these issues, we propose an adversarial approach that redefines the conventional penalty method by simultaneously considering two competing penalty problems--a technique we term the penalty adversarial problem. Within linear settings, our method not only ensures the fulfillment of constraints but also guarantees solvability, leading to more precise solutions compared to traditional approaches. We further reveal that our method effectively performs an automatic adjustment of penalty parameters by leveraging the relationship between the objective and loss functions, thereby obviating the need for manual parameter tuning. Additionally, we extend this adversarial framework to develop a neural network-based solution for optimal control problems governed by linear or nonlinear partial differential equations. We demonstrate the efficacy of this innovative approach through a series of numerical examples.

math.OC

Principal hierarchy of Frobenius manifolds with rational and trigonometric superpotentials

In this paper, we construct the principal hierarchies for Frobenius manifolds with rational and trigonometric superpotentials, as well as their almost dualities. We demonstrate that in both cases, submanifolds with even superpotentials form natural Frobenius submanifolds, and their principal hierarchies can be obtained as restrictions of the principal hierarchies for the original Frobenius manifolds. Furthermore, we introduce a natural rank-1 extension for each of these Frobenius manifolds, providing solutions to the associated open WDVV equations. The principal hierarchy for each extension is also explicitly constructed.

math-ph

Explicit bounds for small prime nonresidues

Let $χ$ be a Dirichlet character modulo a prime~$p$. We give explicit upper bounds on $q_1 0$ such that $q_n\leq Cp^{\frac{1}{4}}(\log p)^{\frac{n+1}{2}}$ whenever $n\leq n_0$ and $p\geq p_0$.

math.NT

On the number of primes for which a polynomial is Eisenstein

Previously Heyman and Shparlinski gave an asymptotic formula with error term for the number of Eisenstein polynomials of fixed degree and bounded height. Let $ψ(f)$ denote the number of primes for which a polynomial $f$ is Eisenstein. We give expressions for the mean and variance of the function $ψ$ for each fixed degree, where the polynomials are ordered according to their height.

math.NT