SearcharxivSearch

arXiv subjects

Fuxiang Yang

Publications and source records attributed to Fuxiang Yang.

7 recordsLinked to original sources

Higher syzygy bundles and the Eisenbud-Huneke-Ulrich conjecture

We study ideals in a polynomial ring with partially linear (virtual) resolutions. We establish an effective bound beyond which their powers coincide with a power of the maximal ideal, proving a slightly weaker version of the Eisenbud-Huneke-Ulrich conjecture for a more general class of ideals. We also obtain an upper bound on their Castelnuovo-Mumford regularity, extending Macaulay's bound and a theorem of Eisenbud-Huneke-Ulrich. We introduce higher syzygy bundles, which generalize the classical Green-Lazarsfeld syzygy bundle. The key ingredient is the relationship between the syzygies of partially linear ideals and the sheaf cohomology of higher syzygy bundles.

math.AC

Chain of World: World Model Thinking in Latent Motion

Vision-Language-Action (VLA) models are a promising path toward embodied intelligence, yet they often overlook the predictive and temporal-causal structure underlying visual dynamics. World-model VLAs address this by predicting future frames, but waste capacity reconstructing redundant backgrounds. Latent-action VLAs encode frame-to-frame transitions compactly, but lack temporally continuous dynamic modeling and world knowledge. To overcome these limitations, we introduce CoWVLA (Chain-of-World VLA), a new "Chain of World" paradigm that unifies world-model temporal reasoning with a disentangled latent motion representation. First, a pretrained video VAE serves as a latent motion extractor, explicitly factorizing video segments into structure and motion latents. Then, during pre-training, the VLA learns from an instruction and an initial frame to infer a continuous latent motion chain and predict the segment's terminal frame. Finally, during co-fine-tuning, this latent dynamic is aligned with discrete action prediction by jointly modeling sparse keyframes and action sequences in a unified autoregressive decoder. This design preserves the world-model benefits of temporal reasoning and world knowledge while retaining the compactness and interpretability of latent actions, enabling efficient visuomotor learning. Extensive experiments on robotic simulation benchmarks show that CoWVLA outperforms existing world-model and latent-action approaches and achieves moderate computational efficiency, highlighting its potential as a more effective VLA pretraining paradigm. The project website can be found at https://fx-hit.github.io/cowvla-io.

cs.CV

FOCA: Frequency-Oriented Cross-Domain Forgery Detection, Localization and Explanation via Multi-Modal Large Language Model

Advances in image tampering techniques, particularly generative models, pose significant challenges to media verification, digital forensics, and public trust. Existing image forgery detection and localization (IFDL) methods suffer from two key limitations: over-reliance on semantic content while neglecting textural cues, and limited interpretability of subtle low-level tampering traces. To address these issues, we propose FOCA, a multimodal large language model-based framework that integrates discriminative features from both the RGB spatial and frequency domains via a cross-attention fusion module. This design enables accurate forgery detection and localization while providing explicit, human-interpretable cross-domain explanations. We further introduce FSE-Set, a large-scale dataset with diverse authentic and tampered images, pixel-level masks, and dual-domain annotations. Extensive experiments show that FOCA outperforms state-of-the-art methods in detection performance and interpretability across both spatial and frequency domains.

cs.CV

Powers of binary forms and derived Hermite reciprocity

For $a,b \ge 1$, Hilbert found in 1886 a collection of polynomial equations that cut out set-theoretically the variety X parametrizing a-th powers of binary forms of degree b. We determine the ideal of all polynomials vanishing on X, showing that it is generated in degree b+1 and that it has a linear minimal free resolution. We do this by generalizing results of Abdesselam and Chipalkatti on an analogue of the Foulkes--Howe map and by establishing a derived analogue of the classical Hermite reciprocity theorem for complexes of ${\rm SL}_2$-representations. In our investigation, we are led to the ideal generated by the subrepresentation ${\rm Sym}^{ab}({\Bbb C}^2) \subset {\rm Sym}^a({\rm Sym}^b {\Bbb C}^2)$. We determine its Castelnuovo--Mumford regularity in general and the minimal free resolution for small values of b.

math.AC

Global-Local Aware Scene Text Editing

Scene Text Editing (STE) involves replacing text in a scene image with new target text while preserving both the original text style and background texture. Existing methods suffer from two major challenges: inconsistency and length-insensitivity. They often fail to maintain coherence between the edited local patch and the surrounding area, and they struggle to handle significant differences in text length before and after editing. To tackle these challenges, we propose an end-to-end framework called Global-Local Aware Scene Text Editing (GLASTE), which simultaneously incorporates high-level global contextual information along with delicate local features. Specifically, we design a global-local combination structure, joint global and local losses, and enhance text image features to ensure consistency in text style within local patches while maintaining harmony between local and global areas. Additionally, we express the text style as a vector independent of the image size, which can be transferred to target text images of various sizes. We use an affine fusion to fill target text images into the editing patch while maintaining their aspect ratio unchanged. Extensive experiments on real-world datasets validate that our GLASTE model outperforms previous methods in both quantitative metrics and qualitative results and effectively mitigates the two challenges.

cs.CV

Scene Style Text Editing

In this work, we propose a task called "Scene Style Text Editing (SSTE)", changing the text content as well as the text style of the source image while keeping the original text scene. Existing methods neglect to fine-grained adjust the style of the foreground text, such as its rotation angle, color, and font type. To tackle this task, we propose a quadruple framework named "QuadNet" to embed and adjust foreground text styles in the latent feature space. Specifically, QuadNet consists of four parts, namely background inpainting, style encoder, content encoder, and fusion generator. The background inpainting erases the source text content and recovers the appropriate background with a highly authentic texture. The style encoder extracts the style embedding of the foreground text. The content encoder provides target text representations in the latent feature space to implement the content edits. The fusion generator combines the information yielded from the mentioned parts and generates the rendered text images. Practically, our method is capable of performing promisingly on real-world datasets with merely string-level annotation. To the best of our knowledge, our work is the first to finely manipulate the foreground text content and style by deeply semantic editing in the latent feature space. Extensive experiments demonstrate that QuadNet has the ability to generate photo-realistic foreground text and avoid source text shadows in real-world scenes when editing text content.

cs.CV

Asymptotic Behavior of Differential Powers

In this paper, we study the differential power operation on ideals. We begin with a focus on monomial ideals in characteristic 0 and find a class of ideals whose differential powers are eventually principal. We also study the containment problem between ordinary and differential powers of ideals, in analogy to earlier work comparing ordinary and symbolic powers of ideals. We further define a possible closure operation on ideals, called the differential closure, in analogy with integral closure and tight closure. We show that this closure operation agrees with taking the radical of an ideal if and only if the ambient ring is a simple $D$-module.

math.AC