SearcharxivSearch

arXiv subjects

Xiaoyan Yang

Publications and source records attributed to Xiaoyan Yang.

At least 19 recordsLinked to original sources

Finitistic dimensions in triangulated categories with a compact silting generator

This paper studies finitistic dimensions in triangulated categories with a compact silting generator. We unify several notions of (big) finitistic dimension appearing in the literature, show that they are essentially equivalent, and relate them to the abelian heart when the generator is tilting. This yields a new categorical perspective on the finitistic dimension conjecture. Furthermore, we establish explicit inequalities for (big) finitistic and global dimensions under recollements, demonstrating that finiteness in the middle category is equivalent to finiteness in the outer categories. These results generalize classical ring-theoretic theorems and apply to triangular matrix rings, exact contexts and trivial extensions.

math.CT

RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing

Recent advances in video editing have been largely driven by large-scale instruction-based datasets. However, existing datasets still suffer from two critical limitations. First, target videos are commonly produced by automatic editing models, which may introduce visible artifacts and unreliable supervision signals. Second, most public datasets rely primarily on textual instructions, while lacking visual references that are crucial for precise, identity-preserving, and controllable editing. To address these limitations, we introduce RefVideo-6M, a large-scale reference-guided editing dataset containing 5 million video editing samples and 1 million image editing samples. To ensure reliable supervision, our dataset uses a construction pipeline that treats artifact-free real videos as editing targets and generates quality-filtered input conditions with multiple editing experts. In addition, it provides approximately 6 million visual references, covering diverse reference types and editing scenarios, thereby enabling models to learn fine-grained visual correspondence beyond text-only instructions. Based on RefVideo-6M, we further train a reference-guided video editing model, Ref-MoT, to evaluate the effectiveness and scalability of the proposed dataset. Extensive experiments demonstrate that RefVideo-6M provides substantially more reliable supervision than existing datasets and enables the training of powerful editing models with improved visual quality, controllability, and reference consistency. The open-source dataset is available at https://huggingface.co/datasets/RefVideo6M/RefVideo6M.

cs.CV

LATTE: Forecasting Peer Anchored Preference Trajectories for Personalized LLM Generation

Personalized generation with frozen large language models requires a conditioning signal that is both compact and current. Existing personalization methods typically retrieve or summarize user histories in text, or compress them into static latent profiles and soft prompts. These approaches are efficient, but they treat a user's past behavior as an aggregate profile and therefore mix stable identity, recent drift, and item content in the same representation. We propose LAtent Trajectory Tracking and Extrapolation (LATTE), a framework that represents personalization as forecasting a peer anchored relative preference state. For each historical session, LATTE subtracts a time masked baseline formed from comparable users who responded to the same item, producing a state that measures how the target user differs from peers under a shared item context. A lightweight sequence predictor then forecasts the next state in this trajectory, and a State to Token Bridge injects the forecast into a frozen instruction tuned LLM through a single anchored soft token. We provide a latent factor analysis showing when peer anchoring cancels shared item variation and why temporal forecasting trades off stale averages against noisy recent states. Experiments on Amazon Reviews 2023 and MemoryCD show that LATTE consistently outperforms retrieval, summary memory, static latent profiles, difference aware latent profiles, and soft prompt compression baselines. On Amazon Reviews 2023, LATTE improves average ROUGE-L from 0.219 for a static latent profile and 0.245 for the strongest added latent compression baseline to 0.259. Additional pairwise comparisons and diagnostic analyses suggest that the improvement is mainly due to forecasting user-specific trajectory information, rather than merely adding a soft prompt interface.

cs.CL

Chains of model structures arising from cotorsion pairs on extriangulated categories

The main aim of this paper is to study chains of model structures arising from cotorsion pairs in extriangulated categories. Starting with a hereditary Hovey triple, we construct further hereditary Hovey triples whose homotopy categories are equivalent under suitable completeness assumptions, thereby refining results due to El Maaouy and Shao-Wang-Zhang. As an application, we consider objects of finite Gorenstein injective dimension with respect to a proper class of $\mathbb{E}$-triangles. Under mild set-theoretic assumptions, we obtain a chain of model structures whose homotopy categories are all triangulated equivalent to a common stable category. This recovers known results for Gorenstein injective modules and yields new examples in the derived category of a ring when the proper class is given by cohomological ghost triangles.

math.RT

Quillen equivalence for chain homotopy categories induced by balanced pairs

For a balanced pair $(\mathcal{X},\mathcal{Y})$ in an abelian category, we investigate when the chain homotopy categories ${\bf K}(\mathcal{X})$ and ${\bf K}(\mathcal{Y})$ are triangulated equivalent. To this end, we realize these chain homotopy categories as homotopy categories of certain model categories and give conditions that ensure the existence of a Quillen equivalence between the model categories in question. We further give applications to cotorsion triples, Gorenstein projective and Gorenstein injective modules, as well as pure projective and pure injective objects.

math.RT

Quasi-projective dimensions of complexes over rings

Quasi-projective dimension of modules over associative rings is generalized in this paper to the one of complexes of modules. Basic properties of this dimension are established, including a comparison result with projective dimension and a derived Auslander-Buchsbaum formula for complexes of finite quasi-projective dimension. Several sufficient conditions are provided for a commutative noetherian local ring to be a complete intersection under the assumption that each finitely generated module has finite quasi-projective dimension. This provides some positive answers to an open question on quasi-projective dimension proposed by Gheibi-Jorgensen-Takahashi. Moreover, the behavior of quasi-projective dimension under taking the quotient of a commutative ring modulo a regular sequence is investigated, and some partial results toward the change-of-rings question on quasi-projective dimension are given.

math.RA

ABC-GS: Alignment-Based Controllable Style Transfer for 3D Gaussian Splatting

3D scene stylization approaches based on Neural Radiance Fields (NeRF) achieve promising results by optimizing with Nearest Neighbor Feature Matching (NNFM) loss. However, NNFM loss does not consider global style information. In addition, the implicit representation of NeRF limits their fine-grained control over the resulting scenes. In this paper, we introduce ABC-GS, a novel framework based on 3D Gaussian Splatting to achieve high-quality 3D style transfer. To this end, a controllable matching stage is designed to achieve precise alignment between scene content and style features through segmentation masks. Moreover, a style transfer loss function based on feature alignment is proposed to ensure that the outcomes of style transfer accurately reflect the global style of the reference image. Furthermore, the original geometric information of the scene is preserved with the depth loss and Gaussian regularization terms. Extensive experiments show that our ABC-GS provides controllability of style transfer and achieves stylization results that are more faithfully aligned with the global style of the chosen artistic reference. Our homepage is available at https://vpx-ecnu.github.io/ABC-GS-website.

cs.CV

Grade and Cohen-Macaulayness for DG-modules

We establish an inequality relating the projective dimension of a DG-module in $\mathrm{D}^\mathrm{b}_\mathrm{f}(A)$ to its grade and introduce the concept of perfect DG-modules as a natural generalization of perfect modules. It is proved that a DG-module $M$ over a local Cohen-Macaulay DG-ring with constant amplitude is Cohen-Macaulay if and only if $M$ is perfect and $\mathrm{amp}M \leq \mathrm{amp}\mathrm{R}Γ_{\bar{\mathfrak{m}}}(M)$. An affirmative answer is provided to Conjecture 2.11 of Yoshida [J. Pure Appl. Algebra 123 (1998) 313--326]. We also study the grade of DG-modules with finite injective dimension and examine the preservation of Cohen-Macaulayness under tensor products.

math.AC

Tele-Omni: a Unified Multimodal Framework for Video Generation and Editing

Recent advances in diffusion-based video generation have substantially improved visual fidelity and temporal coherence. However, most existing approaches remain task-specific and rely primarily on textual instructions, limiting their ability to handle multimodal inputs, contextual references, and diverse video generation and editing scenarios within a unified framework. Moreover, many video editing methods depend on carefully engineered pipelines tailored to individual operations, which hinders scalability and composability. In this paper, we propose Tele-Omni, a unified multimodal framework for video generation and editing that follows multimodal instructions, including text, images, and reference videos, within a single model. Tele-Omni leverages pretrained multimodal large language models to parse heterogeneous instructions and infer structured generation or editing intents, while diffusion-based generators perform high-quality video synthesis conditioned on these structured signals. To enable joint training across heterogeneous video tasks, we introduce a task-aware data processing pipeline that unifies multimodal inputs into a structured instruction format while preserving task-specific constraints. Tele-Omni supports a wide range of video-centric tasks, including text-to-video generation, image-to-video generation, first-last-frame video generation, in-context video generation, and in-context video editing. By decoupling instruction parsing from video synthesis and combining it with task-aware data design, Tele-Omni achieves flexible multimodal control while maintaining strong temporal coherence and visual consistency. Experimental results demonstrate that Tele-Omni achieves competitive performance across multiple tasks.

cs.CV

Point2Insert: Video Object Insertion via Sparse Point Guidance

This paper introduces Point2Insert, a sparse-point-based framework for flexible and user-friendly object insertion in videos, motivated by the growing popularity of accurate, low-effort object placement. Existing approaches face two major challenges: mask-based insertion methods require labor-intensive mask annotations, while instruction-based methods struggle to place objects at precise locations. Point2Insert addresses these issues by requiring only a small number of sparse points instead of dense masks, eliminating the need for tedious mask drawing. Specifically, it supports both positive and negative points to indicate regions that are suitable or unsuitable for insertion, enabling fine-grained spatial control over object locations. The training of Point2Insert consists of two stages. In Stage 1, we train an insertion model that generates objects in given regions conditioned on either sparse-point prompts or a binary mask. In Stage 2, we further train the model on paired videos synthesized by an object removal model, adapting it to video insertion. Moreover, motivated by the higher insertion success rate of mask-guided editing, we leverage a mask-guided insertion model as a teacher to distill reliable insertion behavior into the point-guided model. Extensive experiments demonstrate that Point2Insert consistently outperforms strong baselines and even surpasses models with $\times$10 more parameters.

cs.CV

TeleStyle: Content-Preserving Style Transfer in Images and Videos

Content-preserving style transfer, generating stylized outputs based on content and style references, remains a significant challenge for Diffusion Transformers (DiTs) due to the inherent entanglement of content and style features in their internal representations. In this technical report, we present TeleStyle, a lightweight yet effective model for both image and video stylization. Built upon Qwen-Image-Edit, TeleStyle leverages the base model's robust capabilities in content preservation and style customization. To facilitate effective training, we curated a high-quality dataset of distinct specific styles and further synthesized triplets using thousands of diverse, in-the-wild style categories. We introduce a Curriculum Continual Learning framework to train TeleStyle on this hybrid dataset of clean (curated) and noisy (synthetic) triplets. This approach enables the model to generalize to unseen styles without compromising precise content fidelity. Additionally, we introduce a video-to-video stylization module to enhance temporal consistency and visual quality. TeleStyle achieves state-of-the-art performance across three core evaluation metrics: style similarity, content consistency, and aesthetic quality. Code and pre-trained models are available at https://github.com/Tele-AI/TeleStyle

cs.CV

TeleWorld: Towards Dynamic Multimodal Synthesis with a 4D World Model

World models aim to endow AI systems with the ability to represent, generate, and interact with dynamic environments in a coherent and temporally consistent manner. While recent video generation models have demonstrated impressive visual quality, they remain limited in real-time interaction, long-horizon consistency, and persistent memory of dynamic scenes, hindering their evolution into practical world models. In this report, we present TeleWorld, a real-time multimodal 4D world modeling framework that unifies video generation, dynamic scene reconstruction, and long-term world memory within a closed-loop system. TeleWorld introduces a novel generation-reconstruction-guidance paradigm, where generated video streams are continuously reconstructed into a dynamic 4D spatio-temporal representation, which in turn guides subsequent generation to maintain spatial, temporal, and physical consistency. To support long-horizon generation with low latency, we employ an autoregressive diffusion-based video model enhanced with Macro-from-Micro Planning (MMPL)--a hierarchical planning method that reduces error accumulation from frame-level to segment-level-alongside efficient Distribution Matching Distillation (DMD), enabling real-time synthesis under practical computational budgets. Our approach achieves seamless integration of dynamic object modeling and static scene representation within a unified 4D framework, advancing world models toward practical, interactive, and computationally accessible systems. Extensive experiments demonstrate that TeleWorld achieves strong performance in both static and dynamic world understanding, long-term consistency, and real-time generation efficiency, positioning it as a practical step toward interactive, memory-enabled world models for multimodal generation and embodied intelligence.

cs.CV

Induced complete hereditary cotorsion pairs in D(R) with respect to Cartan-Eilenberg exact sequences

Given a complete hereditary cotorsion pair (A,B) in ModR, we construct a complete hereditary cotorsion pair in the derived category D(R) of unbounded complexes with respect to the proper class ξ of cohomologically ghost triangles induced by the Cartan-Eilenberg exact sequences. More specifically, we prove that, each of the classes of projectively coresolved ξ-Gflat complexes PGF(ξ), ξ-Gflat complexes GF(ξ), ξ-Ginjective complexes GI(ξ), ξ-Gprojective complexes GP(ξ) (the last when R is virtually Gorenstein), forms one half of a complete hereditary cotorsion pair in D(R) with respect to ξ. Moreover, various homological dimensions offer additional way to obtain such cotorsion pairs in D(R) with respect to ξ.

math.CT

UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation

We present UniModel, a unified generative model that jointly supports visual understanding and visual generation within a single pixel-to-pixel diffusion framework. Our goal is to achieve unification along three axes: the model, the tasks, and the representations. At the representation level, we eliminate modality discrepancies by mapping both text and images into a shared visual space: textual prompts are rendered as painted text images on a clean canvas, and all inputs and outputs are treated purely as RGB pixels. This yields a fully vision-native formulation of multimodal learning. At the task level, a broad range of vision-language problems are cast as pixel-to-pixel transformations in this visual space. For understanding tasks, the model takes an RGB image and produces a painted text image that visually encodes the semantic prediction. For generation tasks, painted text images serve as visual conditions that guide realistic and semantically aligned image synthesis. Captioning and text-to-image generation thus become different directions of the same underlying visual translation process. At the model level, we instantiate a single Unified Diffusion Transformer trained with rectified flow in pixel space. A shared backbone jointly learns bidirectional mappings between natural images and painted text images, with lightweight task embeddings to specify the desired direction. Experiments on text-to-image synthesis and image-to-text understanding demonstrate strong cross-modal alignment and emergent controllability such as cycle-consistent image-caption-image loops. Our initial exploration suggests that unifying model, tasks, and representations in a single visual space is a promising paradigm for general-purpose multimodal intelligence.

cs.CV

Improved Bounds for the s-multiplicity

Let (RmR), (SmS) and (TmT) be Noetherian local rings sharing the same residue eld k and prime characteristic p > 0. We establish some formulas relating the h-function and s-multiplicity of the ber product R T S in terms of the h-functions and s-multiplicities of R, T and S. Furthermore, we derive formulas that connect the h-function and s-multiplicity of the idealization ring R M to the corresponding invariants of R and M, where M is a nitely generated R-module. As applications of these results, we derive new estimates for the Taylor-Miller question and the Watanabe-Yoshida conjecture concerning s-multiplicity.

math.AC

Models for chain homotopy category of relative acyclic complexes

Let $(\mathcal{X}, \mathcal{Y})$ be a balanced pair in an abelian category $\mathcal{A}$. Denote by ${\bf K}_{\mathcal{E}\text{-}{\rm ac}}(\mathcal{X})$ the chain homotopy category of right $\mathcal{X}$-acyclic complexes with all items in $\mathcal{X}$, and dually by ${\bf K}_{\mathcal{E}\text{-}{\rm ac}}(\mathcal{Y})$ the chain homotopy category of left $\mathcal{Y}$-acyclic complexes with all items in $\mathcal{Y}$. We establish realizations of ${\bf K}_{\mathcal{E}\text{-}{\rm ac}}(\mathcal{X})$ and ${\bf K}_{\mathcal{E}\text{-}{\rm ac}}(\mathcal{Y})$ as homotopy categories of model categories under mild conditions. Consequently, we obtain relative versions of recollements of Krause and Neeman-Murfet. We further give applications to Gorenstein projective and Gorenstein injective modules.

math.RT

G-dimensions for DG-modules over commutative DG-rings

We define and study a notion of G-dimension for DG-modules over a non-positively graded commutative noetherian DG-ring $A$. Some criteria for the finiteness of the G-dimension of a DG-module are given by applying a DG-version of projective resolution introduced by Minamoto [Israel J. Math. 245 (2021) 409-454]. Moreover, it is proved that the finiteness of G-dimension characterizes the local Gorenstein property of $A$. Applications go in three directions. The first is to establish the connection between G-dimensions and the little finitistic dimensions of &\mathcal{A}&. The second is to characterize Cohen-Macaulay and Gorenstein DG-rings by the relations between the class of maximal local-Cohen-Macaulay DG-modules and a special G-class of DG-modules. The third is to extend the classical Buchwtweiz-Happel Theorem and its inverse from commutative noetherian local rings to the setting of commutative noetherian local DG-rings.Our method is somewhat different from classical commutative ring.

math.AC

DeepCell: Self-Supervised Multiview Fusion for Circuit Representation Learning

We introduce DeepCell, a novel circuit representation learning framework that effectively integrates multiview information from both And-Inverter Graphs (AIGs) and Post-Mapping (PM) netlists. At its core, DeepCell employs a self-supervised Mask Circuit Modeling (MCM) strategy, inspired by masked language modeling, to fuse complementary circuit representations from different design stages into unified and rich embeddings. To our knowledge, DeepCell is the first framework explicitly designed for PM netlist representation learning, setting new benchmarks in both predictive accuracy and reconstruction quality. We demonstrate the practical efficacy of DeepCell by applying it to critical EDA tasks such as functional Engineering Change Orders (ECO) and technology mapping. Extensive experimental results show that DeepCell significantly surpasses state-of-the-art open-source EDA tools in efficiency and performance.

cs.LG