SearcharxivSearch

arXiv subjects

Viacheslav Ivanov

Publications and source records attributed to Viacheslav Ivanov.

3 recordsLinked to original sources

Prompt2Effect: Training-Free Image-to-Video Model Specialization via LoRA Generation

While personalizing Image-to-Video (I2V) diffusion models with specific visual effects is increasingly demanded for high-end generation, current practice requires training a separate Low-Rank Adaptation (LoRA) module for each effect, incurring substantial data curation and iterative optimization costs that hinder interactive control. We present Prompt2Effect, a weight-driven hypernetwork that amortizes per-effect training by directly synthesizing effect-specific LoRA weights in a single forward pass. Unlike prior hypernetworks that regress adapter weights purely from semantics, Prompt2Effect is explicitly conditioned on the frozen base model weights, grounding prediction in the structural geometry of each layer. Furthermore, instead of predicting raw LoRA matrices, we introduce an SVD-canonicalized parameterization that resolves factorization ambiguity and stabilizes large-scale synthesis. Extensive experiments demonstrate that Prompt2Effect achieves on-par or superior video quality and effect alignment compared to conventional LoRA fine-tuning, while reducing the computational cost from 56 GPU training hours to 3.3 seconds of hypernetwork inference. When used as initialization for subsequent fine-tuning, our predicted weights further improve final performance and accelerate optimization by approximately 10x.

cs.CV

AlcheMinT: Fine-grained Temporal Control for Multi-Reference Consistent Video Generation

Recent advances in subject-driven video generation with large diffusion models have enabled personalized content synthesis conditioned on user-provided subjects. However, existing methods lack fine-grained temporal control over subject appearance and disappearance, which are essential for applications such as compositional video synthesis, storyboarding, and controllable animation. We propose AlcheMinT, a unified framework that introduces explicit timestamps conditioning for subject-driven video generation. Our approach introduces a novel positional encoding mechanism that unlocks the encoding of temporal intervals, associated in our case with subject identities, while seamlessly integrating with the pretrained video generation model positional embeddings. Additionally, we incorporate subject-descriptive text tokens to strengthen binding between visual identity and video captions, mitigating ambiguity during generation. Through token-wise concatenation, AlcheMinT avoids any additional cross-attention modules and incurs negligible parameter overhead. We establish a benchmark evaluating multiple subject identity preservation, video fidelity, and temporal adherence. Experimental results demonstrate that AlcheMinT achieves visual quality matching state-of-the-art video personalization methods, while, for the first time, enabling precise temporal control over multi-subject generation within videos. Project page is at https://snap-research.github.io/Video-AlcheMinT

cs.CV

Yangians and degenerate affine Schur algebras

Drinfeld's degenerate affine analog of Schur-Weyl duality relates representations of the degenerate affine Hecke algebra $AH_r$ to representations of the Yangian $Y_n$. One way to understand the construction is to introduce an intermediate algebra $AS(n,r)$, the degenerate affine Schur algebra, which appears both as the endomorphism algebra of an induced tensor space over $AH_r$, and as the image of a homomorphism $D_{n,r}:Y_n \rightarrow AS(n,r)$. In this paper, we describe $D_{n,r}$ using a diagrammatic calculus. Then we use a theorem of Drinfeld to compute $\ker D_{n,r}$ when $n > r$, thereby giving a presentation of $AS(n,r)$ in these cases. We formulate a conjecture in the remaining cases. Finally, we apply results of Arakawa to develop some of the representation theory of $AS(n,r)$.

math.RT