SearcharxivSearch

arXiv subjects

Kimia Shaban

Publications and source records attributed to Kimia Shaban.

5 recordsLinked to original sources

SEGA: Spectral-Energy Guided Attention for Resolution Extrapolation in Diffusion Transformers

Diffusion transformers (DiTs) have emerged as a dominant architecture for text-to-image generation, yet their performance drops when generating at resolutions beyond their training range. Existing training-free approaches mitigate this by modifying inference-time attention behavior, often through Rotary Position Embeddings (RoPE) extrapolation combined with attention scaling. However, these strategies apply a uniform and content-agnostic scaling across RoPE components with distinct frequency characteristics, inducing a trade-off between preserving global structure and recovering fine detail. We introduce SEGA, a training-free method that dynamically scales attention across RoPE components according to the latent's spatial-frequency structure at each denoising step. This adaptive scaling improves both structural coherence and fine-detail fidelity. Experiments show that SEGA consistently improves high-resolution synthesis across multiple target resolutions, outperforming state-of-the-art training-free baselines.

cs.CV

Sizes of witnesses in Covtree

Given a set $\Gamma$ of $k$ unlabelled posets, each of size $n$, we say that a poset $Q$ is a \emph{witness} to $\Gamma$ if $\Gamma$ is the set of downsets of size $n$ of $Q$. We say that $Q$ is a \emph{minimal witness} if it does not contain a proper downset that is itself a witness to $\Gamma$. Motivated by the causal set approach to quantum gravity, we study the upper bound on the size of minimal witnesses as a function of $n$ and $k$. We show that there is no linear upper bound of the form $n+k+c$ for any constant $c$. We introduce the \emph{exchange graph of downsets} as a new tool to study this scenario, and use it to show that all minimal witnesses $Q$ satisfy the bound $|Q|\leq nk-n$, and that when $k=3$ there is at least one minimal witness $Q$ that satisfies the bound $|Q|\leq \frac{3}{2}(n+1)$.

math.CO

When Does RL Help Medical VLMs? Disentangling Vision, SFT, and RL Gains

Reinforcement learning (RL) is increasingly used to post-train medical Vision-Language Models (VLMs), yet it remains unclear whether RL improves medical visual reasoning or mainly sharpens behaviors already induced by supervised fine-tuning (SFT). We present a controlled study that disentangles these effects along three axes: vision, SFT, and RL. Using MedMNIST as a multi-modality testbed, we probe visual perception by benchmarking VLM vision towers against vision-only baselines, quantify reasoning support and sampling efficiency via Accuracy@1 versus Pass@K, and evaluate when RL closes the support gap and how gains transfer across modalities. We find that RL is most effective when the model already has non-trivial support (high Pass@K): it primarily sharpens the output distribution, improving Acc@1 and sampling efficiency, while SFT expands support and makes RL effective. Based on these findings, we propose a boundary-aware recipe and instantiate it by RL post-training an OctoMed-initialized model on a small, balanced subset of PMC multiple-choice VQA, achieving strong average performance across six medical VQA benchmarks.

cs.CV

Statistics and asymptotics of subdivergence-free Feynman integrals in $\phi^4$ theory

Recent algorithmic improvements have made it possible to evaluate subdivergence-free (=primitive=skeleton) Feynman integrals in $\phi^4$ theory numerically up to 18 loops. By now, all such integrals up to 13 loops and several hundred thousand at higher loop order have been computed. This data enables a statistical analysis of the typical behaviour of Feynman integrals at large loop order. We find that the average value grows exponentially, but the observed growth rate is accurately described by its leading asymptotics only upwards of 25 loops. This is in contrast with the $N$-dependence of the $ON(N)$-symmetric $\phi^4$ theory, which is close to its large-order asymptotics already around 10 loops. Secondly, the distribution of integrals has a largely continuous inner part but a few extreme outliers. This makes uniform random sampling inefficient. We find that the value of the integral is correlated with many features of the graph, which can be used for importance sampling. With a naive test implementation we obtained an approximately 1000-fold speedup compared with uniform sampling. This suggests that in future work, Feynman amplitudes at large loop order might be computed numerically with statistical methods, rather than through enumerating and evaluating every individual integral.

hep-th

Predicting Feynman periods in $\phi^4$-theory

We present efficient data-driven approaches to predict Feynman periods in $\phi^4$-theory from properties of the underlying Feynman graphs. We find that the numbers of cuts and cycles determines the period to approximately 2% accuracy. Hepp bound and Martin invariant allow to predict the period with accuracy much better than 1%. In most cases, the period is a multi-linear function of the parameters in question. Besides classical correlation analysis, we also investigate the usefulness of machine-learning algorithms to predict the period. When sufficiently many properties of the graph are used, the period can be predicted with better than 0.05% relative accuracy. We use one of the constructed prediction models for weighted Monte-Carlo sampling of Feynman graphs, and compute the primitive contribution to the beta function of $\phi^4$-theory at $L\in \left \lbrace 13, 14, 15, 16 \right \rbrace $ loops. Our results confirm the previously known numerical estimates of the primitive beta function and improve their accuracy. Compared to uniform random sampling of graphs, our new algorithm reaches 35-fold higher accuracy in fixed runtime, or requires 1000-fold less runtime to reach a given accuracy. The data set of all periods computed for this work, combined with a previous data set, is made publicly available. Besides the physical application, it could serve as a benchmark for graph-based machine learning algorithms.

hep-th