SearcharxivSearch

arXiv subjects

Yuqing Shi

Publications and source records attributed to Yuqing Shi.

5 recordsLinked to original sources

Do We Really Need External Tools to Mitigate Hallucinations? SIRA: Shared-Prefix Internal Reconstruction of Attribution

Large vision-language models (LVLMs) often hallucinate when language priors dominate weak or ambiguous visual evidence. Existing contrastive decoding methods mitigate this problem by comparing predictions from the original image with those from externally perturbed visual inputs, but such references can introduce off-manifold artifacts and require costly extra forward passes. We propose SIRA, a training-free internal contrastive decoding framework that constructs a counterfactual reference inside the same LVLM by exploiting the staged information flow of multimodal transformers. Instead of removing visual information from the input, SIRA first lets image and text tokens interact through a shared prefix, forming an aligned multimodal state that preserves prompt interpretation, decoding history, positional structure, and early visual grounding. It then forks a counterfactual branch in later transformer layers, where attention to image-token positions is masked. This branch retains the shared multimodal context but lacks continued access to fine-grained visual evidence, yielding a language-prior-dominated internal reference for token-level contrast. During decoding, SIRA suppresses tokens that remain strong without late visual access and favors predictions whose advantage depends on the full visual pathway. Experiments on POPE, CHAIR, and AMBER with Qwen2.5-VL and LLaVA-v1.5 show that SIRA consistently reduces hallucinations while preserving descriptive coverage and incurring lower overhead than two-pass contrastive decoding. SIRA requires no training, external verifier, or perturbed input, and applies to open-weight LVLMs with white-box inference access.

cs.CV

On periodic homotopy and homology equivalences of spaces

There are at least two ways to approach the homotopy theory of spaces `at chromatic height $n$': one may localize with respect to $T(n)$-homology or with respect to $v_n$-periodic homotopy groups. It was already observed by Bousfield that these two options yield rather different results. We build on his work to prove precise comparison results between the two notions. A crucial concept is a more robust notion of $T(n)$-equivalence that we call `parametric $T(n)$-equivalence': this is a map of spaces that induces an equivalence on $\infty$-categories of local systems valued in $T(n)$-local spectra. Our results are sharpest in the case of infinite loop spaces, where amongst other things we prove a $T(n)$-local version of a result of Kuhn on the Morava $K$-theory of the Whitehead tower. As a corollary of our results we also produce a formula for the $L_n^f$-localization of an infinite loop space $\Omega^\infty E$ of a spectrum satisfying $L_{n-1}^f E \simeq 0$.

math.AT

Universal property of the Bousfield--Kuhn functor

We present a universal property of the Bousfield--Kuhn functor $\operatorname{\Phi}_h$ of height $h$, for every positive natural number $h$. This result is achieved by proving that the costabilisation of the $\infty$-category of $v_h$-periodic homotopy types is equivalent to the $\infty$-category of $\operatorname{T}(h)$-local spectra. A key component in our proofs is the spectral Lie algebra model for $v_h$-periodic homotopy types (see arXiv:1803.06325): We relate the costabilisation of the $\infty$-category of spectral Lie algebras with the costabilisations of the $\infty$-category of non-unital $\mathcal{E}_{n}$-algebras, via our construction of higher enveloping algebras of spectral Lie algebras.

math.AT

Multi-stage Progressive Reasoning for Dunhuang Murals Inpainting

Dunhuang murals suffer from fading, breakage, surface brittleness and extensive peeling affected by prolonged environmental erosion. Image inpainting techniques are widely used in the field of digital mural inpainting. Generally speaking, for mural inpainting tasks with large area damage, it is challenging for any image inpainting method. In this paper, we design a multi-stage progressive reasoning network (MPR-Net) containing global to local receptive fields for murals inpainting. This network is capable of recursively inferring the damage boundary and progressively tightening the regional texture constraints. Moreover, to adaptively fuse plentiful information at various scales of murals, a multi-scale feature aggregation module (MFA) is designed to empower the capability to select the significant features. The execution of the model is similar to the process of a mural restorer (i.e., inpainting the structure of the damaged mural globally first and then adding the local texture details further). Our method has been evaluated through both qualitative and quantitative experiments, and the results demonstrate that it outperforms state-of-the-art image inpainting methods.

cs.CV

Goodwillie's cosimplicial model for the space of long knots and its applications

We work out the details of a correspondence observed by Goodwillie between cosimplicial spaces and good functors from a category of open subsets of the interval to the category of compactly generated weak Hausdorff spaces. Using this, we compute the first page of the integral Bousfield--Kan homotopy spectral sequence of the tower of fibrations given by the Taylor tower of the embedding functor associated to the space of long knots. Based on the methods in [Con08], we give a combinatorial interpretation of the differentials $d^1$ mapping into the diagonal terms, by introducing the notion of $(i, n)$-marked unitrivalent graphs.

math.AT