SearcharxivSearch

arXiv subjects

Peihang Wu

Publications and source records attributed to Peihang Wu.

5 recordsLinked to original sources

PathSelect: Sequential Token Selection for Whole Slide Pathology

Gigapixel Whole-Slide Images (WSIs) present a fundamental computational bottleneck for vision-language models (VLMs) due to extreme sequence lengths. Existing approaches predominantly rely on spatial sampling or training-free pruning, which risk diluting weak but informative signals, leading to the loss of critical diagnostic evidence due to the spatially diffuse nature of pathological cues. We reformulate WSI token pruning as a sequential selection process, enabling the model to autonomously learn an optimal routing strategy rather than relying on static heuristics. We herein propose a decoupled routing framework integrated as an active plugin into the fully pre-trained SlideChat base model, leaving both the slide encoder and large language model frozen. To provide continuous gradients for the non-differentiable pruning operation during training, we introduce PathSelect. PathSelect employs a variance-preserving noise gate to modulate each patch's information flow via a differentiable Soft Top-K operator, paired with a diagonal-attention Denoiser that recovers the perturbed representations without semantic leakage. At inference, the PathSelect module is entirely detached. Relying solely on the trained Scorer, a deterministic Hard Top-K operator executes adaptive, data-dependent trajectory termination, significantly accelerating downstream generative processing with exceptionally low sequential token selection latency. Driven by an empirical average of only 44.86 tokens under a maximum constraint of K = 128, our framework achieves 74.00% overall accuracy on SlideBench (TCGA), representing an approximate 36.6x spatial token reduction relative to the uncompressed baseline average while consistently outperforming sampling-based counterparts.

cs.CV

Canonical extensions of $p$-adic shtukas on toroidal compactifications of Shimura varieties

We construct canonical extensions of $p$-adic shtukas on integral models of toroidal compactifications of abelian-type Shimura varieties with quasi-parahoric levels at any prime number $p$. More precisely, we define the notion of a log diamond as a $v$-sheaf associated with a log scheme over $\mathbb{Z}_p$ and construct a $p$-adic log shtuka over the log diamond of an integral toroidal compactification of an abelian-type Shimura variety by studying the ``degeneration'' of the shtuka at the boundary. Moreover, we provide a definition of canonical integral models of toroidal and minimal compactifications in the sense of Pappas and Rapoport, and verify it in the same generality as above. Applications include the canonicity and functoriality of integral toroidal compactifications, as well as an axiomatic proof of the well-positionedness of all well-known stratifications on the special fiber.

math.NT

Multimodal Model for Computational Pathology:Representation Learning and Image Compression

Whole slide imaging (WSI) has transformed digital pathology by enabling computational analysis of gigapixel histopathology images. Recent foundation model advances have accelerated progress in computational pathology, facilitating joint reasoning across pathology images, clinical reports, and structured data. Despite this progress, challenges remain: the extreme resolution of WSIs creates computational hurdles for visual learning; limited expert annotations constrain supervised approaches; integrating multimodal information while preserving biological interpretability remains difficult; and the opacity of modeling ultra-long visual sequences hinders clinical transparency. This review comprehensively surveys recent advances in multimodal computational pathology. We systematically analyze four research directions: (1) self-supervised representation learning and structure-aware token compression for WSIs; (2) multimodal data generation and augmentation; (3) parameter-efficient adaptation and reasoning-enhanced few-shot learning; and (4) multi-agent collaborative reasoning for trustworthy diagnosis. We specifically examine how token compression enables cross-scale modeling and how multi-agent mechanisms simulate a pathologist's "Chain of Thought" across magnifications to achieve uncertainty-aware evidence fusion. Finally, we discuss open challenges and argue that future progress depends on unified multimodal frameworks integrating high-resolution visual data with clinical and biomedical knowledge to support interpretable and safe AI-assisted diagnosis.

cs.CV

Arithmetic compactifications of integral models of Shimura varieties of abelian type

In this paper, we construct good toroidal and minimal compactifications in the sense of Lan-Stroh for integral models of abelian-type Shimura varieties. We start with finding suitable types of cusp labels and cone decompositions which are compatible with those of the associated Hodge-type Shimura varieties. We then study the action of $\mathbb{Q}$-points of the adjoint group on boundary charts and toroidal compactifications of Hodge-type integral models. In particular, we extend the twisting construction of Kisin and Pappas to boundary charts. Finally, up to taking refinements of cone decompositions, we construct an abelian-type toroidal compactification as an open and closed algebraic subspace of a quotient from a disjoint union of Hodge-type toroidal compactifications and construct minimal compactifications with a similar method. Furthermore, we show results on nearby cycles of these compactifications and verify Pink's formula when the level at $p$ is an intersection of $n$ quasi-parahoric subgroups.

math.NT