SearcharxivSearch

arXiv subjects

Evan W. Damron

Publications and source records attributed to Evan W. Damron.

4 recordsLinked to original sources

DALE-CT: Depth-Aware 2D Slice Encoders Learn an Anatomical World Model of Chest CT

Chest CT is among the highest-volume imaging exams in medicine, yet expert voxel-level annotations are scarce and costly, motivating encoders that learn directly from unlabeled scans. We present DALE-CT, a family of 2D slice-based Vision Transformers trained from scratch on chest CT with the heuristics-free LeJEPA objective. We introduce depth-aware slab sampling, which draws self-supervised views from across a physical $z$-axis slab rather than a single slice, implicitly tasking the 2D encoder with representing how anatomy changes between neighboring slices. The frozen representations trace each scan as a smooth anatomical trajectory, recover cranio-caudal slice ordering without labels, and distinguish slices by the anatomy they contain rather than by position alone. This anatomical world model emerges without any 3D or positional supervision, and an otherwise-identical encoder trained on slices in isolation never develops it. Building on this backbone, we introduce dense auxiliary supervision into the pretraining objective, using anatomical and abnormality masks to supervise patch and slice tokens alongside the self-supervised loss, and we compare the resulting variants against a DINOv2 baseline continually pretrained on CT-RATE. Among nine public and in-house models evaluated under the same protocol, DALE-CT-2S is the strongest 2D model in-domain, reaching 0.825 Macro AUROC on CT-RATE, within 0.024 of COLIPRI-CRM and without any text supervision. We subsequently scale the supervision-free configuration to a $\sim$287k-scan multi-source pool, to our knowledge the largest reported chest-CT pretraining corpus. The resulting DALE-CT-0-L posts the best 2D external-transfer point estimates, and we release it as our recommended backbone with the full model family, training code, and evaluation pipeline.

cs.CV

Lesion Detection in CT with Frozen Self-Distilled Features: SALT, a Spatially Adaptive Label-Guided Temperature

Self-supervised pretraining objectives are spatially uniform: the teacher temperature and the per-patch loss weight are identical everywhere in the image, so a lesion a few patches wide contributes no more to the training signal than the surrounding parenchyma. Prior work biases the views toward annotated regions, which changes what the model sees but adds no pressure on the objective. We instead condition the targets of self-distillation, a method we call SALT (Spatially Adaptive Label-guided Temperature). Weak, box-derived labels, available only during pretraining, define a compact region on the encoder's patch grid, inside which the teacher's softmax temperature is sharpened and the masked-patch loss is up-weighted. The objectives, the masking policy and the centering statistics are otherwise unchanged, and at every downstream use the encoder is a plain feature extractor with no labels and no conditioning. We evaluate by freezing the encoder and training only a lightweight multi-depth CenterNet-style head, detecting lesions in 3D on four CT cohorts, and we isolate the mechanism against a backbone identical in architecture, pretraining data, schedule and label-guided cropping but with no target conditioning. We report patch-level separability, 3D detection stratified by cohort and by lesion size, box quality, and a detector-free probe in which a single frozen patch embedding re-identifies a lesion in a follow-up scan without registration, masks or fine-tuning. Because the conditioning is expressed through a spatial indicator rather than through label semantics, the formulation admits any weak spatial annotation; we instantiate and validate it for lesions.

cs.CV

Curriculum-Driven 3D CT Report Generation via Language-Free Visual Grafting and Zone-Constrained Compression

Automated radiology report generation from 3D computed tomography (CT) volumes is challenging due to extreme sequence lengths, severe class imbalance, and the tendency of large language models (LLMs) to ignore visual tokens in favor of linguistic priors. We present Ker-VLJEPA-3B, a four-phase curriculum learning framework for free-text report generation from thoracic CT volumes. A phased training curriculum progressively adapts a Llama 3.2 3B decoder to ground its output in visual features from a frozen, self-supervised encoder. Our visual backbone (LeJEPA ViT-Large) is trained via self-supervised joint-embedding prediction on unlabeled CTs, without text supervision. Unlike contrastive models (CLIP, BiomedCLIP), this language-free backbone yields modality-pure representations. Vision-language alignment is deferred to the curriculum's bridge and generation phases. This modality-agnostic design can integrate any self-supervised encoder into an LLM without paired text during foundation training. Methodological innovations include: (1) zone-constrained cross-attention compressing slice embeddings into 32 spatially-grounded visual tokens; (2) PCA whitening of anisotropic LLM embeddings; (3) a positive-findings-only strategy eliminating posterior collapse; (4) warm bridge initialization transferring projection weights; and (5) selective cross-attention freezing with elastic weight consolidation to prevent catastrophic forgetting. Evaluated on the CT-RATE benchmark (2,984 validation volumes, 18 classes), Ker-VLJEPA-3B achieves a macro F1 of 0.429, surpassing the state-of-the-art (U-VLM, macro F1 = 0.414) by 3.6%, and reaching 0.448 (+8.2%) with threshold optimization. Ablation studies confirm 56.6% of generation quality derives from patient-specific visual content. Code and weights are available.

cs.CV

Vision Foundry: A System for Training Foundational Vision AI Models

Self-supervised learning (SSL) leverages vast unannotated medical datasets, yet steep technical barriers limit adoption by clinical researchers. We introduce Vision Foundry, a code-free, HIPAA-compliant platform that democratizes pre-training, adaptation, and deployment of foundational vision models. The system integrates the DINO-MX framework, abstracting distributed infrastructure complexities while implementing specialized strategies like Magnification-Aware Distillation (MAD) and Parameter-Efficient Fine-Tuning (PEFT). We validate the platform across domains, including neuropathology segmentation, lung cellularity estimation, and coronary calcium scoring. Our experiments demonstrate that models trained via Vision Foundry significantly outperform generic baselines in segmentation fidelity and regression accuracy, while exhibiting robust zero-shot generalization across imaging protocols. By bridging the gap between advanced representation learning and practical application, Vision Foundry enables domain experts to develop state-of-the-art clinical AI tools with minimal annotation overhead, shifting focus from engineering optimization to clinical discovery.

q-bio.QM