SearcharxivSearch

arXiv subjects

Daniel Lin

Publications and source records attributed to Daniel Lin.

6 recordsLinked to original sources

Improving Cross-Site Whole-Heart Segmentation

Whole-heart segmentation from CT and MRI is essential for quantitative cardiac image analysis, but remains challenging under multi-center and multi-modality distribution shift. In the CARE whole-heart segmentation task, models must generalize from limited labeled sites to unseen acquisition distributions, where variation in spacing, intensity, reconstruction texture, and anatomy can degrade out-of-distribution performance. We propose a modality-routed 3D cardiac segmentation pipeline that combines TotalSegmentator-initialized nnU-Netv2 models with site-characterized, label-preserving appearance augmentation. We first characterize the available sites using measurable image properties and use this analysis to motivate candidate data-space generalization routes. The final retained recipe applies Bias Field + Bezier appearance augmentation, combining smooth spatial intensity perturbation with nonlinear intensity remapping, followed by lightweight class-wise largest-connected-component cleanup. On the primary held-out-site validation splits, the final configuration improves CT mean Dice from 0.8350 to 0.9135 and MRI mean Dice from 0.7695 to 0.7830, while also reducing HD95. These results suggest that site-motivated appearance augmentation is a practical strategy for improving cross-site robustness in limited-data whole-heart segmentation. Our code can be found in https://github.com/Purdue-M2/Improving-Cross-Site-Whole-Heart-Segmentation

eess.IV

olmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Models

PDF documents have the potential to provide trillions of novel, high-quality tokens for training language models. However, these documents come in a diversity of types with differing formats and visual layouts that pose a challenge when attempting to extract and faithfully represent the underlying content for language model use. Traditional open source tools often produce lower quality extractions compared to vision language models (VLMs), but reliance on the best VLMs can be prohibitively costly (e.g., over 6,240 USD per million PDF pages for GPT-4o) or infeasible if the PDFs cannot be sent to proprietary APIs. We present olmOCR, an open-source toolkit for processing PDFs into clean, linearized plain text in natural reading order while preserving structured content like sections, tables, lists, equations, and more. Our toolkit runs a fine-tuned 7B vision language model (VLM) trained on olmOCR-mix-0225, a sample of 260,000 pages from over 100,000 crawled PDFs with diverse properties, including graphics, handwritten text and poor quality scans. olmOCR is optimized for large-scale batch processing, able to scale flexibly to different hardware setups and can convert a million PDF pages for only 176 USD. To aid comparison with existing systems, we also introduce olmOCR-Bench, a curated set of 1,400 PDFs capturing many content types that remain challenging even for the best tools and VLMs, including formulas, tables, tiny fonts, old scans, and more. We find olmOCR outperforms even top VLMs including GPT-4o, Gemini Flash 2 and Qwen-2.5-VL. We openly release all components of olmOCR: our fine-tuned VLM model, training code and data, an efficient inference pipeline that supports vLLM and SGLang backends, and benchmark olmOCR-Bench.

cs.CL

The Semantic Scholar Open Data Platform

The volume of scientific output is creating an urgent need for automated tools to help scientists keep up with developments in their field. Semantic Scholar (S2) is an open data platform and website aimed at accelerating science by helping scholars discover and understand scientific literature. We combine public and proprietary data sources using state-of-the-art techniques for scholarly PDF content extraction and automatic knowledge graph construction to build the Semantic Scholar Academic Graph, the largest open scientific literature graph to-date, with 200M+ papers, 80M+ authors, 550M+ paper-authorship edges, and 2.4B+ citation edges. The graph includes advanced semantic features such as structurally parsed text, natural language summaries, and vector embeddings. In this paper, we describe the components of the S2 data processing pipeline and the associated APIs offered by the platform. We will update this living document to reflect changes as we add new data offerings and improve existing services.

cs.DL

Presheaves over a join restriction category

Just as the presheaf category is the free cocompletion of any small category, there is an analogous notion of free cocompletion for any small restriction category. In this paper, we extend the work on restriction presheaves to presheaves over join restriction categories, and show that the join restriction category of join restriction presheaves is equivalent to some partial map category of sheaves. We then use this to show that the Yoneda embedding exhibits the category of join restriction presheaves as the free cocompletion of any small join restriction category.

math.CT

Cocompletion of restriction categories

Restriction categories were introduced as a way of generalising the notion of partial map categories. In this paper, we define cocomplete restriction category, and give the free cocompletion of a small restriction category as a suitably defined category of restriction presheaves. We also consider the case where our restriction category is locally small.

math.CT

Slant-gap plasmonic nanoantenna for optical chirality enhancement

We present a new design of plasmonic nanoantenna with a slant gap for optical chirality engineering. At resonance, the slant gap provides highly enhanced electric field parallel to external magnetic field with a phase delay of 90 degree, resulting in enhanced optical chirality. We show by numerical simulations that upon linearly polarized excitation our achiral nanoantenna can generate near field with enhanced optical chirality that can be tuned by the slant angle and resonance condition. Our design can be easily realized and may find applications in circular dichroism enhancement.

cond-mat.mes-hall