Searcharxiv⌕ Search

arXiv · 2609.23019

SomaNet: Weakly Supervised Learning for Instance Soma Segmentation in 3D Electron Microscopy with Partial Annotations

Abstract

Soma instance segmentation, i.e., identifying and delineating individual cell somas as distinct instances, is crucial for cellular analysis and connectomic reconstruction. Three-dimensional electron microscopy (3D EM) provides nanometer-scale resolution for capturing fine-grained soma morphology. However, dense instance-level manual annotation is prohibitively costly, limiting the scalability of fully supervised methods. To address this challenge, we propose SomaNet, a weakly supervised framework for 3D EM soma instance segmentation under partial annotation constraints. SomaNet adopts a teacher--student learning paradigm tailored to partial labels. The teacher is trained using partially annotated data to generate pseudo-labels, while the student jointly learns from the partial ground-truth annotations and the generated pseudo-labels, progressively recovering dense instance segmentations. To accurately delineate cell somas under limited supervision with varying instance counts, SomaNet incorporates affinity learning, which encourages high similarity within instances and low similarity across instance boundaries. Semantic-guided affinity decoding and 2D-to-3D reconstruction then produce volumetrically consistent 3D soma instances while preserving the large receptive fields of 2D backbones. The framework is architecture-flexible and supports diverse backbones, including vision transformers and foundation models, enabling direct transfer of pretrained visual representations to volumetric EM segmentation. Experiments on 3D EM brain datasets demonstrate that SomaNet achieves accurate and robust soma instance segmentation across regions with diverse soma morphologies under partial annotation. Code is available at https://github.com/mkhateri/SomaNet.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mohammad Khateri, Morteza Ghahremani, Jussi Tohka, Alejandra Sierra. 2026-09-19. SomaNet: Weakly Supervised Learning for Instance Soma Segmentation in 3D Electron Microscopy with Partial Annotations. https://arxiv.org/abs/2609.23019

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Rigid Motion Estimation using Accelerated Iterative Coordinate Descent (REACT) for MR Imaging

Purpose: To develop a computationally viable autofocus method for estimating 3D rigid motion in MR imaging. Theory and Methods: The proposed method, REACT, assumes a piecewise-constant motion trajectory and estimates the rigid motion parameters of individual temporal segments by optimizing an image-quality metric. Coordinate descent is adopted to decompose the high-dimensional optimization problem into a series of subproblems, each updating the motion parameters of a single temporal segment. The cost function of each subproblem is assumed to be approximately locally convex under suitable acquisition conditions. Each subproblem is then solved using a derivative-free solver, thereby avoiding an exhaustive grid search. Numerical simulations investigated the local convexity assumption and data acquisition requirements. REACT was evaluated for respiratory motion correction on in vivo free-breathing coronary MR angiography datasets. Coronary artery sharpness was quantified using unbounded image edge profile acutance (u-IEPA). Results: In numerical simulations, the objective surfaces of the subproblems were approximately locally convex when the current motion estimate was sufficiently close to the desired solution, and REACT required the data collected within each temporal segment to be sufficiently distributed across k-space. In the in vivo study, REACT yielded higher u-IEPA for both the left anterior descending artery (LAD) and the right coronary artery than did a conventional translational motion-estimation method using image-based navigators. REACT also yielded higher u-IEPA for the LAD than did a conventional autofocus nonrigid motion correction method. Conclusion: This study demonstrates the feasibility of coordinate descent for autofocus motion correction in MR imaging.

eess.IV↗

Projected Energy Matching for Generative 3D Priors

Transport-based generative models, which learn a time-dependent vector field that moves noise to data, have become a dominant paradigm. However, these models typically do not explicitly encode the data distribution. Energy-based models (EBMs) instead represent the data distribution explicitly through a scalar energy landscape, which Energy Matching learns by combining transport learning with contrastive refinement. Its transport objective, however, fits energy gradients to stochastic targets, whose variance degrades the training signal at scale. We introduce Projected Energy Matching, which learns this landscape through a more stable route: we first train a time-independent transport teacher, then freeze it and fit the negative energy gradient to its predicted velocities. This projection replaces noisy transport targets with deterministic supervision, while contrastive refinement shapes the landscape near the data manifold. On CIFAR-10, gradient-noise analysis reveals a cleaner training signal, accompanied by faster convergence than Energy Matching at matched, teacher-free training budgets. In the latent space of CT volumes, our method enables 3D CT generation and achieves better FID scores than flow models. The learned scalar potential serves as a zero-shot prior for the ill-posed inverse problem of sparse-view cone-beam CT reconstruction. By making explicit energy landscapes practical at volumetric scale, this work opens a path to wider adoption of energy-based formulations, bringing their flexibility to high-dimensional generation and inverse problems.

eess.IV↗

Computer vision enabled oxygen sensing

Luminescence-based chemical sensors are almost universally read as point detectors with the signal inverted through a single calibration model. Here we reframe optical oxygen sensing as a computer vision problem, in which the sensing film acts as a spatially heterogeneous encoder and a pretrained Temporal Vision Transformer (TViT) as its decoder. The heterogeneity in film thickness, diffusion path length and illumination translate to pixels that provide complementary information on the underlying diffusion dynamics. We achieved up to 54% lower mean absolute error (MAE) by increasing the internal diversity in performance of a pixel group; a gain that linear models cannot reproduce. Using a low-cost platform comprising a Raspberry Pi camera, a UV LED and a porous PtOEP/polystyrene film in tandem with a TViT architecture yielded an MAE of ~6.7 μmol/L, a 96% reduction from the two-site Stern-Volmer (SV) model. We showed that this framework can computationally mitigate the fundamental trade-off between mechanical robustness and temporal response in diffusion limited oxygen sensing by reducing T90 response times by 91%, while exhibiting physically plausible dynamics under a Rauch-Tung-Striebel smoother (2.3% flag rate). The framework was applied across setups, environments and biofilm states, establishing an IoT-compatible paradigm for computationally compensated diffusion-limited chemical sensing.

eess.IV↗