SearcharxivSearch

arXiv subjects

Yuxiang Xing

Publications and source records attributed to Yuxiang Xing.

15 recordsLinked to original sources

LabOSBench: Benchmarking Computer Use Agents for Scientific Instrument Control

Current computer-use benchmarks primarily focus on software operation tasks in virtualized systems, whereas scientific instrumentation scenarios require coordinated control over complex interfaces, and feedback-driven parameter adjustment. However, directly evaluating agents on physical high-precision instruments is impractical due to high cost, safety risks, limited accessibility, and difficulty in ensuring reproducible evaluation. This motivates the need for a simulated yet realistic testbed that preserves the operational challenges of scientific instruments while enabling scalable and safe benchmarking. To this end, we introduce LabOSBench, a challenging benchmark for multimodal GUI agents built on a suite of web-based scientific-instrument simulators. Operating directly via a browser, LabOSBench avoids resource-heavy OS virtualization while supporting flexible task configuration and execution-based evaluation. Specifically, LabOSBench constructs 96 subtasks across eight instrument simulators, covering workflows from sample loading, alignment, parameter tuning, and data acquisition to result inspection. We evaluate general-purpose vision-language models, specialized GUI agent models, and advanced agentic frameworks at both subtask and end-to-end levels. Our experiments reveal that while existing agents can complete many structured GUI subtasks, they still struggle with feedback-driven operations and long-horizon workflow execution. Overall, LabOSBench provides a reproducible, low-cost testbed for advancing computer-using agents toward scientific-instrument control.

cs.AI

Energy-Threshold Bias Calculator: A Physics-Model Based Adaptive Correction Scheme for Photon-Counting CT

Photon-counting detector based computed tomography (PCCT) has greatly advanced in recent years. However, spectral inconsistency, referring to inter-pixel variations in detected counts per energy bin, can easily leads to ring or band artifacts and inaccuracies in CT reconstructed images. This work proposes a novel physics-model based method to correct for spectral inconsistency by modeling it through two terms: (1) a fixed spectral skew term (energy threshold-independent filtration function) determined at a given energy threshold, and (2) a variable energy-threshold bias term that can be directly calculated by using our spectral model as the threshold changes. After the two terms being computed out in the calibration stage, they will be incorporated into our spectral model to adaptively generate the spectral correction vectors as well as the material decomposition vectors if needed, pixel-by-pixel for PCCT projection data. Using a minimum set of parameters with explicit physics meaning, such an energy-threshold bias calculator (ETB-Cal) has advantages of computational efficiency, robustness in implementation, and convenience with no need of X-ray fluorescence materials in calibration. To validate our method, both numerical simulations and physical experiments using multiple phantoms were carried out on a tabletop PCCT system, with preliminary results showing a significant reduction in non-uniformity, from 29.3 to 5.8 HU for Gammex multi-energy phantom versus no correction (comparatively, 8.3 HU was achieved by a polynomial-involving model-based approach with no explicit modeling and calculating of energy threshold bias but more calibration data required), and from 27.9 to 3.2 HU for the Kyoto head phantom.

physics.med-ph

Nonperiodic dynamic CT reconstruction using backward-warping INR with regularization of diffeomorphism (BIRD)

Dynamic computed tomography (CT) reconstruction faces significant challenges in addressing motion artifacts, particularly for nonperiodic rapid movements such as cardiac imaging with fast heart rates. Traditional methods struggle with the extreme limited-angle problems inherent in nonperiodic cases. Deep learning methods have improved performance but face generalization challenges. Recent implicit neural representation (INR) techniques show promise through self-supervised deep learning, but have critical limitations: computational inefficiency due to forward-warping modeling, difficulty balancing DVF complexity with anatomical plausibility, and challenges in preserving fine details without additional patient-specific pre-scans. This paper presents a novel INR-based framework, BIRD, for nonperiodic dynamic CT reconstruction. It addresses these challenges through four key contributions: (1) backward-warping deformation that enables direct computation of each dynamic voxel with significantly reduced computational cost, (2) diffeomorphism-based DVF regularization that ensures anatomically plausible deformations while maintaining representational capacity, (3) motion-compensated analytical reconstruction that enhances fine details without requiring additional pre-scans, and (4) dimensional-reduction design for efficient 4D coordinate encoding. Through various simulations and practical studies, including digital and physical phantoms and retrospective patient data, we demonstrate the effectiveness of our approach for nonperiodic dynamic CT reconstruction with enhanced details and reduced motion artifacts. The proposed framework enables more accurate dynamic CT reconstruction with potential clinical applications, such as one-beat cardiac reconstruction, cinematic image sequences for functional imaging, and motion artifact reduction in conventional CT scans.

cs.CV

DLBayesian: An Alternative Bayesian Reconstruction of Limited-view CT by Optimizing Deep Learning Parameters

Limited-view computed tomography (CT) presents significant potential for reducing radiation exposure and expediting the scanning process. While deep learning (DL) methods have exhibited promising results in mitigating streaking artifacts caused by a reduced number of projection views, their generalization remains challenging. In this work, we proposed a DL-driven alternative Bayesian reconstruction method (DLBayesian) that efficiently integrates data-driven priors and data consistency constraints. DLBayesian comprises three stages: group-level embedding, significance evaluation, and individual-level consistency adaptation. Firstly, DL network parameters are optimized to learn how to eliminate the general limited-view artifacts on a large-scale paired dataset. Then, we introduced a significance score to quantitatively evaluate the contribution of parameters in DL models as a guide for the subsequent individual-level adaptation. Finally, in the Bayesian adaptation stage, an alternative Bayesian reconstruction further optimizes the DL network parameters precisely according to the projection data of the target case. We validated DLBayesian with sparse-view (90 views) projections from a circular trajectory CT and a special data missing case from a multi-segment linear trajectory CT. The results underscore DLBayesian's superior generalization capabilities across variations in patients, anatomic structures, and data distribution, as well as excelling in contextual structure recovery compared to networks solely trained via supervised loss. Real experiments on a dead rat demonstrate its capability in practical CT scans.

physics.med-ph

ComptoNet: An End-to-End Deep Learning Framework for Scatter Estimation in Multi-Source Stationary CT

Multi-source stationary computed tomography (MSS-CT) offers significant advantages in medical and industrial applications due to its gantry-less scan architecture and/or capability of simultaneous multi-source emission. However, the lack of anti-scatter grid deployment in MSS-CT results in severe forward and/or cross scatter contamination, presenting a critical challenge that necessitates an accurate and efficient scatter correction. In this work, ComptoNet, an innovative end-to-end deep learning framework for scatter estimation in MSS-CT, is proposed, which integrates Compton-scattering physics with deep learning techniques to address the challenges of scatter estimation effectively. Central to ComptoNet is the Compton-map, a novel concept that captures the distribution of scatter signals outside the scan field of view, primarily consisting of large-angle Compton scatter. In ComptoNet, a reference Compton-map and/or spare detector data are used to guide the physics-driven deep estimation of scatter from simultaneous emissions by multiple sources. Additionally, a frequency attention module is employed for enhancing the low-frequency smoothness. Such a multi-source deep scatter estimation framework decouples the cross and forward scatter. It reduces network complexity and ensures a consistent low-frequency signature with different photon numbers of simulations, as evidenced by mean absolute percentage errors (MAPEs) that are less than $1.26\%$. Conducted by using data generated from Monte Carlo simulations with various phantoms, experiments demonstrate the effectiveness of ComptoNet, with significant improvements in scatter estimation accuracy (a MAPE of $0.84\%$). After scatter correction, nearly artifact-free CT images are obtained, further validating the capability of our proposed ComptoNet in mitigating scatter-induced errors.

physics.med-ph

A square cross-section FOV rotational CL (SC-CL) and its analytical reconstruction method

Rotational computed laminography (CL) has broad application potential in three-dimensional imaging of plate-like objects, as it only needs x-ray to pass through the tested object in the thickness direction during the imaging process. In this study, a square cross-section FOV rotational CL (SC-CL) was proposed. Then, the FDK-type analytical reconstruction algorithm applicable to the SC-CL was derived. On this basis, the proposed method was validated through numerical experiments.

eess.IV

Generalized-Equiangular Geometry CT: Concept and Shift-Invariant FBP Algorithms

With advanced X-ray source and detector technologies being continuously developed, non-traditional CT geometries have been widely explored. Generalized-Equiangular Geometry CT (GEGCT) architecture, in which an X-ray source might be positioned radially far away from the focus of arced detector array that is equiangularly spaced, is of importance in many novel CT systems and designs. GEGCT, unfortunately, has no theoretically exact and shift-invariant analytical image reconstruction algorithm in general. In this study, to obtain fast and accurate reconstruction from GEGCT and to promote its system design and optimization, an in-depth investigation on a group of approximate Filtered BackProjection (FBP) algorithms with a variety of weighting strategies has been conducted. The architecture of GEGCT is first presented and characterized by using a normalized-radial-offset distance (NROD). Next, shift-invariant weighted FBP-type algorithms are derived in a unified framework, with pre-filtering, filtering, and post-filtering weights. Three viable weighting strategies are then presented including a classic one developed by Besson in the literature and two new ones generated from a curvature fitting and from an empirical formula, where all of the three weights can be expressed as certain functions of NROD. After that, an analysis of reconstruction accuracy is conducted with a wide range of NROD. We further stretch the weighted FBP-type algorithms to GEGCT with dynamic NROD. Finally, the weighted FBP algorithm for GEGCT is extended to a three-dimensional form in the case of cone-beam scan with a cylindrical detector array.

physics.med-ph

Extraction-based Deep Learning Reconstruction of Interior Tomography

Interior tomography is a typical strategy for radiation dose reduction in computed tomography, where only a certain region-of-interest (ROI) is scanned. However, given the truncated projection data, ROI reconstruction by conventional analytical algorithms may suffer from severe cupping artifacts. In this paper, we proposed a new extraction-based deep learning method for the reconstruction of interior tomography. Our approach works in dual domains where a sinogram-domain network (SDNet) estimates the contribution of the exterior region to the truncated projection and an image-domain network (IDNet) further mitigates artifacts. Unlike the previous extrapolation-based methods, SDNet is intended to obtain a complete ROI-only sinogram via extraction instead of a fully non-truncated sinogram for both the ROI and exterior regions. Our experiments validated the proposed method and the results indicate that the proposed method can disclose more reliable structures. It achieved better image quality with better generalization performance than extrapolation-based methods.

physics.med-ph

Investigation of domain gap problem in several deep-learning-based CT metal artefact reduction methods

Metal artefacts in CT images may disrupt image quality and interfere with diagnosis. Recently many deep-learning-based CT metal artefact reduction (MAR) methods have been proposed. Current deep MAR methods may be troubled with domain gap problem, where methods trained on simulated data cannot perform well on practical data. In this work, we experimentally investigate two image-domain supervised methods, two dual-domain supervised methods and two image-domain unsupervised methods on a dental dataset and a torso dataset, to explore whether domain gap problem exists or is overcome. We find that I-DL-MAR and DudoNet are effective for practical data of the torso dataset, indicating the domain gap problem is solved. However, none of the investigated methods perform satisfactorily on practical data of the dental dataset. Based on the experimental results, we further analyze the causes of domain gap problem for each method and dataset, which may be beneficial for improving existing methods or designing new ones. The findings suggest that the domain gap problem in deep MAR methods remains to be addressed.

cs.CV

Fluence Adaptation for Task-based Dose Optimization in X-ray Phase-Contrast Imaging

Purpose: Grating-based imaging (GBI) and edge-illumination (EI) are two promising types of XPCI as the conventional x-ray sources can be directly utilized. For GBI and EI systems, the phase-stepping acquisition with multiple exposures at a constant fluence is usually adopted in the literature. This work, however, attempts to challenge such a constant fluence concept during the phase-stepping process and proposes a fluence adaptation mechanism for dose reduction. Method: Recently, analytic multi-order moment analysis has been proposed to improve the computing efficiency. In these algorithms, multiple contrasts can be calculated by summing together the weighted phase-stepping curves (PSCs) with some kernel functions, which suggests us that the raw data at different steps have different contributions for the noise in retrieved contrasts. Based on analytic retrieval formulas and the Gaussian noise model for detected signals, we derived an optimal adaptive fluence distribution, which is proportional to the absolute weighting kernel functions and the root of original sample PSCs acquired under the constant fluence. Results: To validate our analyses, simulations and experiments are conducted for GBI and EI systems. Simulated results demonstrate that the dose reduction ratio between our proposed fluence distributions and the typical constant one can be about 20% for the phase contrast, which is consistent with our theoretical predictions. Although the experimental noise reduction ratios are a little smaller than the theoretical ones, synthetic and real experiments both observe better noise performance by our proposed method. Our simulated results also give out the effective ranges of the parameters of the PSCs, such as the visibility in GBI, the standard deviation and the mean value in EI, providing a guidance for the use of our proposed approach in practice.

physics.med-ph

A novel deep learning-based method for monochromatic image synthesis from spectral CT using photon-counting detectors

With the growing technology of photon-counting detectors (PCD), spectral CT is a widely concerned topic which has the potential of material differentiation. However, due to some non-ideal factors such as cross talk and pulse pile-up of the detectors, direct reconstruction from detected spectrum without any corrections will get a wrong result. Conventional methods try to model these factors using calibration and make corrections accordingly, but depend on the preciseness of the model. To solve this problem, in this paper, we proposed a novel deep learning-based monochromatic image synthesis method working in sinogram domain. Different from previous deep learning-based methods aimed at this problem, we designed a novel network architecture according to the physical model of cross talk, and it can solve this problem better in an ingenious way. Our method was tested on a cone-beam CT (CBCT) system equipped with a PCD. After using FDK algorithm on the corrected projection, we got quite more accurate results with less noise, which showed the feasibility of monochromatic image synthesis by our method.

physics.med-ph

Truncated analytic moment analysis and its hybrid-field contrast in grating-based x-ray phase contrast imaging

For grating-based x-ray phase contrast imaging (GPCI), a multi-order moment analysis (MMA) has been recently developed to obtain multiple contrasts from the ultra-small-angle x-ray scattering distribution, as a novel information retrieval approach that is totally different from the conventional Fourier components analysis (FCA). In this paper, we present an analytic form of MMA in theory that can retrieve multiple contrasts directly from raw phase-stepping images, with no scattering distribution involved. For practical implementation, a truncated analytic analysis (called as TA-MMA) is adopted and it is hundreds of times faster in computation than the original deconvolution-based MMA (called as DB-MMA). More importantly, TA-MMA is proved to establish a quantitative connection between FCA and MMA, i.e., the first-order moment computed by TA-MMA is essentially the product of the phase contrast and the dark-field contrast retrieved by FCA, providing a new physical parameter for GPCI. The new physical parameter, in fact, can be treated as a ``hybrid-field'' contrast as it fuses the original phase contrast and dark-field contrast in a straightforward manner, which may be the first physical fusion contrast in GPCI to our knowledge and may have a potential to be directly used in practical applications.

physics.med-ph

A cascaded dual-domain deep learning reconstruction method for sparsely spaced multidetector helical CT

Helical CT has been widely used in clinical diagnosis. Sparsely spaced multidetector in z direction can increase the coverage of the detector provided limited detector rows. It can speed up volumetric CT scan, lower the radiation dose and reduce motion artifacts. However, it leads to insufficient data for reconstruction. That means reconstructions from general analytical methods will have severe artifacts. Iterative reconstruction methods might be able to deal with this situation but with the cost of huge computational load. In this work, we propose a cascaded dual-domain deep learning method that completes both data transformation in projection domain and error reduction in image domain. First, a convolutional neural network (CNN) in projection domain is constructed to estimate missing helical projection data and converting helical projection data to 2D fan-beam projection data. This step is to suppress helical artifacts and reduce the following computational cost. Then, an analytical linear operator is followed to transfer the data from projection domain to image domain. Finally, an image domain CNN is added to improve image quality further. These three steps work as an entirety and can be trained end to end. The overall network is trained using a simulated lung CT dataset with Poisson noise from 25 patients. We evaluate the trained network on another three patients and obtain very encouraging results with both visual examination and quantitative comparison. The resulting RRMSE is 6.56% and the SSIM is 99.60%. In addition, we test the trained network on the lung CT dataset with different noise level and a new dental CT dataset to demonstrate the generalization and robustness of our method.

physics.med-ph

A Model-based Deep Learning Reconstruction for X-ray CT

Low dose CT is of great interest in these days. Dose reduction raises noise level in projections and decrease image quality in reconstructions. Model based image reconstruction can combine statistical noise model together with prior knowledge into an Bayesian optimization problem so that significantly reduce noise and artefacts. In this work, we propose a model-base deep learning for CT reconstruction so that a reconstruction network can be trained with no ground-truth images needed. Instead of minimizing cost function for each image, the network learns to minimize an ensemble cost function for the whole training set. No iteration will be needed for real data reconstruction using such a trained network. We experimented with a penalized weighted least-squares (PWLS) cost function for low dose CT reconstruction and tested on data from a practical dental CT. Very encouraging results with great noise reductions are obtained.

physics.med-ph

Comparison of projection domain, image domain, and comprehensive deep learning for sparse-view X-ray CT image reconstruction

X-ray Computed Tomography (CT) imaging has been widely used in clinical diagnosis, non-destructive examination, and public safety inspection. Sparse-view (sparse view) CT has great potential in radiation dose reduction and scan acceleration. However, sparse view CT data is insufficient and traditional reconstruction results in severe streaking artifacts. In this work, based on deep learning, we compared image reconstruction performance for sparse view CT reconstruction with projection domain network, image domain network, and comprehensive network combining projection and image domains. Our study is executed with numerical simulated projection of CT images from real scans. Results demonstrated deep learning networks can effectively reconstruct rich high frequency structural information without streaking artefact commonly seen in sparse view CT. A comprehensive network combining deep learning in both projection domain and image domain can get best results.

physics.med-ph