SearcharxivSearch

arXiv subjects

Jie Deng

Publications and source records attributed to Jie Deng.

At least 19 recordsLinked to original sources

Physics-Assisted Deep Learning Denoising for Stabilized IMPULSED dMRI Microenvironment Parameter Fitting

Diffusion-weighted MRI (dMRI) is a powerful tool for quantifying cellular microenvironment parameters. This study proposes a physics-assisted deep learning (DL)-based denoising framework designed to enhance dMRI signal quality and improve the robustness of subsequent biophysical model fitting. A dataset of paired noise-free and Rician-noise-corrupted dMRI signals was generated using the IMPULSED-dMRI signal model. Three denoising architectures were evaluated: Convolutional Neural Networks (CNN), Multilayer Perceptron (MLP), and Long Short-Term Memory (LSTM) networks. Denoised signals were then fitted to estimate cell diameter $d$, intracellular volume fraction $V_{\mathrm{in}}$, and extracellular apparent diffusion coefficient $D_\mathrm{ex}$. DL-based processing substantially improved dMRI signal denoising. The MLP and LSTM achieved similar performance, with the LSTM slightly better overall, and both outperformed the CNN. In the subsequent model fitting step, the LSTM produced modest reductions in parameter MAE. The dominant benefit was fitting stabilization, with the overall fitting failure rate reduced from 57.6\% to 17.7\%. The proposed framework improves dMRI signal quality and stabilizes subsequent IMPULSED-based microenvironmental parameter fitting.

physics.med-ph

An integrated diffusion-weighted imaging processing and interpretation platform for MR-guided radiotherapy

Background: Magnetic resonance imaging-guided linear accelerators (MR-Linacs) allow diffusion-weighted imaging (DWI) to be acquired at every treatment fraction, but converting these low-signal-to-noise-ratio acquisitions into clinical decisions requires both reliable quantitative processing and an interpretation that reconciles a scattered and often contradictory literature. Purpose: To describe and evaluate an integrated, web-based platform that carries raw MR-Linac DWI to a structured, literature-grounded clinical interpretation, and to assess its retrieval-augmented generation (RAG) interpretation module by independent expert rating. Methods: The platform couples a deep-learning processing pipeline, comprising distortion correction, denoising, and intravoxel incoherent motion (IVIM)/apparent diffusion coefficient (ADC) fitting, with longitudinal region-of-interest analysis and a RAG interpretation agent. The agent reasons over a two-layer knowledge base of curated publications (a structured catalog index plus line-indexed full text), delegates arithmetic to deterministic tools, and is designed to trace each statement to a source document, section, and line range. One medical physicist and one physician independently rated the agent's reports for nine longitudinal glioblastoma cases on a 1-5 scale across three metrics: clinical-reasoning soundness, literature-citation quality, and overall clinical utility. Results: Across 54 ratings, the pooled mean was 4.65 +/- 0.80, with 93% of ratings >= 4; metric means were 4.6 (reasoning), 4.5 (citation), and 4.8 (utility), and raters agreed within one point on 85% of paired ratings. Conclusions: A single platform can integrate MR-Linac DWI post-processing with traceable, expert-evaluated clinical interpretation, while highlighting the safeguards needed to verify LLM-generated reasoning in radiation oncology.

physics.med-ph

Integrated Heat and Power System Scheduling with Continuous-Time Thermal Dynamics via Bernstein-Galerkin Optimization

Coordinated scheduling of district heating networks (DHNs) and electric power systems can improve operational flexibility and reduce costs by exploiting thermal inertia. Most existing formulations rely on simplified discrete-time DHN models, which may inadequately represent continuous spatiotemporal thermal dynamics and can lead to biased flexibility estimation and suboptimal schedules. In this paper, an integrated heat and power system scheduling framework that explicitly incorporates the continuous-time thermal dynamics of DHNs is proposed. A Bernstein-Galerkin transform method is developed to convert the underlying partial-differential thermal-dynamics constraints into a finite set of algebraic constraints, enabling tractable optimization while retaining dynamic fidelity. The resulting model transforms the original infinite-dimensional variational problem into a finite-dimensional coefficient optimization that can be solved using optimization solvers. Compared with conventional discretization approaches, the proposed method provides a more accurate representation of thermal dynamics and yields schedules with improved economic performance and reliability.

eess.SY

Distribution-free false-alarm calibration and chance-corrected spatial evaluation for industrial anomaly detection

Studies of industrial visual inspection commonly report the area under the receiver operating characteristic curve (AUROC) and the overlap between anomaly maps and defect masks. Neither measure specifies the false-alarm rate at a selected threshold, while recurrent defect locations and mask geometry can inflate overlap. We combine a distribution-free upper tolerance threshold with a paired-minus-crossed spatial test. This test compares each detector's score-contributing locations with the matched defect mask and with masks from other images; the difference in rates defines spatial-evidence lift relative to the empirical chance-overlap rate. We evaluate three detectors on 120 point-defect images from three ISP-AD modalities and three fixed data splits. Of 378 alarms, 230 overlap the matched mask. Paired and crossed rates are nevertheless similar in eight of nine detector--modality cells; only DINOv2--ASM has a positive 95\% bootstrap lower bound (lift 0.259, 95\% interval 0.159--0.347). On the independent Magnetic Tile Defect dataset, the same analysis gives lifts of 0.203 (0.169--0.236) for Wide ResNet-50 (WRN50) patch memory and 0.231 (0.202--0.262) for Vision Transformer B/16 (ViT-B/16) patch memory, with one-sided permutation $p=10^{-5}$ for both. When crossed masks are restricted to the same defect class, the lifts remain 0.185 and 0.210. Exact sample planning shows that, with 150 calibration normals, a 95\%-confidence distribution-free claim is supported only for target false-positive rates of 1.98\% or higher; a 1\% target requires at least 299 normals. The results support reporting operating-point performance and chance-corrected spatial evidence alongside AUROC and raw mask overlap.

cs.CV

When Can Fraud Operations Authorize Automation? A Decision-Support Framework for Fresh Audit Evidence and Review Workload

Fraud operations must allocate events among automatic approval, analyst review, and automatic blocking even though the labels needed to evaluate these actions are selective and delayed. Predictive scores order cases, but they do not show whether the evidence is current and representative enough to delegate an action to the model. We develop freshness-constrained audit capacity (FCAC), a decision-support framework that treats automation as an authorization decision constrained by action risk, evidence freshness, and shared review capacity. It evaluates candidate action regions from mature randomized audits and a prespecified temporal allowance. Supported regions are automated; unsupported regions remain in review. The resulting decision record reports evidence age, audit demand, total review workload, value exposure, and compatible temporal change. We show that current action risk is unidentified without restricting unobserved label evolution. Under representative randomized audits, label-independent evidence windows, and a prespecified condition linking historical and current action risk, we derive simultaneous finite-sample control of unsafe authorization. Chronological evaluations with simulated audits on IEEE-CIS, ULB-Worldline, and Elliptic++ yield zero-drift automation rates of 84.4%, 67.4%, and 81.3%, with total review workloads of 24.1%, 46.0%, and 43.1%. The experiments reveal an audit-capacity trade-off: sparse auditing delays authorization, whereas intensive auditing eventually increases workload. A separately specified BAF stress test further indicates that fallback thresholds must reflect candidate-specific evidence rather than a common fraction of the risk limit. These findings identify audit freshness and analyst capacity as joint design considerations for fraud decision support.

cs.LG

IR275K: A Benchmark for Infrared Multi-Frame Super-Resolution Toward Efficient Remote Sensing

Efficient processing is becoming increasingly important in infrared remote sensing, where satellite constellations produce large volumes of observations under constrained detector resolution, power, and downlink bandwidth. Multi-frame super-resolution (MFSR) offers a software-based route to spatial enhancement, but its evaluation in infrared sensing remains fragmented across private datasets and ad-hoc protocols. Existing benchmarks do not explicitly capture the thermal contrast, sensor noise, weak texture, and platform-induced frame-to-frame variation that characterize infrared video. We introduce IR275K, a curated benchmark containing 594 infrared video sequences and 275,196 frames. It provides sequence-level train/validation/test splits and a reproducible X4 evaluation protocol. As an initial architectural probe, we further evaluate CGMamba, a lightweight state-space model with 10.90M parameters and 112.14G FLOPs. CGMamba combines 2D rotary position encoding (2D~RoPE) with center-guided cross-Mamba (CGCM) fusion for implicit multi-frame reconstruction. It achieves 33.19dB PSNR, outperforming infrared single-image super-resolution references by 0.35--0.52~dB at substantially lower computational cost. Ablation results show that removing 2D~RoPE from CGCM causes a 1.53dB drop and severe grid-like artifacts. This indicates that explicit spatial anchoring is critical for stabilizing SSM-based cross-frame gating under infrared conditions. IR275K provides a reproducible foundation for accuracy--efficiency evaluation of infrared MFSR methods, while the architectural analysis offers a concrete starting point for spatially aware SSM design under resource-constrained infrared sensing. Dataset and evaluation resources are available at: https://github.com/InfraRecon7/IR275K.

cs.CV

Investigating the Uncertainty of Cellular Microenvironment Parameter Estimations via Diffusion MRI Cytometry

This study aims to identify cell microenvironment parameters that can be robustly estimated from IMPULSED diffusion MRI signals and to develop a reliable mapping-based estimation framework. Diffusion MRI signals were simulated using the established IMPULSED model with one pulsed gradient spin echo sequence and two oscillating gradient spin echo sequences at different frequencies. Five cellular parameters were considered: cell diameter ($d$), intracellular diffusion coefficient ($D_{in}$), intracellular volume fraction ($V_{in}$), extracellular diffusion coefficient ($D_{ex}$), and the frequency-dependent slope of $D_{ex}$ ($\beta_{ex}$). Parameter uncertainty was quantified using Jacobian-based sensitivity analysis at an SNR of 30, representing clinically achievable conditions on a 1.5T MRI scanner. To enable direct parameter mapping, signals were logarithmically transformed, reduced in dimension using principal component analysis, and then used to estimate parameters with linear regression, fourth-order polynomial regression, and a fully connected four-layer neural network. Model validation was performed in vitro using MC38 cell lines. Uncertainty analysis identified $d$, $V_{in}$, and $D_{ex}$ as robustly derivable parameters, each with relative uncertainty below 1.0. Among the tested models, the four-layer neural network performed best, with mean absolute errors of 1.7 $\mu$m for $d$, 5.06% for $V_{in}$, and 0.28 $\mu$m$^2$/ms for $D_{ex}$. In vitro validation showed a 6.7% error in cell diameter estimation. These results demonstrate that IMPULSED dMRI can support robust estimation of key cell microenvironment parameters and provide a practical framework for noninvasive assessment of tumor microenvironment changes during radiation therapy response monitoring.

physics.med-ph

EvoDrive: Pareto Evolution for Safety-Critical Autonomous Driving via Self-Improving LLM Agents

Generating safety-critical scenarios is essential for validating and improving autonomous driving systems, yet it inherently requires maximizing adversariality to expose failures while preserving realism. Existing methods usually manage this trade-off with handcrafted heuristics, confining generation to known priors and overlooking underexplored patterns. While recent open-ended agentic evolution can push this limit, unconstrained general agents lack strict simulator grounding and tend to collapse the multi-objective tension into single-scalar maximization. Here we present EvoDrive, the first automated, LLM-based agentic evolution framework for multi-objective scenario generation. EvoDrive employs a simulator-grounded actor-critic architecture where a memory-driven actor iteratively proposes improvements to the generators and critics filter out implausible candidates, and a self-evolving world evaluator routes promising proposals to optimize simulation budgets. EvoDrive further maintains a Pareto archive of evaluated candidates to preserve diverse attack-realism trade-offs and guide future evolution via simulation feedback. Benchmark results on MetaDrive and CARLA show that EvoDrive not only significantly expands the Pareto frontier across various generators, but also produces valuable scenarios for policy training.

cs.AI

Thermal conductivity of seifertite and pyrite-type SiO$_2$: A comparative study

Thermal conductivity is a fundamental material property that plays a crucial role in understanding the dynamics and evolution of planetary interiors. Despite its importance, the thermal conductivity of seifertite and pyrite-type SiO$_2$ remains unknown. Here, we calculate the lattice thermal conductivities of seifertite and pyrite-type SiO$_2$ using the Green-Kubo method based on molecular dynamics (MD) simulations driven by two machine learning potentials (MLPs) constructed from the SCAN and PBEsol exchange-correlation functionals, with $\textit{ab initio}$-level accuracy. To demonstrate our methodology, we also compute thermal conductivities using the phonon quasiparticle approach for comparison. Overall, the Green-Kubo method predicts up to 119 % higher thermal conductivity with a temperature dependence close to $T^{-1}$, as it fully captures diffusion-like phonons at high temperatures that are missed by the phonon quasiparticle approach. The 19 % reduction in thermal conductivity across the phase transition from seifertite to the pyrite-type phase suggests the potential formation of a thermally insulating layer in the mantle of super-Earths.

cond-mat.mtrl-sci

Optimizing IMPULSED Acquisition Protocols for Clinical 3T Scanners Through Bayesian Experimental Design

To optimize diffusion MRI acquisition protocols for IMPULSED model at clinical 3T scanner using Bayesian experimental design, enabling accurate cellular-scale parameter estimation under realistic scan time and scanner hardware constraints. Expected Information Gain (EIG) was used as the optimization objective to maximize the information content of acquired measurements for IMPULSED model fitting. Bayesian optimization with Gaussian process surrogates efficiently searched the high-dimensional acquisition parameter space, including pulse types (PGSE, OGSEn1, and OGSEn2), diffusion times, and b-values. Optimized protocols were systematically evaluated against a heuristically designed baseline protocol through simulation studies assessing classification accuracy and parameter estimation performance across SNR levels of 5-40. Robustness to optimization assumptions was examined by varying prior distributions and assumed SNR. In-vivo validation was performed using canine tumor data acquired at 3T. The optimized protocol eliminated OGSEn2 acquisitions, concentrated measurements at high b-values, employing concurrently optimized diffusion timing. Compared to the baseline protocol, the optimized design achieved superior classification accuracy for distinguishing cell populations and reduced parameter estimation error across biologically relevant parameter ranges at various SNRs. Performance advantages were consistent across diverse optimization scenarios, demonstrating robustness to prior knowledge and noise assumptions. In-vivo parameter maps showed substantially improved quality and smoothness. Bayesian optimization substantially improves IMPULSED acquisition design for clinical 3T scanners. This principled, algorithm-agnostic framework enables accurate diffusion MRI cytometry under clinical constraints, with potential applications to tumor characterization and treatment monitoring.

physics.med-ph

Spatiotemporal Gaussian representation-based dynamic reconstruction and motion estimation framework for time-resolved volumetric MR imaging (DREME-GSMR)

Time-resolved volumetric MR imaging that reconstructs a 3D MRI within sub-seconds to resolve deformable motion is essential for motion-adaptive radiotherapy. Representing patient anatomy and associated motion fields as 3D Gaussians, we developed a spatiotemporal Gaussian representation-based framework (DREME-GSMR), which enables time-resolved dynamic MRI reconstruction from a pre-treatment 3D MR scan without any prior anatomical/motion model. DREME-GSMR represents a reference MRI volume and a corresponding low-rank motion model (as motion-basis components) using 3D Gaussians, and incorporates a dual-path MLP/CNN motion encoder to estimate temporal motion coefficients of the motion model from raw k-space-derived signals. Furthermore, using the solved motion model, DREME-GSMR can infer motion coefficients directly from new online k-space data, allowing subsequent intra-treatment volumetric MR imaging and motion tracking (real-time imaging). A motion-augmentation strategy is further introduced to improve robustness to unseen motion patterns during real-time imaging. DREME-GSMR was evaluated on the XCAT digital phantom, a physical motion phantom, and MR-LINAC datasets acquired from 6 healthy volunteers and 20 patients (with independent sequential scans for cross-evaluation). DREME-GSMR reconstructs MRIs of a ~400ms temporal resolution, with an inference time of ~10ms/volume. In XCAT experiments, DREME-GSMR achieved mean(s.d.) SSIM, tumor center-of-mass-error(COME), and DSC of 0.92(0.01)/0.91(0.02), 0.50(0.15)/0.65(0.19) mm, and 0.92(0.02)/0.92(0.03) for dynamic reconstruction/real-time imaging. For the physical phantom, the mean target COME was 1.19(0.94)/1.40(1.15) mm for dynamic/real-time imaging, while for volunteers and patients, the mean liver COME for real-time imaging was 1.31(0.82) and 0.96(0.64) mm, respectively.

physics.med-ph

Uniform estimates and Brezis-Merle type inequalities for the $k$-Hessian equation

In this paper, we prove a Brezis-Merle type inequality for $k$-convex functions vanishing on the boundary. As an application, we establish an Alexandrov-Bakelman-Pucci type estimate for the intermediate Hessian equation. Furthermore, we establish a concentration-compactness principle for the blow-up behavior of solutions to the mean field type $k$-Hessian equation.

math.AP

Adapting Segment Anything Model 3 for Concept-Driven Lesion Segmentation in Medical Images: An Experimental Study

Accurate lesion segmentation is essential in medical image analysis, yet most existing methods are designed for specific anatomical sites or imaging modalities, limiting their generalizability. Recent vision-language foundation models enable concept-driven segmentation in natural images, offering a promising direction for more flexible medical image analysis. However, concept-prompt-based lesion segmentation, particularly with the latest Segment Anything Model 3 (SAM3), remains underexplored. In this work, we present a systematic evaluation of SAM3 for lesion segmentation. We assess its performance using geometric bounding boxes and concept-based text and image prompts across multiple modalities, including multiparametric MRI, CT, ultrasound, dermoscopy, and endoscopy. To improve robustness, we incorporate additional prior knowledge, such as adjacent-slice predictions, multiparametric information, and prior annotations. We further compare different fine-tuning strategies, including partial module tuning, adapter-based methods, and full-model optimization. Experiments on 13 datasets covering 11 lesion types demonstrate that SAM3 achieves strong cross-modality generalization, reliable concept-driven segmentation, and accurate lesion delineation. These results highlight the potential of concept-based foundation models for scalable and practical medical image segmentation. Code and trained models will be released at: https://github.com/apple1986/lesion-sam3

eess.IV

Ab initio electronic conductivity of Fe-bearing post-perovskite

The electrical conductivity of high-pressure silicates profoundly influences the interior dynamics of rocky planets. Employing the Kubo-Greenwood formalism, we perform ab initio calculations of electronic conductivity in Fe-bearing post-perovskite under super-Earth mantle conditions, up to 4000 K and 500 GPa. Electronic structures are obtained via many-body perturbation theory, incorporating dynamical screening and correlations among localized Fe-3d orbitals. In contrast to (Fe,Mg)O, for which metallization has been reported at comparable conditions, our results indicate that post-perovskite with Earth-like Fe contents is unlikely to metallize in super-Earth mantles via band-gap closure, yielding negligible low-frequency conductivity. Any substantial conductivity would require non-electronic mechanisms, such as thermally activated small-polaron hopping, which fall beyond the scope of band conduction.

cond-mat.mtrl-sci

VisRefiner: Learning from Visual Differences for Screenshot-to-Code Generation

Screenshot-to-code generation aims to translate user interface screenshots into executable frontend code that faithfully reproduces the target layout and style. Existing multimodal large language models perform this mapping directly from screenshots but are trained without observing the visual outcomes of their generated code. In contrast, human developers iteratively render their implementation, compare it with the design, and learn how visual differences relate to code changes. Inspired by this process, we propose VisRefiner, a training framework that enables models to learn from visual differences between rendered predictions and reference designs. We construct difference-aligned supervision that associates visual discrepancies with corresponding code edits, allowing the model to understand how appearance variations arise from implementation changes. Building on this, we introduce a reinforcement learning stage for self-refinement, where the model improves its generated code by observing both the rendered output and the target design, identifying their visual differences, and updating the code accordingly. Experiments show that VisRefiner substantially improves single-step generation quality and layout fidelity, while also endowing models with strong self-refinement ability. These results demonstrate the effectiveness of learning from visual differences for advancing screenshot-to-code generation.

cs.CV

Beyond Rejection Sampling: Trajectory Fusion for Scaling Mathematical Reasoning

Large language models (LLMs) have made impressive strides in mathematical reasoning, often fine-tuned using rejection sampling that retains only correct reasoning trajectories. While effective, this paradigm treats supervision as a binary filter that systematically excludes teacher-generated errors, leaving a gap in how reasoning failures are modeled during training. In this paper, we propose TrajFusion, a fine-tuning strategy that reframes rejection sampling as a structured supervision construction process. Specifically, TrajFusion forms fused trajectories that explicitly model trial-and-error reasoning by interleaving selected incorrect trajectories with reflection prompts and correct trajectories. The length of each fused sample is adaptively controlled based on the frequency and diversity of teacher errors, providing richer supervision for challenging problems while safely reducing to vanilla rejection sampling fine-tuning (RFT) when error signals are uninformative. TrajFusion requires no changes to the architecture or training objective. Extensive experiments across multiple math benchmarks demonstrate that TrajFusion consistently outperforms RFT, particularly on challenging and long-form reasoning problems.

cs.CL

ConPress: Learning Efficient Reasoning from Multi-Question Contextual Pressure

Large reasoning models (LRMs) typically solve reasoning-intensive tasks by generating long chain-of-thought (CoT) traces, leading to substantial inference overhead. We identify a reproducible inference-time phenomenon, termed Self-Compression: when multiple independent and answerable questions are presented within a single prompt, the model spontaneously produces shorter reasoning traces for each question. This phenomenon arises from multi-question contextual pressure during generation and consistently manifests across models and benchmarks. Building on this observation, we propose ConPress (Learning from Contextual Pressure), a lightweight self-supervised fine-tuning approach. ConPress constructs multi-question prompts to induce self-compression, samples the resulting model outputs, and parses and filters per-question traces to obtain concise yet correct reasoning trajectories. These trajectories are directly used for supervised fine-tuning, internalizing compressed reasoning behavior in single-question settings without external teachers, manual pruning, or reinforcement learning. With only 8k fine-tuning examples, ConPress reduces reasoning token usage by 59% on MATH500 and 33% on AIME25, while maintaining competitive accuracy.

cs.CL

Optimized Slice-Phase Control of Mirror Pulse in Cold-Atom Interferometry with Finite Response Time

Atom interferometers require both high efficiency and robust performance in their mirror pulses under experimental inhomogeneities. In this work, we demonstrated that quantum optimal control designed mirror pulse significantly enhance interferometer performance by using novel adaptive sliced structure. Using gradient ascent pulse engineering (GRAPE), optimized mirror pulse for a Mach-Zehnder light-pulse atom interferometer was designed by discretizing the control into non-uniform phase slices. This design broadened the tolerence to experimentally relevant variations in detuning $[-\Omega_0,\Omega_0]$ and Rabi frequency $[0.1\times\Omega_0,1.9\times\Omega_0]$ ($\Omega_0=2\pi\times25$ kHz), while maintaining high transfer efficiency even when the response-time delays up to 1.6 $\rm{\mu s}$. The optimized pulse was found to be robust to coupling inhomogeneity and velocity spread, offering a significant improvement in robustness over conventional pulse. The adaptive pulse slicing method provides a minimalist strategy that reduces experimental complexity while enhancing robustness and scalability, offering an innovative scheme for quantum optimal control in high precision atom interferometry.

quant-ph