SearcharxivSearch

arXiv subjects

Lei Tian

Publications and source records attributed to Lei Tian.

At least 19 recordsLinked to original sources

Site-specific Channel Modeling Based on Remote-Sensing Maps for 6G Space--Air--Ground Digital Twins

Site-specific channel models are essential for wireless digital twins of 6G space--air--ground communication systems. However, 3D maps are difficult to obtain over wide areas, which limits large-area site-specific channel modeling. To address this issue, this paper proposes a remote-sensing-based augmented ray-tracing channel modeling framework. The framework comprises a deterministic RT branch, a measurement-statistical branch, and an RT augmentation branch. To overcome the difficulty of acquiring large-area 3D maps, the deterministic RT branch reconstructs a 3D RT scene from satellite remote-sensing imagery and calibrates its electromagnetic material parameters using measured path loss. To provide the statistical parameters required for RT augmentation, the measurement-statistical branch establishes the marginal distributions and interparameter dependence models of the channel parameters. Specifically, a wideband UAV channel measurement campaign is conducted at 4.60 GHz, and a proposed multipath estimation method estimates the complex amplitudes, delays, and Doppler shifts of the measured multipath. To bridge the gap between RT predictions and measurements, the RT augmentation branch organizes the RT multipath into LoS, LoS-tail, and NLoS components, generates additional short-delay LoS-tail paths, and reallocates the component and path powers according to the measurement-derived statistics while preserving the total RT received power. The validation results show that the proposed framework reduces the path loss RMSE from 5.45 to 4.35 dB and, relative to calibrated RT, decreases the RMS delay spread and normalized Doppler spread RMSEs by 53.03 and 26.48, respectively. The proposed framework provides a site-specific channel modeling approach for 6G space--air--ground digital-twin studies.

eess.SP

Point Spread Function Engineering Using Implicit Neural Representations

Point spread function (PSF) engineering through pupil plane modulation is a technique used in microscopy to achieve specific imaging properties, such as depth encoding or extended depth of field. Existing PSF design methods often rely on extensive domain knowledge and task-specific basis functions, making it difficult to generalize across different applications. We treat the PSF engineering task as a phase retrieval problem and propose a neural field pupil design method that optimizes a phase profile for any arbitrary, user-defined 3D PSF distribution. This provides a flexible framework for 3D PSF engineering for various applications with implicit regularization that proves robust to initialization compared to pixel-wise optimization methods

physics.optics

The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy

The growing ability of large language models and vision-language models to jointly interpret and reason over images and text is reshaping medical imaging AI, moving it from task-specific predictors toward autonomous agents that perceive, reason, plan, remember, and act in clinical environments. This survey departs from the capability-first perspective of existing literature and instead begins from clinical deployment, asking what tasks, contamination-resistant benchmarks, and interactive training environments are required before medical agents can be trusted in practice. Medical agents are formalized as sequential decision-making systems under partial observability, together with a three-level autonomy taxonomy spanning assisted, cooperative, and fully autonomous operation. The field is organized along a unified scaling spine consisting of framework scaling, capability scaling, and environment scaling. Within this framework, clinical environment scaling, the integration of tools, data, and clinical gyms, is identified as the most actionable yet underexplored direction for agents operating in PACS, EHR, and FHIR ecosystems. Clinical self-evolution, where agents improve through interaction with their environments rather than parameter scaling alone, is further positioned as a key research frontier, drawing insights from self-improving agents, agent gyms, and test-time compute scaling. Applications across radiology, pathology, ophthalmology, and hospital workflows are examined together with deployment challenges including hallucination, cascading failures, and fairness. By consolidating more than 300 references, with particular emphasis on advances from 2025 to 2026, this survey provides a roadmap toward trustworthy, self-improving medical imaging systems for real clinical practice.

cs.AI

DeepFilters: Scattering-Aware Pupil Engineering with Learned Digital Filter Reconstruction for Extended Depth of Field Microscopy

Extended depth of field microscopy encodes axial information into a single acquisition through engineered point spread functions, but conventional and deep optics approaches are subject to degradation in scattering tissue. We introduce DeepFilters, a scattering-aware deep optics framework that jointly optimizes a parameterized pupil filter and a digital-filter-based reconstruction network through a calibrated differentiable forward model to achieve broad generalization without retraining. Incorporating empirical scattering kernels, physics-guided regularization, and a hybrid genetic-gradient initialization strategy, DeepFilters extends the PSF from 16 micron to >400 micron in clear media and enables signal recovery beyond 120 micron deep in biological tissues, validated across fixed brain slices and sea urchin embryos.

physics.optics

Coordinate-conditioned Deconvolution for Scalable Spatially Varying High-Throughput Imaging

Wide-field fluorescence microscopy with compact optics often suffers from spatially varying blur due to field-dependent aberrations, vignetting, and sensor truncation, while finite sensor sampling imposes an inherent trade-off between field of view (FOV) and resolution. Computational Miniaturized Mesoscope (CM2) alleviate the sampling limit by multiplexing multiple sub-views onto a single sensor, but introduce view crosstalk and a highly ill-conditioned inverse problem compounded by spatially variant point spread functions (PSFs). Prior learning-based spatially varying (SV) reconstruction methods typically rely on global SV operators with fixed input sizes, resulting in memory and training costs that scale poorly with image dimensions. We propose SV-CoDe (Spatially Varying Coordinate-conditioned Deconvolution), a scalable deep learning framework that achieves uniform, high-resolution reconstruction across a 6.5 mm FOV. Unlike conventional methods, SV-CoDe employs coordinate-conditioned convolutions to locally adapt reconstruction kernels; this enables patch-based training that decouples parameter count from FOV size. SV-CoDe achieves the best image quality in both simulated and experimental measurements while requiring 10x less model size and 10x less training data than prior baselines. Trained purely on physics-based simulations, the network robustly generalizes to bead phantoms, weakly scattering brain slices, and freely moving C. elegans. SV-CoDe offers a scalable, physics-aware solution for correcting SV blur in compact optical systems and is readily extendable to a broad range of biomedical imaging applications.

eess.IV

Mid-Infrared Photothermal Relaxation Intensity Diffraction Tomography for Video-rate Volumetric Chemical Imaging

Three-dimensional molecular imaging of living cells is essential for unraveling cellular metabolism and response to therapies. However, existing volumetric methods, including fluorescence microscopy and quantitative phase imaging, either require fluorescent labels or lack chemical specificity. Mid-infrared (mid-IR) photothermal microscopy provides label-free spectroscopic contrast with sub-micrometer resolution but is limited by slow acquisition rates, precluding 3D live-cell studies. Here, we present a photothermal relaxation intensity diffraction tomography (PRIDT) system that encodes mid-IR absorption induced refractive index change via a photothermal relaxation scheme and recovers it through intensity diffraction tomography. PRIDT achieves video-rate volumetric chemical imaging with up to 15 Hz per wavelength and offers lateral and axial resolutions of 264 nm and 1.12 um over a volumetric field of view of 50x50x10 um3. We showcase high-speed PRIDT imaging of protein and lipid metabolism in ovarian cancer cells and lipid-droplet dynamics in live cells. PRIDT opens new avenues for rapid, quantitative, three-dimensional molecular imaging in living systems.

physics.optics

Dual-wavelength Fourier Ptychographic Topography

We introduce a dual-wavelength Fourier ptychographic topography (FPT) method that extends the lambda/2 height-range limit of single-wavelength FPT. By reconstructing complex fields at two illumination wavelengths and exploiting their phase difference, the method achieves an effective synthetic wavelength lambda_s and an unambiguous range of lambda_s/2 without reducing lateral resolution. A noise-robust wrapped-number search is used to select per-pixel integer pairs (k1, k2), and a global refinement with circular TV regularization and soft bounds improves stability and preserves height discontinuities. The approach is validated through rigorous scattering-model-based simulations and experiments on structured silicon samples, demonstrating accurate height recovery in regimes where single-wavelength FPT exhibits phase wrapping. We analyze the limits of the FPT forward model and identify aspect ratio (AR) and phase modulation transfer function (ph-MTF) as key predictors of reconstruction fidelity. Simulations and experiments show that increasing AR beyond a practical threshold causes loss of high-frequency phase transfer and destabilizes dual-wavelength unwrapping. Within this AR range, dual-wavelength FPT provides robust, high-resolution topography suitable for semiconductor and industrial metrology.

physics.optics

A Comprehensive Survey of 3GPP Release 19 ISAC Channel Modeling: From Empirical Features to Unified Methodology and Standardized Simulator

Integrated Sensing and Communication (ISAC) has been identified as a key 6G application by ITU and 3GPP. Channel measurement and modeling is a prerequisite for ISAC system design and has attracted widespread attention from both academia and industry. 3GPP Release 19 initiated the ISAC channel study item in December 2023 and finalized its modeling specification in May 2025 after extensive technical discussions. However, a comprehensive survey that provides a systematic overview,from empirical channel features to modeling methodologies and standardized simulators,remains unavailable. In this paper, the key requirements and challenges in ISAC channel research are first analyzed, followed by a structured overview of the standardization workflow throughout the 3GPP Release 19 process. Then, critical aspects of ISAC channels, including physical objects, target channels, and background channels, are examined in depth, together with additional features such as spatial consistency, environment objects, Doppler characteristics, and shared clusters, supported by measurement-based analysis. To establish a unified ISAC channel modeling framework, an Extended Geometry-based Stochastic Model (E-GBSM) is proposed, incorporating all the aforementioned ISAC channel characteristics. Finally, a standardized simulator is developed based on E-GBSM, and a two-phase calibration procedure aligned with 3GPP Release 19 is conducted to validate both the model and the simulator, demonstrating close agreement with industrial reference results. Overall, this paper provides a systematic survey of 3GPP Release 19 ISAC channel standardization and offers insights into best practices for new feature characterization, unified modeling methodology, and standardized simulator implementation, which can effectively supporting ISAC technology evaluation and future 6G standardization.

eess.SP

DCL-SE: Dynamic Curriculum Learning for Spatiotemporal Encoding of Brain Imaging

High-dimensional neuroimaging analyses for clinical diagnosis are often constrained by compromises in spatiotemporal fidelity and by the limited adaptability of large-scale, general-purpose models. To address these challenges, we introduce Dynamic Curriculum Learning for Spatiotemporal Encoding (DCL-SE), an end-to-end framework centered on data-driven spatiotemporal encoding (DaSE). We leverage Approximate Rank Pooling (ARP) to efficiently encode three-dimensional volumetric brain data into information-rich, two-dimensional dynamic representations, and then employ a dynamic curriculum learning strategy, guided by a Dynamic Group Mechanism (DGM), to progressively train the decoder, refining feature extraction from global anatomical structures to fine pathological details. Evaluated across six publicly available datasets, including Alzheimer's disease and brain tumor classification, cerebral artery segmentation, and brain age prediction, DCL-SE consistently outperforms existing methods in accuracy, robustness, and interpretability. These findings underscore the critical importance of compact, task-specific architectures in the era of large-scale pretrained networks.

cs.CV

Transfer-Function Approach to Substrate-Enhanced Diffraction Tomography

Forward and backward scattering provide complementary volumetric and interfacial information, yet conventional three-dimensional (3D) imaging typically accesses only one. In this Letter, we present a substrate-enhanced diffraction tomography approach that simultaneously recovers both channels under multi-angle epi-illumination.This geometry captures one forward- and two backward-scattering bands in axially symmetric Fourier regions, where their complementary coverage enables phase-absorption separation in a non-Hermitian spectrum. Explicit 3D transfer functions are derived for both channels, and an axial Kramers-Kronig relation is established to incorporate substrate-induced boundary conditions in a unified framework. Our results establish a label-free, high-resolution 3D imaging modality that surpasses the limits of existing methods.

physics.optics

Theoretical Analysis of Near-Field MIMO Channel Capacity and Mid-Band Experimental Validation

With the increase of multiple-input-multiple-output (MIMO) array size and carrier frequency, near-field MIMO communications will become crucial in 6G wireless networks. Due to the increase of MIMO near-field range, the research of near-field MIMO capacity has aroused wide interest. In this paper, we focus on the theoretical analysis and empirical study of near-field MIMO capacity. First, the near-field channel model is characterized from the electromagnetic information perspective. Second, with the uniform planar array (UPA), the channel capacity based on effective degree of freedom (EDoF) is analyzed theoretically, and the closed-form analytical expressions are derived in detail. Finally, based on the numerical verification of near-field channel measurement experiment at 13 GHz band, we reveal that the channel capacity of UPA-type MIMO systems decreases continuously with the communication distance increasing. It can be observed that the near-field channel capacity gain is relatively obvious when large-scale MIMO is adopted at both receiving and transmitter ends, but the near-field channel capacity gain may be limited in the actual communication system with the small antenna array at receiving end. This work will give some reference to the near-field communication systems.

eess.SP

Astra: Toward General-Purpose Mobile Robots via Hierarchical Multimodal Learning

Modern robot navigation systems encounter difficulties in diverse and complex indoor environments. Traditional approaches rely on multiple modules with small models or rule-based systems and thus lack adaptability to new environments. To address this, we developed Astra, a comprehensive dual-model architecture, Astra-Global and Astra-Local, for mobile robot navigation. Astra-Global, a multimodal LLM, processes vision and language inputs to perform self and goal localization using a hybrid topological-semantic graph as the global map, and outperforms traditional visual place recognition methods. Astra-Local, a multitask network, handles local path planning and odometry estimation. Its 4D spatial-temporal encoder, trained through self-supervised learning, generates robust 4D features for downstream tasks. The planning head utilizes flow matching and a novel masked ESDF loss to minimize collision risks for generating local trajectories, and the odometry head integrates multi-sensor inputs via a transformer encoder to predict the relative pose of the robot. Deployed on real in-house mobile robots, Astra achieves high end-to-end mission success rate across diverse indoor environments.

cs.RO

CCL-LGS: Contrastive Codebook Learning for 3D Language Gaussian Splatting

Recent advances in 3D reconstruction techniques and vision-language models have fueled significant progress in 3D semantic understanding, a capability critical to robotics, autonomous driving, and virtual/augmented reality. However, methods that rely on 2D priors are prone to a critical challenge: cross-view semantic inconsistencies induced by occlusion, image blur, and view-dependent variations. These inconsistencies, when propagated via projection supervision, deteriorate the quality of 3D Gaussian semantic fields and introduce artifacts in the rendered outputs. To mitigate this limitation, we propose CCL-LGS, a novel framework that enforces view-consistent semantic supervision by integrating multi-view semantic cues. Specifically, our approach first employs a zero-shot tracker to align a set of SAM-generated 2D masks and reliably identify their corresponding categories. Next, we utilize CLIP to extract robust semantic encodings across views. Finally, our Contrastive Codebook Learning (CCL) module distills discriminative semantic features by enforcing intra-class compactness and inter-class distinctiveness. In contrast to previous methods that directly apply CLIP to imperfect masks, our framework explicitly resolves semantic conflicts while preserving category discriminability. Extensive experiments demonstrate that CCL-LGS outperforms previous state-of-the-art methods. Our project page is available at https://epsilontl.github.io/CCL-LGS/.

cs.CV

Whitened Score Diffusion: A Structured Prior for Imaging Inverse Problems

Conventional score-based diffusion models (DMs) may struggle with anisotropic Gaussian diffusion processes due to the required inversion of covariance matrices in the denoising score matching training objective \cite{vincent_connection_2011}. We propose Whitened Score (WS) diffusion models, a novel framework based on stochastic differential equations that learns the Whitened Score function instead of the standard score. This approach circumvents covariance inversion, extending score-based DMs by enabling stable training of DMs on arbitrary Gaussian forward noising processes. WS DMs establish equivalence with flow matching for arbitrary Gaussian noise, allow for tailored spectral inductive biases, and provide strong Bayesian priors for imaging inverse problems with structured noise. We experiment with a variety of computational imaging tasks using the CIFAR, CelebA ($64\times64$), and CelebA-HQ ($256\times256$) datasets and demonstrate that WS diffusion priors trained on anisotropic Gaussian noising processes consistently outperform conventional diffusion priors based on isotropic Gaussian noise. Our code is open-sourced at \href{https://github.com/jeffreyalido/wsdiffusion}{\texttt{github.com/jeffreyalido/wsdiffusion}}.

eess.IV

Empirical Study on Near-Field and Spatial Non-Stationarity Modeling for THz XL-MIMO Channel in Indoor Scenario

Terahertz (THz) extremely large-scale MIMO (XL-MIMO) is considered a key enabling technology for 6G and beyond due to its advantages such as wide bandwidth and high beam gain. As the frequency and array size increase, users are more likely to fall within the near-field (NF) region, where the far-field plane-wave assumption no longer holds. This also introduces spatial non-stationarity (SnS), as different antenna elements observe distinct multipath characteristics. Therefore, this paper proposes a THz XL-MIMO channel model that accounts for both NF propagation and SnS, validated using channel measurement data. In this work, we first conduct THz XL-MIMO channel measurements at 100 GHz and 132 GHz using 301- and 531-element ULAs in indoor environments, revealing pronounced NF effects characterized by nonlinear inter-element phase variations, as well as element-dependent delay and angle shifts. Moreover, the SnS phenomenon is observed, arising not only from blockage but also from inconsistent reflection or scattering. Based on these observations, a hybrid NF channel modeling approach combining the scatterer-excited point-source model and the specular reflection model is proposed to capture nonlinear phase variation. For SnS modeling, amplitude attenuation factors (AAFs) are introduced to characterize the continuous variation of path power across the array. By analyzing the statistical distribution and spatial autocorrelation properties of AAFs, a statistical rank-matching-based method is proposed for their generation. Finally, the model is validated using measured data. Evaluation across metrics such as entropy capacity, condition number, spatial correlation, channel gain, Rician K-factor, and RMS delay spread confirms that the proposed model closely aligns with measurements and effectively characterizes the essential features of THz XL-MIMO channels.

eess.SP

A Unified Deterministic Channel Model for Multi-Type RIS with Reflective, Transmissive, and Polarization Operations

Reconfigurable Intelligent Surface (RIS) technologies have been considered as a promising enabler for 6G, enabling advantageous control of electromagnetic (EM) propagation. RIS can be categorized into multiple types based on their reflective/transmissive modes and polarization control capabilities, all of which are expected to be widely deployed in practical environments. A reliable RIS channel model is essential for the design and development of RIS communication systems. While deterministic modeling approaches such as ray-tracing (RT) offer significant benefits, a unified model that accommodates all RIS types is still lacking. This paper addresses this gap by developing a high-precision deterministic channel model based on RT, supporting multiple RIS types: reflective, transmissive, hybrid, and three polarization operation modes. To achieve this, a unified EM response model for the aforementioned RIS types is developed. The reflection and transmission coefficients of RIS elements are derived using a tensor-based equivalent impedance approach, followed by calculating the scattered fields of the RIS to establish an EM response model. The performance of different RIS types is compared through simulations in typical scenarios. During this process, passive and lossless constraints on the reflection and transmission coefficients are incorporated to ensure fairness in the performance evaluation. Simulation results validate the framework's accuracy in characterizing the RIS channel, and specific cases tailored for dual-polarization independent control and polarization rotating RISs are highlighted as insights for their future deployment. This work can be helpful for the evaluation and optimization of RIS-enabled wireless communication systems.

eess.SP

Uncertainty-Aware Large Language Models for Explainable Disease Diagnosis

Explainable disease diagnosis, which leverages patient information (e.g., signs and symptoms) and computational models to generate probable diagnoses and reasonings, offers clear clinical values. However, when clinical notes encompass insufficient evidence for a definite diagnosis, such as the absence of definitive symptoms, diagnostic uncertainty usually arises, increasing the risk of misdiagnosis and adverse outcomes. Although explicitly identifying and explaining diagnostic uncertainties is essential for trustworthy diagnostic systems, it remains under-explored. To fill this gap, we introduce ConfiDx, an uncertainty-aware large language model (LLM) created by fine-tuning open-source LLMs with diagnostic criteria. We formalized the task and assembled richly annotated datasets that capture varying degrees of diagnostic ambiguity. Evaluating ConfiDx on real-world datasets demonstrated that it excelled in identifying diagnostic uncertainties, achieving superior diagnostic performance, and generating trustworthy explanations for diagnoses and uncertainties. To our knowledge, this is the first study to jointly address diagnostic uncertainty recognition and explanation, substantially enhancing the reliability of automatic diagnostic systems.

cs.CL

High-Resolution Multipath Angle Estimation Based on Power-Angle-Delay Profile for Directional Scanning Sounding

Directional scanning sounding (DSS) has become widely adopted for high-frequency channel measurements because it effectively compensates for severe path loss. However, the resolution of existing multipath component (MPC) angle estimation methods is constrained by the DSS angle sampling interval. Therefore, this communication proposes a high-resolution MPC angle estimation method based on power-angle-delay profile (PADP) for DSS. By exploiting the mapping relationship between the power difference of adjacent angles in the PADP and MPC offset angle, the resolution of MPC angle estimation is refined, significantly enhancing the accuracy of MPC angle and amplitude estimation without increasing measurement complexity. Numerical simulation results demonstrate that the proposed method reduces the mean squared estimation errors of angle and amplitude by one order of magnitude compared to traditional omnidirectional synthesis methods. Furthermore, the estimation errors approach the Cram\'er-Rao Lower Bounds (CRLBs) derived for wideband DSS, thereby validating its superior performance in MPC angle and amplitude estimation. Finally, experiments conducted in an indoor scenario at 37.5 GHz validate the excellent performance of the proposed method in practical applications.

eess.SP