SearcharxivSearch

arXiv subjects

Liu Li

Publications and source records attributed to Liu Li.

At least 19 recordsLinked to original sources

Wireless Physical-Layer Foundation Models: Architectures, Learning Paradigms, Applications, and Deployment

Foundation models, i.e., large neural networks pretrained on broad unlabeled data and adapted to many downstream tasks, have reshaped natural language processing and computer vision and are now being explored for the wireless physical layer. Wireless Physical-Layer Foundation Models (WPFMs) aim to learn transferable representations of signals such as channel state information (CSI), in-phase and quadrature (IQ) samples, and spectrograms so that a single pretrained backbone can support tasks ranging from channel estimation and prediction to localization and sensing while using limited task-specific data. This paper provides a dedicated review of WPFMs from learning design to practical deployment. We first establish the theoretical background, covering the neural architectures used for wireless signals, the self-supervised pretraining paradigms of masked modeling, contrastive learning, and generative pretraining, and the fine-tuning strategies that adapt pretrained models to downstream tasks. We then introduce a taxonomy that organizes existing models along five dimensions: architecture family, input modality and tokenization, pretraining objective, model scale and deployment target, and generalization capability. Building on this basis, we review applications across telecommunications, localization, and sensing, and, for each domain, analyze deployment feasibility by mapping model size to the memory, compute, and latency budgets of representative wireless hardware. Finally, we discuss model compression and efficient deployment, summarize the cross-cutting challenges, and outline open research directions. Our goal is to provide a reference that connects pretraining, architecture, and fine-tuning with the practical constraints of wireless systems.

eess.SP

A Dual-Track Framework for Template-Constrained LaTeX Conversion

With the increasing demands for advanced document conversion, mapping structured Markdown drafts into template-compliant formats like LaTeX remains a challenge. Existing approaches largely depend on either deterministic rule-based converters or pure end-to-end Large Language Model (LLM) generation. The former fails to correctly handle asset insertions and template-specific constraints, while the latter tends to induce semantic drift, leading to hallucinations that are difficult to debug. To address these limitations, we introduce a robust Dual-Track Framework that systematically decouples template formatting from document processing: an offline track extracts template constraints into a reusable manifest, while an online track implements a hybrid execution pipeline. This pipeline confines LLM usage exclusively to reasoning-intensive components (e.g., semantic metadata, bibliographic references, and complex visual/tabular layouts) while delegating rule-based engines for deterministic processing. Empirical evaluation across 7 LaTeX templates and 56 published research papers demonstrates that our method preserves better structural fidelity, satisfies diverse layout constraints, and achieves a higher compilation success rate compared to the previous baselines.

cs.CL

Minimal Sufficient Representations for Self-interpretable Deep Neural Networks

Deep neural networks (DNNs) achieve remarkable predictive performance but remain difficult to interpret, largely due to overparameterization that obscures the minimal structure required for interpretation. Here we introduce DeepIn, a self-interpretable neural network framework that adaptively identifies and learns the minimal representation necessary for preserving the full expressive capacity of standard DNNs. We show that DeepIn can correctly identify the minimal representation dimension, select relevant variables, and recover the minimal sufficient network architecture for prediction. The resulting estimator achieves optimal non-asymptotic error rates that adapt to the learned minimal dimension, demonstrating that recovering minimal sufficient structure fundamentally improves generalization error. Building on these guarantees, we further develop hypothesis testing procedures for both selected variables and learned representations, bridging deep representation learning with formal statistical inference. Across biomedical and vision benchmarks, DeepIn improves both predictive accuracy and interpretability, reducing error by up to 30% on real-world datasets while automatically uncovering human-interpretable discriminative patterns. Our results suggest that interpretability and statistical rigor can be embedded directly into deep architectures without sacrificing performance.

stat.ME

Ultracompact high-Q whispering gallery mode microresonator in a non-closed waveguide path

Integrated photonic circuits are foundational for versatile applications, where high-performance traveling-wave optical resonators are critical. Conventional whispering-gallery mode microresonators (WGMRs) confine light in closed-loop waveguide paths, thus inevitably occupy large footprints. Here, we report an ultracompact high loaded Q silicon photonic WGMR in an open curved path instead. By leveraging spatial mode multiplexing, low-loss mode converter-based photonic routers enable reentrant photon recycling in a single non-closed waveguide. The fabricated device achieves a measured loaded Q-factor of 1.78*10^5 at 1554.3 nm with a 1.05 nm free spectral range in a ultracompact footprint of 0.00137 mm^2-6*smaller than standard WGMRs while delivering 100*higher Q-factor than photonic crystal counterparts. This work pioneers dense integration of high-performance WGMR arrays through open-path mode recirculation.

physics.optics

Integrated Silicon Photonic Multichannel Optical Hybrid for Broadband Parallel Coherent Reception

We design and demonstrate a monolithically integrated silicon photonic multichannel optical hybrid for versatile broadband coherent reception, addressing the critical limitations of current wavelength multiplexed systems in scalability and power efficiency. The device combines a phase-compensated 90-degree optical hybrid with four robust three-stage Mach-Zehnder interferometer lattice filters, enabling 34-port functionality (two inputs and 32 outputs) for simultaneous analog and digital signal processing. Leveraging multimode interferometer designs,the chip achieves a broadband response with sub-dB passband uniformity across eight 200 GHz-spaced wavelength channels, while maintaining phase errors below 4 degrees over a 13.5 nm (1539-1552.5 nm) bandwidth with only 2.5 mW thermal tuning power.Experimentally, we validate its parallel-processing capability through RF channelizer reception (showing an average spurious-free dynamic range of 80.8 dB*Hz2/3 and image rejection ratio of 33.26 dB) and coherent optical communication (achieving 1.024 Tb/s data rate for 32-QAM signals with bit error rates far below the 20% SD-FEC threshold). The scheme enhances system performance with fully passive wavelength multiplexing integration, supporting high-fidelity uniformity and projecting scalability to 1.468 Tb/s. This work promises advancements in high-performance optoelectronic devices for next-generation AI-driven data centers and 5G-XG networks.

physics.optics

Miniaturized Computational Dispersion-Engineered Silicon Photonic Vernier Caliper Spectrometer

The development of miniaturized spectrometers for cost-effective mobile applications remains challenging, as small footprints fundamentally degrade bandwidth and resolution. Typically, achieving high resolution necessitates extended and sophisticated optical paths for spectral decorrelation. These restrict bandwidth both physically (through resonant wavelength periodicity constraints) and mathematically (due to resulting ill-conditioned large matrix factorizations). Here, we report a spectrometer using a computational dispersion-engineered silicon photonic Vernier caliper. This deterministic design enables periodicity-suppressed orthogonal measurements by nature, thus overcoming the bandwidth-resolution-footprint limit of current chip-scale spectrometers. Leveraging the dispersion-engineered Vernier subwavelength grating microrings and factorization-free matrix computation, a spectral resolution of 1.4 pm is achieved throughout a bandwidth of >160 nm with a footprint of <55*35 {\mu}m2 in a single detection channel,establishing the highest bandwidth-to-resolution-to-footprint ratio (>57 {\mu}m-2) demonstrated to date. Furthermore, broadband densely overlapped molecular absorption spectra of hydrogen cyanide are precisely measured, resolving 49 R- and P-branch lines with linewidths ranging from 15 to 86 pm which is fundamentally challenging for compressive sensing approaches. Our chip-scale spectrometer provides a new path toward precise and real-time multi-species spectral analysis and facilitates their commercialization.

physics.optics

Versatile and reconfigurable integrated silicon nitride photonic microresonator

Unlocking the full potential of integrated photonics requires versatile, multi-functional devices that can adapt to diverse application demands. However, confronting this challenge with conventional single-function resonators often results in tedious and complex systems. We present an elegant solution: a versatile and reconfigurable dual-polarization Si3N4 microresonator that represents a paradigm shift in on-chip photonic designs. Our device, based on a binary-star orbital architecture, can be dynamically reconfigured into three distinct topologies: a M\"obius-like microcavity, a Fabry-P\'erot resonator, and a microring resonator. This unprecedented functionality is enabled by a tunable balanced Mach-Zehnder interferometer that facilitates controllable mutual mode coupling of counterpropagating lights using a single control knob. We experimentally demonstrate that the device not only supports polarization-diverse operation on a compact footprint but also gives rise to a rich variety of physical phenomena, including a standing wave cavity, a traveling wave cavity, free spectral range multiplication, and the photonic pinning effect. These behaviors are accurately modeled using the Transfer Matrix Method and intuitively explained by Temporal Coupled Mode Theory. Our results underscore the profound potential for a chip-scale platform to realize reconfigurable reconstructive spectrometers and on-chip synthetic dimensions for topological physics.

physics.optics

Million-Q Dual-Polarization Micro-Fabry-Perot Resonators in Silicon Nitride Photonic Integrated Circuits

Miniaturized Fabry-Perot standing-wave resonators and whispering-gallery travelling wave resonators constitute foundational building blocks for photonic integrated circuits. While both architectures offer transformative potential through high quality factors and dual-polarization operation, integrated Fabry-Perot resonators face significant challenges in simultaneously achieving ultra-high Q-factors and broadband thermal tunability for fundamental transverse magnetic (TM0) and transverse electric (TE0) modes within a compact footprint-primarily due to polarization-dependent losses in conventional chip-scale reflectors. Here, we overcome this limitation by demonstrating an integrated silicon nitride dual-polarization micro-Fabry-Perot resonator with polarization-insensitive Sagnac loop reflectors and multimode waveguides to effectively suppress losses and enable high-performances for both fundamental transverse magnetic (TM0) and transverse electric (TE0) modes. The device achieves record loaded quality factors of 2.38*106 (TM0) and 3.48*105 (TE0) respectively and intrinsic quality factors will be even higher. Moreover, both two modes are tuned over the whole free spectral range of around 0.111 nm (TM0) and 0.112 nm (TE0) with the thermal tuning efficiencies of approximately 1.04 pm/mW (TM0) and 1.24 pm/mW (TE0). These advances establish a new benchmark for compact, high-performance dual-polarization resonators in optical sensors, nonlinear and integrated quantum photonics.

physics.optics

Topology Optimization in Medical Image Segmentation with Fast Euler Characteristic

Deep learning-based medical image segmentation techniques have shown promising results when evaluated based on conventional metrics such as the Dice score or Intersection-over-Union. However, these fully automatic methods often fail to meet clinically acceptable accuracy, especially when topological constraints should be observed, e.g., continuous boundaries or closed surfaces. In medical image segmentation, the correctness of a segmentation in terms of the required topological genus sometimes is even more important than the pixel-wise accuracy. Existing topology-aware approaches commonly estimate and constrain the topological structure via the concept of persistent homology (PH). However, these methods are difficult to implement for high dimensional data due to their polynomial computational complexity. To overcome this problem, we propose a novel and fast approach for topology-aware segmentation based on the Euler Characteristic ($\chi$). First, we propose a fast formulation for $\chi$ computation in both 2D and 3D. The scalar $\chi$ error between the prediction and ground-truth serves as the topological evaluation metric. Then we estimate the spatial topology correctness of any segmentation network via a so-called topological violation map, i.e., a detailed map that highlights regions with $\chi$ errors. Finally, the segmentation results from the arbitrary network are refined based on the topological violation maps by a topology-aware correction network. Our experiments are conducted on both 2D and 3D datasets and show that our method can significantly improve topological correctness while preserving pixel-wise segmentation accuracy.

eess.IV

Advances in Automated Fetal Brain MRI Segmentation and Biometry: Insights from the FeTA 2024 Challenge

Accurate fetal brain tissue segmentation and biometric analysis are essential for studying brain development in utero. The FeTA Challenge 2024 advanced automated fetal brain MRI analysis by introducing biometry prediction as a new task alongside tissue segmentation. For the first time, our diverse multi-centric test set included data from a new low-field (0.55T) MRI dataset. Evaluation metrics were also expanded to include the topology-specific Euler characteristic difference (ED). Sixteen teams submitted segmentation methods, most of which performed consistently across both high- and low-field scans. However, longitudinal trends indicate that segmentation accuracy may be reaching a plateau, with results now approaching inter-rater variability. The ED metric uncovered topological differences that were missed by conventional metrics, while the low-field dataset achieved the highest segmentation scores, highlighting the potential of affordable imaging systems when paired with high-quality reconstruction. Seven teams participated in the biometry task, but most methods failed to outperform a simple baseline that predicted measurements based solely on gestational age, underscoring the challenge of extracting reliable biometric estimates from image data alone. Domain shift analysis identified image quality as the most significant factor affecting model generalization, with super-resolution pipelines also playing a substantial role. Other factors, such as gestational age, pathology, and acquisition site, had smaller, though still measurable, effects. Overall, FeTA 2024 offers a comprehensive benchmark for multi-class segmentation and biometry estimation in fetal brain MRI, underscoring the need for data-centric approaches, improved topological evaluation, and greater dataset diversity to enable clinically robust and generalizable AI tools.

cs.CV

In-Context Learning Distillation for Efficient Few-Shot Fine-Tuning

We applied few-shot in-context learning on the OPT-1.3B model for the natural language inference task and employed knowledge distillation to internalize the context information, reducing model parameter from 1.3B to 125M and achieving a size reduction from 2.5GB to 0.25GB. Compared to using in-context learning alone on similarly sized models, this context distillation approach achieved a nearly 50% improvement in out-of-domain accuracy, demonstrating superior knowledge transfer capabilities over prompt-based methods. Furthermore, this approach reduced memory consumption by up to 60% while delivering a 20% improvement in out-of-domain accuracy compared to conventional pattern-based fine-tuning.

cs.CL

Metasurface-generated large and arbitrary analog convolution kernels for accelerated machine vision

In the rapidly evolving field of artificial intelligence, convolutional neural networks are essential for tackling complex challenges such as machine vision and medical diagnosis. Recently, to address the challenges in processing speed and power consumption of conventional digital convolution operations, many optical components have been suggested to replace the digital convolution layer in the neural network, accelerating various machine vision tasks. Nonetheless, the analog nature of the optical convolution kernel has not been fully explored. Here, we develop a spatial frequency domain training method to create arbitrarily shaped analog convolution kernels using an optical metasurface as the convolution layer, with its receptive field largely surpassing digital convolution kernels. By employing spatial multiplexing, the multiple parallel convolution kernels with both positive and negative weights are generated under the incoherent illumination condition. We experimentally demonstrate a 98.59% classification accuracy on the MNIST dataset, with simulations showing 92.63% and 68.67% accuracy on the Fashion-MNIST and CIFAR-10 datasets with additional digital layers. This work underscores the unique advantage of analog optical convolution, offering a promising avenue to accelerate machine vision tasks, especially in edge devices.

physics.optics

Universal Topology Refinement for Medical Image Segmentation with Polynomial Feature Synthesis

Although existing medical image segmentation methods provide impressive pixel-wise accuracy, they often neglect topological correctness, making their segmentations unusable for many downstream tasks. One option is to retrain such models whilst including a topology-driven loss component. However, this is computationally expensive and often impractical. A better solution would be to have a versatile plug-and-play topology refinement method that is compatible with any domain-specific segmentation pipeline. Directly training a post-processing model to mitigate topological errors often fails as such models tend to be biased towards the topological errors of a target segmentation network. The diversity of these errors is confined to the information provided by a labelled training set, which is especially problematic for small datasets. Our method solves this problem by training a model-agnostic topology refinement network with synthetic segmentations that cover a wide variety of topological errors. Inspired by the Stone-Weierstrass theorem, we synthesize topology-perturbation masks with randomly sampled coefficients of orthogonal polynomial bases, which ensures a complete and unbiased representation. Practically, we verified the efficiency and effectiveness of our methods as being compatible with multiple families of polynomial bases, and show evidence that our universal plug-and-play topology refinement network outperforms both existing topology-driven learning-based and post-processing methods. We also show that combining our method with learning-based models provides an effortless add-on, which can further improve the performance of existing approaches.

eess.IV

Reconfigurable unitary transformations of optical beam arrays

Spatial transformations of light are ubiquitous in optics, with examples ranging from simple imaging with a lens to quantum and classical information processing in waveguide meshes. Multi-plane light converter (MPLC) systems have emerged as a platform that promises completely general spatial transformations, i.e., a universal unitary. However until now, MPLC systems have demonstrated transformations that are far from general, e.g., converting from a Gaussian to Laguerre-Gauss mode. Here, we demonstrate the promise of an MLPC, the ability to impose an arbitrary unitary transformation that can be reconfigured dynamically. Specifically, we consider transformations on superpositions of parallel free-space beams arranged in an array, which is a common information encoding in photonics. We experimentally test the full gamut of unitary transformations for a system of two parallel beams and make a map of their fidelity. We obtain an average transformation fidelity of $0.85 \pm 0.03$. This high-fidelity suggests MPLCs are a useful tool implementing the unitary transformations that comprise quantum and classical information processing.

quant-ph

Stability and Generalizability in SDE Diffusion Models with Measure-Preserving Dynamics

Inverse problems describe the process of estimating the causal factors from a set of measurements or data. Mapping of often incomplete or degraded data to parameters is ill-posed, thus data-driven iterative solutions are required, for example when reconstructing clean images from poor signals. Diffusion models have shown promise as potent generative tools for solving inverse problems due to their superior reconstruction quality and their compatibility with iterative solvers. However, most existing approaches are limited to linear inverse problems represented as Stochastic Differential Equations (SDEs). This simplification falls short of addressing the challenging nature of real-world problems, leading to amplified cumulative errors and biases. We provide an explanation for this gap through the lens of measure-preserving dynamics of Random Dynamical Systems (RDS) with which we analyse Temporal Distribution Discrepancy and thus introduce a theoretical framework based on RDS for SDE diffusion models. We uncover several strategies that inherently enhance the stability and generalizability of diffusion models for inverse problems and introduce a novel score-based diffusion framework, the \textbf{D}ynamics-aware S\textbf{D}E \textbf{D}iffusion \textbf{G}enerative \textbf{M}odel (D$^3$GM). The \textit{Measure-preserving property} can return the degraded measurement to the original state despite complex degradation with the RDS concept of \textit{stability}. Our extensive experimental results corroborate the effectiveness of D$^3$GM across multiple benchmarks including a prominent application for inverse problems, magnetic resonance imaging. Code and data will be publicly available.

cs.AI

Weakly Supervised Learning of Cortical Surface Reconstruction from Segmentations

Existing learning-based cortical surface reconstruction approaches heavily rely on the supervision of pseudo ground truth (pGT) cortical surfaces for training. Such pGT surfaces are generated by traditional neuroimage processing pipelines, which are time consuming and difficult to generalize well to low-resolution brain MRI, e.g., from fetuses and neonates. In this work, we present CoSeg, a learning-based cortical surface reconstruction framework weakly supervised by brain segmentations without the need for pGT surfaces. CoSeg introduces temporal attention networks to learn time-varying velocity fields from brain MRI for diffeomorphic surface deformations, which fit an initial surface to target cortical surfaces within only 0.11 seconds for each brain hemisphere. A weakly supervised loss is designed to reconstruct pial surfaces by inflating the white surface along the normal direction towards the boundary of the cortical gray matter segmentation. This alleviates partial volume effects and encourages the pial surface to deform into deep and challenging cortical sulci. We evaluate CoSeg on 1,113 adult brain MRI at 1mm and 2mm resolution. CoSeg achieves superior geometric and morphological accuracy compared to existing learning-based approaches. We also verify that CoSeg can extract high-quality cortical surfaces from fetal brain MRI on which traditional pipelines fail to produce acceptable results.

eess.IV

The Developing Human Connectome Project: A Fast Deep Learning-based Pipeline for Neonatal Cortical Surface Reconstruction

The Developing Human Connectome Project (dHCP) aims to explore developmental patterns of the human brain during the perinatal period. An automated processing pipeline has been developed to extract high-quality cortical surfaces from structural brain magnetic resonance (MR) images for the dHCP neonatal dataset. However, the current implementation of the pipeline requires more than 6.5 hours to process a single MRI scan, making it expensive for large-scale neuroimaging studies. In this paper, we propose a fast deep learning (DL) based pipeline for dHCP neonatal cortical surface reconstruction, incorporating DL-based brain extraction, cortical surface reconstruction and spherical projection, as well as GPU-accelerated cortical surface inflation and cortical feature estimation. We introduce a multiscale deformation network to learn diffeomorphic cortical surface reconstruction end-to-end from T2-weighted brain MRI. A fast unsupervised spherical mapping approach is integrated to minimize metric distortions between cortical surfaces and projected spheres. The entire workflow of our DL-based dHCP pipeline completes within only 24 seconds on a modern GPU, which is nearly 1000 times faster than the original dHCP pipeline. The qualitative assessment demonstrates that for 82.5% of the test samples, the cortical surfaces reconstructed by our DL-based pipeline achieve superior (54.2%) or equal (28.3%) surface quality compared to the original dHCP pipeline.

eess.IV

Polarization-entangled photon pair generation from an epsilon-near-zero metasurface

Polarization-entangled photon pair sources are essential for diverse quantum technologies, such as quantum communication, computation, and imaging. However, the generation of complex polarization-entangled quantum states has long been constrained by the available nonlinear susceptibility tensor of natural nonlinear crystals, necessitating a cumbersome and intricate setup for additional coherent superposition or post-selection. In this study, we introduce and experimentally demonstrate a nanoscale polarization-entangled photon pair source utilizing an artificially-engineered metamaterial platform. This platform is based on a plasmonic metasurface that is strongly coupled to an epsilon-near-zero (ENZ) material. By precisely engineering resonances at both pump and signal/idler wavelengths, and leveraging the field enhancement provided by the ENZ effect, the photon pair generation efficiency of the 68-nm-thick metasurface is significantly boosted. More notably, the ENZ metasurface platform facilitates versatile manipulation of the system's anisotropic second-order nonlinear susceptibility tensor, enabling direct control over the polarization states of the photon pairs, which leads to the generation of a polarization-entangled Bell state without the need for additional components. Our approach opens a new avenue for the simultaneous photon pair generation and quantum state engineering in a compact platform.

physics.optics