SearcharxivSearch

arXiv subjects

Zhixiong Chen

Publications and source records attributed to Zhixiong Chen.

At least 19 recordsLinked to original sources

SceneGTMM: A Conformal Mapping-based Scene-Aware Transferable GNN-Transformer Dual-Graph Interaction Framework for Map Matching

Map matching is a key technology connecting positioning data with high precision road networks, but it faces challenges in noise robustness, cross regional transfer, and interpretability. To addr ess the limitations of existing methods in local global fusion, dynamic road network adaptation, and reliance on black box mod els, this paper proposes SceneGTMM, a transferable GNN Transformer dual graph interaction map matching framework based on a conformal mapping based scene relative strategy. 1) Conformal mapping based scene relative strategy: constructs trajectory centric local coordinate systems to reduce dependence on the training road network, supporting cross regional transfer and dynamic road network updates; 2) GNN Transformer dual graph interaction architecture: a GNN modeled road graph captures local topological constraints, while a Transformer modeled trajectory graph captures global temporal dependencies, and cross graph attention achieves noise suppression and semantic alignment; 3) CRF enhanced structured prediction: combines the global context of the Transformer with the topological transition constraints of CRF to improve path connectivity and robustness. Experiments show that SceneGTM achieves over 80% accuracy on multi source trajectories with positioning errors of 16 50 meters, representing a 5.3% improvement over HMM. In cross city transfer scenarios, it outperforms MTrajRec, GraphMM, and TMM, and enhances interpretability through attention and relative coordinate visualization. This study provides a new paradigm for high precision, transferable map matching for real time traffic perception and autonomous driving path planning.

cs.CV

Latent Sculpting for Zero-Shot Generalization: A Manifold Learning Approach to Out-of-Distribution Anomaly Detection

Detecting previously unseen attacks remains a major challenge for machine learning-based intrusion detection systems. Deep models trained on network traffic often achieve high accuracy on known attacks but fail under distributional shift because their decision boundaries are tightly coupled to the training data distribution. We introduce Latent Sculpting, a two-stage anomaly detection framework that improves robustness by explicitly structuring the latent representation before density estimation. The first stage trains a Transformer-based tabular encoder using a novel Binary Latent Sculpting loss, which encourages benign traffic to form a compact latent cluster while enforcing separation from anomalous patterns. The second stage fits a Masked Autoregressive Flow to the resulting latent space to produce calibrated probabilistic anomaly scores. Under a strict zero-shot evaluation protocol on the CIC-IDS-2017 benchmark, Stage 1 attains an F1-score of 0.98 on known attacks, while Stage 2 -- evaluated at the balanced threshold (85th-percentile) -- achieves a zero-shot OOD F1-score of 0.867 and AUROC of 0.913. The model successfully detects difficult distribution shifts including stealthy infiltration attacks (78.7% recall, peaking at 97.2%) and low-volume DoS variants (>94% recall), scenarios where conventional approaches often fail. Our results suggest that explicitly separating latent geometry learning from density modeling provides a stable approach for detecting zero-day cyber threats.

cs.LG

Sampling-Free Diffusion Transformers for Low-Complexity MIMO Channel Estimation

Diffusion model-based channel estimators have shown impressive performance but suffer from high computational complexity because they rely on iterative reverse sampling. This paper proposes a sampling-free diffusion transformer (DiT) for low-complexity MIMO channel estimation, termed SF-DiT-CE. Exploiting angular-domain sparsity of MIMO channels, we train a lightweight DiT to directly predict the clean channels from their perturbed observations and noise levels. At inference, the least square (LS) estimate and estimation noise condition the DiT to recover the channel in a single forward pass, eliminating iterative sampling. Numerical results demonstrate that our method achieves superior estimation accuracy and robustness with significantly lower complexity than state-of-the-art baselines.

eess.SP

Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities

Large language models (LLMs) have advanced rapidly, emerging as versatile tools across fields thanks to their exceptional language understanding, generation, and reasoning capabilities. However, performing LLM inference at the network edge remains challenging due to their large memory and compute demands. This survey outlines the challenges specific to LLM edge inference and provides a comprehensive overview of recent progress, covering system architectures, model optimization and deployment, and resource management and scheduling. By synthesizing state-of-the-art techniques and mapping future directions, this survey aims to unlock the potential of LLMs in resource-constrained edge environments.

cs.DC

Scalable Back-Propagation-Free Training of Optical Physics-Informed Neural Networks

Physics intelligence and digital twins often require rapid and repeated performance evaluation of various engineering systems (e.g. robots, autonomous vehicles, semiconductor chips) to enable (almost) real-time actions or decision making. This has motivated the development of accelerated partial differential equation (PDE) solvers, in resource-constrained scenarios if the PDE solvers are to be deployed on the edge. Physics-informed neural networks (PINNs) have shown promise in solving high-dimensional PDEs, but the training time on state-of-the-art digital hardware (e.g., GPUs) is still orders-of-magnitude longer than the latency required for enabling real-time decision making. Photonic computing offers a potential solution to address this huge latency gap because of its ultra-high operation speed. However, the lack of photonic memory and the large device sizes prevent training real-size PINNs on photonic chips. This paper proposes a completely back-propagation-free (BP-free) and highly salable framework for training real-size PINNs on silicon photonic platforms. Our approach involves three key innovations: (1) a sparse-grid Stein derivative estimator to avoid the BP in the loss evaluation of a PINN, (2) a dimension-reduced zeroth-order optimization via tensor-train decomposition to achieve better scalability and convergence in BP-free training, and (3) a scalable on-chip photonic PINN training accelerator design using photonic tensor cores. We validate our numerical methods on both low- and high-dimensional PDE benchmarks. Through pre-silicon simulation based on real device parameters, we further demonstrate the significant performance benefit (e.g., real-time training, huge chip area reduction) of our photonic accelerator.

cs.LG

Coordinate-conditioned Deconvolution for Scalable Spatially Varying High-Throughput Imaging

Wide-field fluorescence microscopy with compact optics often suffers from spatially varying blur due to field-dependent aberrations, vignetting, and sensor truncation, while finite sensor sampling imposes an inherent trade-off between field of view (FOV) and resolution. Computational Miniaturized Mesoscope (CM2) alleviate the sampling limit by multiplexing multiple sub-views onto a single sensor, but introduce view crosstalk and a highly ill-conditioned inverse problem compounded by spatially variant point spread functions (PSFs). Prior learning-based spatially varying (SV) reconstruction methods typically rely on global SV operators with fixed input sizes, resulting in memory and training costs that scale poorly with image dimensions. We propose SV-CoDe (Spatially Varying Coordinate-conditioned Deconvolution), a scalable deep learning framework that achieves uniform, high-resolution reconstruction across a 6.5 mm FOV. Unlike conventional methods, SV-CoDe employs coordinate-conditioned convolutions to locally adapt reconstruction kernels; this enables patch-based training that decouples parameter count from FOV size. SV-CoDe achieves the best image quality in both simulated and experimental measurements while requiring 10x less model size and 10x less training data than prior baselines. Trained purely on physics-based simulations, the network robustly generalizes to bead phantoms, weakly scattering brain slices, and freely moving C. elegans. SV-CoDe offers a scalable, physics-aware solution for correcting SV blur in compact optical systems and is readily extendable to a broad range of biomedical imaging applications.

eess.IV

From n-systems to Lie and Courant algebroids

This paper introduces a method for constructing pure algebroids, dull algebroids, and Lie algebroids. The construction relies on what we deffned as n-systems on vector bundles, and we provide explicit computations for all resulting structure maps. Analogously, metric n-systems deffned on metric vector bundles allow us to construct metric algebroids, pre-Courant algebroids, and Courant algebroids.

math.DG

Efficient LLM Inference over Heterogeneous Edge Networks with Speculative Decoding

Large language model (LLM) inference at the network edge is a promising serving paradigm that leverages distributed edge resources to run inference near users and enhance privacy. Existing edge-based LLM inference systems typically adopt autoregressive decoding (AD), which only generates one token per forward pass. This iterative process, compounded by the limited computational resources of edge nodes, results in high serving latency and constrains the system's ability to support multiple users under growing demands.To address these challenges, we propose a speculative decoding (SD)-based LLM serving framework that deploys small and large models across heterogeneous edge nodes to collaboratively deliver inference services. Specifically, the small model rapidly generates draft tokens that the large model verifies in parallel, enabling multi-token generation per forward pass and thus reducing serving latency. To improve resource utilization of edge nodes, we incorporate pipeline parallelism to overlap drafting and verification across multiple inference tasks. Based on this framework, we analyze and derive a comprehensive latency model incorporating both communication and inference latency. Then, we formulate a joint optimization problem for speculation length, task batching, and wireless communication resource allocation to minimize total serving latency. To address this problem, we derive the closed-form solutions for wireless communication resource allocation, and develop a dynamic programming algorithm for joint batching and speculation control strategies. Experimental results demonstrate that the proposed framework achieves lower serving latency compared to AD-based serving systems. In addition,the proposed joint optimization method delivers up to 44.9% latency reduction compared to benchmark schemes.

eess.SY

Large Language Model-Empowered Channel Prediction and Predictive Beamforming for LEO Satellite Communications

Accurate channel prediction and effective beamforming are essential for low Earth orbit (LEO) satellite communications to enhance system capacity and enable high-speed connectivity. Most existing channel prediction and predictive beamforming methods are limited by model generalization capabilities and struggle to adapt to time-varying wireless propagation environments. Inspired by the remarkable generalization and reasoning capabilities of large language models (LLMs), this work proposes an LLM-based channel prediction framework, namely CPLLM, to forecast future channel state information (CSI) for LEO satellites based on historical CSI data. In the proposed CPLLM, a dedicated CSI encoder is designed to map raw CSI data into the textual embedding space, effectively bridging the modality gap and enabling the LLM to perform reliable reasoning over CSI data. Additionally, a CSI decoder is introduced to simultaneously predict CSI for multiple future time slots, substantially reducing the computational burden and inference latency associated with the inherent autoregressive decoding process of LLMs. Then, instead of training the LLM from scratch, we adopt a parameter-efficient fine-tuning strategy, i.e., LoRA, for CPLLM, where the pretrained LLM remains frozen and trainable low-rank matrices are injected into each Transformer decoder layer to enable effective fine-tuning. Furthermore, we extend CPLLM to directly generate beamforming strategies for future time slots based on historical CSI data, namely BFLLM. This extended framework retains the same architecture as CPLLM, while introducing a dedicated beamforming decoder to output beamforming strategies. Finally, extensive simulation results validate the effectiveness of the proposed approaches in channel prediction and predictive beamforming for LEO satellite communications.

eess.SP

On the $N$th $2$-adic complexity of binary sequences identified with algebraic $2$-adic integers

We identify a binary sequence $\mathcal{S}=(s_n)_{n=0}^\infty$ with the $2$-adic integer $G_\mathcal{S}(2)=\sum\limits_{n=0}^\infty s_n2^n$. In the case that $G_\mathcal{S}(2)$ is algebraic over $\mathbb{Q}$ of degree $d\ge 2$, we prove that the $N$th $2$-adic complexity of $\mathcal{S}$ is at least $\frac{N}{d}+O(1)$, where the implied constant depends only on the minimal polynomial of $G_\mathcal{S}(2)$. This result is an analog of the bound of Mérai and the second author on the linear complexity of automatic sequences, that is, sequences with algebraic $G_\mathcal{S}(X)$ over the rational function field $\mathbb{F}_2(X)$. We further discuss the most important case $d=2$ in both settings and explain that the intersection of the set of $2$-adic algebraic sequences and the set of automatic sequences is the set of (eventually) periodic sequences. Finally, we provide some experimental results supporting the conjecture that $2$-adic algebraic sequences can have also a desirable $N$th linear complexity and automatic sequences a desirable $N$th $2$-adic complexity, respectively.

math.NT

CyberMentor: AI Powered Learning Tool Platform to Address Diverse Student Needs in Cybersecurity Education

Many non-traditional students in cybersecurity programs often lack access to advice from peers, family members and professors, which can hinder their educational experiences. Additionally, these students may not fully benefit from various LLM-powered AI assistants due to issues like content relevance, locality of advice, minimum expertise, and timing. This paper addresses these challenges by introducing an application designed to provide comprehensive support by answering questions related to knowledge, skills, and career preparation advice tailored to the needs of these students. We developed a learning tool platform, CyberMentor, to address the diverse needs and pain points of students majoring in cybersecurity. Powered by agentic workflow and Generative Large Language Models (LLMs), the platform leverages Retrieval-Augmented Generation (RAG) for accurate and contextually relevant information retrieval to achieve accessibility and personalization. We demonstrated its value in addressing knowledge requirements for cybersecurity education and for career marketability, in tackling skill requirements for analytical and programming assignments, and in delivering real time on demand learning support. Using three use scenarios, we showcased CyberMentor in facilitating knowledge acquisition and career preparation and providing seamless skill-based guidance and support. We also employed the LangChain prompt-based evaluation methodology to evaluate the platform's impact, confirming its strong performance in helpfulness, correctness, and completeness. These results underscore the system's ability to support students in developing practical cybersecurity skills while improving equity and sustainability within higher education. Furthermore, CyberMentor's open-source design allows for adaptation across other disciplines, fostering educational innovation and broadening its potential impact.

cs.CY

Enhancing Computer Programming Education with LLMs: A Study on Effective Prompt Engineering for Python Code Generation

Large language models (LLMs) and prompt engineering hold significant potential for advancing computer programming education through personalized instruction. This paper explores this potential by investigating three critical research questions: the systematic categorization of prompt engineering strategies tailored to diverse educational needs, the empowerment of LLMs to solve complex problems beyond their inherent capabilities, and the establishment of a robust framework for evaluating and implementing these strategies. Our methodology involves categorizing programming questions based on educational requirements, applying various prompt engineering strategies, and assessing the effectiveness of LLM-generated responses. Experiments with GPT-4, GPT-4o, Llama3-8b, and Mixtral-8x7b models on datasets such as LeetCode and USACO reveal that GPT-4o consistently outperforms others, particularly with the "multi-step" prompt strategy. The results show that tailored prompt strategies significantly enhance LLM performance, with specific strategies recommended for foundational learning, competition preparation, and advanced problem-solving. This study underscores the crucial role of prompt engineering in maximizing the educational benefits of LLMs. By systematically categorizing and testing these strategies, we provide a comprehensive framework for both educators and students to optimize LLM-based learning experiences. Future research should focus on refining these strategies and addressing current LLM limitations to further enhance educational outcomes in computer programming instruction.

cs.AI

The geometric constraints on Filippov algebroids

Filippov n-algebroids are introduced by Grabowski and Marmo as a natural generalization of Lie algebroids. On this note, we characterized Filippov n-algebroid structures by considering certain multi-input connections, which we called Filippov connections, on the underlying vector bundle. Through this approach, we could express the n-ary bracket of any Filippov n-algebroid using a torsion-free type formula. Additionally, we transformed the generalized Jacobi identity of the Filippov n-algebroid into the Bianchi-Filippov identity. Furthermore, in the case of rank n vector bundles, we provided a characterization of linear Nambu-Poisson structures using Filippov connections.

math.RA

Real-Time FJ/MAC PDE Solvers via Tensorized, Back-Propagation-Free Optical PINN Training

Solving partial differential equations (PDEs) numerically often requires huge computing time, energy cost, and hardware resources in practical applications. This has limited their applications in many scenarios (e.g., autonomous systems, supersonic flows) that have a limited energy budget and require near real-time response. Leveraging optical computing, this paper develops an on-chip training framework for physics-informed neural networks (PINNs), aiming to solve high-dimensional PDEs with fJ/MAC photonic power consumption and ultra-low latency. Despite the ultra-high speed of optical neural networks, training a PINN on an optical chip is hard due to (1) the large size of photonic devices, and (2) the lack of scalable optical memory devices to store the intermediate results of back-propagation (BP). To enable realistic optical PINN training, this paper presents a scalable method to avoid the BP process. We also employ a tensor-compressed approach to improve the convergence and scalability of our optical PINN training. This training framework is designed with tensorized optical neural networks (TONN) for scalable inference acceleration and MZI phase-domain tuning for \textit{in-situ} optimization. Our simulation results of a 20-dim HJB PDE show that our photonic accelerator can reduce the number of MZIs by a factor of $1.17\times 10^3$, with only $1.36$ J and $1.15$ s to solve this equation. This is the first real-size optical PINN training framework that can be applied to solve high-dimensional PDEs.

cs.LG

Tensor-Compressed Back-Propagation-Free Training for (Physics-Informed) Neural Networks

Backward propagation (BP) is widely used to compute the gradients in neural network training. However, it is hard to implement BP on edge devices due to the lack of hardware and software resources to support automatic differentiation. This has tremendously increased the design complexity and time-to-market of on-device training accelerators. This paper presents a completely BP-free framework that only requires forward propagation to train realistic neural networks. Our technical contributions are three-fold. Firstly, we present a tensor-compressed variance reduction approach to greatly improve the scalability of zeroth-order (ZO) optimization, making it feasible to handle a network size that is beyond the capability of previous ZO approaches. Secondly, we present a hybrid gradient evaluation approach to improve the efficiency of ZO training. Finally, we extend our BP-free training framework to physics-informed neural networks (PINNs) by proposing a sparse-grid approach to estimate the derivatives in the loss function without using BP. Our BP-free training only loses little accuracy on the MNIST dataset compared with standard first-order training. We also demonstrate successful results in training a PINN for solving a 20-dim Hamiltonian-Jacobi-Bellman PDE. This memory-efficient and BP-free approach may serve as a foundation for the near-future on-device training on many resource-constraint platforms (e.g., FPGA, ASIC, micro-controllers, and photonic chips).

cs.LG

Maximum-order complexity and $2$-adic complexity

The $2$-adic complexity has been well-analyzed in the periodic case. However, we are not aware of any theoretical results on the $N$th $2$-adic complexity of any promising candidate for a pseudorandom sequence of finite length $N$ or results on a part of the period of length $N$ of a periodic sequence, respectively. Here we introduce the first method for this aperiodic case. More precisely, we study the relation between $N$th maximum-order complexity and $N$th $2$-adic complexity of binary sequences and prove a lower bound on the $N$th $2$-adic complexity in terms of the $N$th maximum-order complexity. Then any known lower bound on the $N$th maximum-order complexity implies a lower bound on the $N$th $2$-adic complexity of the same order of magnitude. In the periodic case, one can prove a slightly better result. The latter bound is sharp which is illustrated by the maximum-order complexity of $\ell$-sequences. The idea of the proof helps us to characterize the maximum-order complexity of periodic sequences in terms of the unique rational number defined by the sequence. We also show that a periodic sequence of maximal maximum-order complexity must be also of maximal $2$-adic complexity.

cs.IT

Weak $(p,k)$-Dirac manifolds

In this paper, we introduce the notion of a weak $(p,k)$-Dirac structure in $TM\oplus Λ^pT^*M$, where $0\leq k \leq p-1$. The weak $(p,k)$-Lagrangian condition has more informations than the $(p,k)$-Lagrangian condition and contains the $(p,k)$-Lagrangian condition. The weak $(p,0)$-Dirac structures are exactly the higher Dirac structures of order p introduced by N. Martinez Alba and H. Bursztyn in [23] and [6], respectively. The regular weak $(p,p-1)$-Dirac structure together with $(p,p-1)$-Lagrangian subspace at each point $m\in M$ have the multisymplectic foliation. Finally, we introduce the notion of weak $(p,k)$-Dirac morphism. We give the condition that a weak $(p,k)$-Dirac manifold is also a weak $(p,k)$-Dirac manifold after pulling back.

math.DG

On (co-)morphisms of $n$-Lie-Rinehart algebras with applications to Nambu-Poisson manifolds

In this paper, we give a unified description of morphisms and comorphisms of $n$-Lie-Rinehart algebras. We show that these morphisms and comorphisms can be regarded as two subalgebras of the $ψ$-sum of $n$-Lie-Rinehart algebras. We also provide similar descriptions for morphisms and comorphisms of $n$-Lie algebroids. It is proved that the category of vector bundles with Nambu-Poisson structures of rank $n$ and the category of their dual bundles with $n$-Lie algebroid structures of rank $n$ are equivalent to each other.

math.RA