SearcharxivSearch

arXiv subjects

Jie Luo

Publications and source records attributed to Jie Luo.

At least 19 recordsLinked to original sources

Boundary Geometry and Surjective Linear Isometries of Weighted Hardy Spaces

Let $D\subset\mathbb C^n$, $n\geq 2$, be a bounded pseudoconvex domain with smooth boundary, and let $H^p_\omega(D)$ be the Hardy space defined using a weighted boundary measure $\omega\,d\sigma$, where $\omega$ is bounded above and bounded away from zero. For every $0<p<\infty$, $p\neq2$, we prove that each surjective linear isometry $T$ of $H^p_\omega(D)$ has the rigid form \( Tf=T(1)(f\circ\varphi), \) where $\varphi\in\operatorname{Aut}(D)$. This extends the classical Forelli-type classification beyond highly symmetric or polynomially convex domains to arbitrary smoothly bounded pseudoconvex domains. The principal difficulty is not the construction of a holomorphic symbol, but proving that this symbol takes values in $D$ and is in fact biholomorphic. We overcome this difficulty by combining equimeasurability methods of Rudin and Schneider with boundary uniqueness, holomorphic approximation, plurisubharmonic exhaustion functions, and removable-singularity arguments across analytic sets. We also solve the complementary geometric problem of determining when an automorphism of $D$ gives rise to an isometry. The answer depends decisively on the boundary measure. We construct two natural measures for which every automorphism induces an isometry: one obtained from an invariant defining function when $\operatorname{Aut}(D)$ is compact, and the other given by Fefferman's invariant surface measure. In sharp contrast, we exhibit domains with noncompact automorphism group---including domains biholomorphic to the unit ball---for which the analogous conclusion fails for ordinary Euclidean surface measure. Thus the isometric structure of Hardy spaces detects not only the biholomorphic geometry of the domain, but also the finer interaction between that geometry and the chosen boundary measure.

math.CV

Manifold-Constrained PET Reconstruction with Learned Flow-Matching Priors

Image reconstruction for positron emission tomography (PET) is an ill-posed Poisson inverse problem that often suffers from severe noise amplification and artifacts. In this work, we introduce an unsupervised, optimization-based reconstruction framework that employs a flow-matching generative model as a learned manifold prior. We train the flow-matching model on high-quality PET images to learn a deterministic ordinary differential equation transport from a Gaussian latent distribution to the empirical PET image distribution, yielding a differentiable generator of anatomically plausible images. We incorporate this generator as an explicit manifold constraint into a regularized Poisson likelihood formulation. We solve the resulting optimization problem using an alternating direction method of multipliers algorithm, in which an expectation-maximization-type surrogate update enforces data consistency and a gradient-based latent-space projection enforces manifold proximity. We evaluate the proposed method on both simulated and real PET datasets, assessing dose-level robustness, lesion-insertion generalization, and cross-scanner transfer. Compared with conventional reconstruction methods and state-of-the-art deep learning baselines, the proposed method provides superior noise suppression, structural preservation, and quantitative accuracy, while maintaining high computational efficiency.

eess.IV

Counting Schreier Sets Under Neighborhood Conditions

We count Schreier sets that satisfy a neighborhood condition, including $k$-clustered, $k$-consecutive-free, $k$-neighbored, $k$-isolated, and closed under integral $2$-averages. For the first four conditions, we determine the initial counts and prove linear recurrence relations. For the last condition, we prove a recurrence that involves the divisor counting function.

math.CO

Beyond Noisy Signals: Dual-Level Denoising for Multi-modal Sequential Recommendation

Multi-modal Sequential Recommendation (SR) incorporates rich side information (e.g., textual and visual features) to enhance dynamic user preference modeling. However, existing frameworks inevitably suffer from a Dual-Noise Dilemma: (1) Feature-level redundancy stemming from the semantic gap between generic pre-trained representations and fine-grained recommendation intent; and (2) Sequence-level stochasticity induced by spurious interactions such as accidental clicks. To break this bottleneck, we propose DDMSR, a novel Dual-level Denoising Multi-modal Sequential Recommendation framework that systematically purifies signals from both feature-topological and sequence-frequency perspectives. Specifically, we first design a graph-based feature denoising module that leverages Laplacian smoothing on item semantic graphs as a structural low-pass filter, effectively suppressing high-frequency semantic noise while preserving salient features. For sequence purification, we introduce a frequency-domain sequence denoising module, utilizing the Fast Fourier Transform and a learnable frequency filter to adaptively modulate the interaction spectrum and attenuate anomalous signals. Furthermore, a multi-modal contrastive alignment objective is incorporated to bridge the heterogeneity gap and enforce cross-modal semantic consistency. Extensive experiments on four public benchmark datasets demonstrate that DDMSR consistently outperforms state-of-the-art baselines, providing a highly robust and efficient solution for multi-modal sequential recommendation. The source code is available at: https://github.com/jluo00/DDMSR.

cs.IR

Self-Referential $K$-SAT and the Finite Analogue of G\"odel's Incompleteness Theorem

Self-reference and solution independence are core properties underlying intractability. This paper establishes a finite combinatorial analogue of G\"odel's incompleteness theorems within Boolean $K$-SAT. While standard random $K$-SAT has assignment correlations that disrupt solution independence, we resolve this via a logarithmic-width ensemble ($K = O(\log N)$). Here, satisfying assignments converge to a Poisson distribution, letting unsatisfiable and uniquely satisfiable formulas coexist. By executing a single-clause substitution conditioned on the unique solution, we construct structurally irreducible SAT/UNSAT pairs that are indistinguishable via local evaluation. Using algorithmic information theory and Shannon channels, we prove that deductive pipelines restricted to a sublinear window suffer from an informational blind spot, forcing a descriptive lower bound of $K(\mathcal{A}) \geq \Omega(N^{1-\delta})$. This deficit forces any Resolution refutation of the UNSAT instance to utilize wide clauses ($w(\pi) \geq \Omega(N^{1-\delta})$), triggering an exponential proof-tree explosion ($S(\phi) \geq \exp(\Omega(N^{1-2\delta}))$). As $\delta \rightarrow 0^+$, this bound converges to the worst-case $2^N$ threshold, reframing the Strong Exponential Time Hypothesis (SETH) as a direct projection of G\"odel incompleteness onto finite computation. We diagnose the decades-long stagnation in complexity theory. Transitioning from Turing's class separation to a G\"odelian paradigm of instance indistinguishability, we introduce a multi-dimensional comparative framework that contrasts these two historical lineages across distinct perspectives. The self-referential hardness exhibits physical invariance: it precludes quantum shortcuts due to the necessity of global semantic analysis and delineates a scaling bottleneck for machine learning architectures operating on lossy, local compression.

cs.CC

Arbitrary-Order Scattering Exceptional Points in Configurable Non-Hermitian Zero-Index Materials

Scattering exceptional points (EPs) are non-Hermitian degeneracies where the eigenvalues and eigenvectors of scattering matrices coalesce, enabling many intriguing phenomena in optical systems. Higher-order scattering EPs are particularly notable for their ultrasensitive response to perturbations, yet achieving flexible, arbitrary-order control remains challenging. Here, we propose a configurable non-Hermitian zero-index material (ZIM) network that enables arbitrary-order scattering EPs, as rigorously proved theoretically and validated numerically. Specifically, we show that in an N-port non-Hermitian ZIM network embedded with loss/gain dopants, the maximum achievable EP order is N, and the order can be flexibly tuned from 2 to N or completely eliminated by adjusting the dopants. Furthermore, we compare conventional coherent perfect absorption with absorbing EPs of different orders. Although both achieve perfect absorption of all incident waves, a second-order EP already outperforms coherent perfect absorption, and higher-order EPs provide further power-law enhancement. These findings establish a pathway toward realizing arbitrary-order EPs in open scattering systems, holding significant promise for advanced sensing applications.

physics.optics

QuantSR+: Pushing the Limit of Quantized Image Super-Resolution Networks

Low-bit quantization is widely used to compress super-resolution (SR) models and reduce storage and computation costs for deployment on resource-limited devices. However, when SR models are pushed to ultra-low precision (2-4 bits), performance can drop sharply due to diminished representational capacity and the detail-sensitive nature of SR. To address these issues, we propose QuantSR+, a unified framework that improves quantization operators, network design, and training optimization, achieving better trade-offs between accuracy and efficiency than prior low-bit SR methods. QuantSR+ mainly relies on three technical contributions: (1) Redistribution-driven Bit Determination (RBD), which reshapes quantization distributions in both forward and backward passes to preserve representation fidelity; (2) Quantized Slimmable Architecture (QSA), which begins with an over-parameterized model and progressively prunes less critical blocks to meet efficiency budgets while pushing the accuracy performance; and (3) Slimming-guided Function-localized Distillation (SFD), which enforces block-aware feature alignment via a direct loss and a progressive, function-local training schedule to capture quantization effects better and speed up convergence. Extensive experiments show that QuantSR+ achieves state-of-the-art performance against both specialized quantized SR methods and generic quantization approaches. For SwinIR-S on Urban100 (x4), it improves PSNR by 0.29 dB over the 2-bit SOTA baseline. Meanwhile, it delivers strong efficiency gains at 2-bit, reducing operations by up to 87.9% and storage by 89.4%. QuantSR+ is effective for both convolutional and transformer-based SR models, indicating broad applicability.

cs.CV

Enjoy Your Layer Normalization with the Computational Efficiency of RMSNorm

Layer normalization (LN) is a fundamental component in modern deep learning, but its per-sample centering and scaling introduce non-negligible inference overhead. RMSNorm improves efficiency by removing the centering operation, yet this may discard benefits associated with centering. This paper propose a framework to determine whether an LN in an arbitrary DNN can be replaced by RMSNorm without changing the model function. The key idea is to fold LN's centering operation into upstream general linear layers by enforcing zero-mean outputs through the column-centered constraint (CCC) and column-based weight centering (CBWC). We extend the analysis to arbitrary DNNs, define such LNs as foldable LNs, and develop a graph-based detection algorithm. Our analysis shows that many LNs in widely used architectures are foldable, enabling exact inference-time conversion and end-to-end acceleration of 2% to 12% without changing model predictions. Experiments across multiple task families further show that, when exact equivalence is partially broken in practical training settings, our method remains competitive with vanilla LN while improving efficiency.

cs.LG

Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters

We introduce Step 3.5 Flash, a sparse Mixture-of-Experts (MoE) model that bridges frontier-level agentic intelligence and computational efficiency. We focus on what matters most when building agents: sharp reasoning and fast, reliable execution. Step 3.5 Flash pairs a 196B-parameter foundation with 11B active parameters for efficient inference. It is optimized with interleaved 3:1 sliding-window/full attention and Multi-Token Prediction (MTP-3) to reduce the latency and cost of multi-round agentic interactions. To reach frontier-level intelligence, we design a scalable reinforcement learning framework that combines verifiable signals with preference feedback, while remaining stable under large-scale off-policy training, enabling consistent self-improvement across mathematics, code, and tool use. Step 3.5 Flash demonstrates strong performance across agent, coding, and math tasks, achieving 85.4% on IMO-AnswerBench, 86.4% on LiveCodeBench-v6 (2024.08-2025.05), 88.2% on tau2-Bench, 69.0% on BrowseComp (with context management), and 51.0% on Terminal-Bench 2.0, comparable to frontier models such as GPT-5.2 xHigh and Gemini 3.0 Pro. By redefining the efficiency frontier, Step 3.5 Flash provides a high-density foundation for deploying sophisticated agents in real-world industrial environments.

cs.CL

Time-Aware Adaptive Side Information Fusion for Sequential Recommendation

Incorporating item-side information, such as category and brand, into sequential recommendation is a well-established and effective approach for improving performance. However, despite significant advancements, current models are generally limited by three key challenges: they often overlook the fine-grained temporal dynamics inherent in timestamps, exhibit vulnerability to noise in user interaction sequences, and rely on computationally expensive fusion architectures. To systematically address these challenges, we propose the Time-Aware Adaptive Side Information Fusion (TASIF) framework. TASIF integrates three synergistic components: (1) a simple, plug-and-play time span partitioning mechanism to capture global temporal patterns; (2) an adaptive frequency filter that leverages a learnable gate to denoise feature sequences adaptively, thereby providing higher-quality inputs for subsequent fusion modules; and (3) an efficient adaptive side information fusion layer, this layer employs a "guide-not-mix" architecture, where attributes guide the attention mechanism without being mixed into the content-representing item embeddings, ensuring deep interaction while ensuring computational efficiency. Extensive experiments on four public datasets demonstrate that TASIF significantly outperforms state-of-the-art baselines while maintaining excellent efficiency in training. Our source code is available at https://github.com/jluo00/TASIF.

cs.IR

Arbitrary Reflectionless Optical Routing via Non-Hermitian Zero-Index Networks

Optical routers are fundamental to photonic systems, but their performance is often limited by unwanted reflections and constrained functionalities. Existing design strategies generally lack complete control over reflectionless pathways and typically require computationally intensive iterative optimization. A general analytical framework for the inverse design of arbitrary reflectionless routing has remained unavailable. Here, we present an analytical inverse-design approach based on non-Hermitian zero-index networks, which enables arbitrary reflectionless routing for nearly any desired scattering response. By establishing a direct algebraic mapping between target scattering responses and the network's physical parameters, we transform the design process from iterative optimization into deterministic calculation. This approach enables the precise engineering of arbitrary reflectionless optical routing. We demonstrate its broad utility by designing devices from unicast and multicast routers with full amplitude and phase control to coherent beam combiners and spatial mode demultiplexers in four-port and six-port networks. Our work provides a systematic and analytical route to designing advanced light-control devices.

physics.optics

EfficientECG: Cross-Attention with Feature Fusion for Efficient Electrocardiogram Classification

Electrocardiogram is a useful diagnostic signal that can detect cardiac abnormalities by measuring the electrical activity generated by the heart. Due to its rapid, non-invasive, and richly informative characteristics, ECG has many emerging applications. In this paper, we study novel deep learning technologies to effectively manage and analyse ECG data, with the aim of building a diagnostic model, accurately and quickly, that can substantially reduce the burden on medical workers. Unlike the existing ECG models that exhibit a high misdiagnosis rate, our deep learning approaches can automatically extract the features of ECG data through end-to-end training. Specifically, we first devise EfficientECG, an accurate and lightweight classification model for ECG analysis based on the existing EfficientNet model, which can effectively handle high-frequency long-sequence ECG data with various leading types. On top of that, we next propose a cross-attention-based feature fusion model of EfficientECG for analysing multi-lead ECG data with multiple features (e.g., gender and age). Our evaluations on representative ECG datasets validate the superiority of our model against state-of-the-art works in terms of high precision, multi-feature fusion, and lightweights.

cs.LG

Unified Error Analysis for Synchronous and Asynchronous Two-User Random Access

We consider a two-user random access system in which each user independently selects a coding scheme from a finite set for every message, without sharing these choices with the other user or with the receiver. The receiver aims to decode only user 1 message but may also decode user 2 message when beneficial. In the synchronous setting, the receiver employs two parallel sub-decoders: one dedicated to decoding user 1 message and another that jointly decodes both users messages. Their outputs are synthesized to produce the final decoding or collision decision. For the asynchronous setting, we examine a time interval containing $L$ consecutive codewords from each user. The receiver deploys $2^{2L}$ parallel sub-decoders, each responsible for decoding a subset of the message-code index pairs. In both synchronous and asynchronous cases, every sub-decoder partitions the coding space into three disjoint regions: operation, margin, and collision, and outputs either decoded messages or a collision report according to the region in which the estimated code index vector lies. Error events are defined for each sub-decoder and for the overall receiver whenever the expected output is not produced. We derive achievable upper bounds on the generalized error performance, defined as a weighted sum of incorrect-decoding, collision, and miss-detection probabilities, for both synchronous and asynchronous scenarios.

cs.IT

Mip-NeWRF: Enhanced Wireless Radiance Field with Hybrid Encoding for Channel Prediction

Recent work on wireless radiance fields represents a promising deep learning approach for channel prediction, however, in complex environments these methods still exhibit limited robustness, slow convergence, and modest accuracy due to insufficiently refined modeling. To address this issue, we propose Mip-NeWRF, a physics-informed neural framework for accurate indoor channel prediction based on sparse channel measurements. The framework operates in a ray-based pipeline with coarse-to-fine importance sampling: frustum samples are encoded, processed by a shared multilayer perceptron (MLP), and the outputs are synthesized into the channel frequency response (CFR). Prior to MLP input, Mip-NeWRF performs conical-frustum sampling and applies a scale-consistent hybrid positional encoding to each frustum. The scale-consistent normalization aligns positional encodings across scene scales, while the hybrid encoding supplies both scale-robust, low-frequency stability to accelerate convergence and fine spatial detail to improve accuracy. During training, a curriculum learning schedule is applied to stabilize and accelerate convergence of the shared MLP. During channel synthesis, the MLP outputs, including predicted virtual transmitter presence probabilities and amplitudes, are combined with modeled pathloss and surface interaction attenuation to enhance physical fidelity and further improve accuracy. Simulation results demonstrate the effectiveness of the proposed approach: in typical scenarios, the normalized mean square error (NMSE) is reduced by 14.3 dB versus state-of-the-art baselines.

eess.SP

Geometry-Aware Cross Modal Alignment for Light Field-LiDAR Semantic Segmentation

Semantic segmentation serves as a cornerstone of scene understanding in autonomous driving but continues to face significant challenges under complex conditions such as occlusion. Light field and LiDAR modalities provide complementary visual and spatial cues that are beneficial for robust perception; however, their effective integration is hindered by limited viewpoint diversity and inherent modality discrepancies. To address these challenges, the first multimodal semantic segmentation dataset integrating light field data and point cloud data is proposed. Based on this dataset, we proposed a multi-modal light field point-cloud fusion segmentation network(Mlpfseg), incorporating feature completion and depth perception to segment both camera images and LiDAR point clouds simultaneously. The feature completion module addresses the density mismatch between point clouds and image pixels by performing differential reconstruction of point-cloud feature maps, enhancing the fusion of these modalities. The depth perception module improves the segmentation of occluded objects by reinforcing attention scores for better occlusion awareness. Our method outperforms image-only segmentation by 1.71 Mean Intersection over Union(mIoU) and point cloud-only segmentation by 2.38 mIoU, demonstrating its effectiveness.

cs.CV

Vulnerable Agent Identification in Large-Scale Multi-Agent Reinforcement Learning

Partial agent failure becomes inevitable when systems scale up, making it crucial to identify the subset of agents whose failure causes worst-case system performance degradations. We study this Vulnerable Agent Identification (VAI) problem in large-scale multi-agent reinforcement learning (MARL). We frame VAI as a Hierarchical Adversarial Decentralized Mean Field Control (HAD-MFC), where the upper level selects vulnerable agents as an NP-hard task and the lower level learns their worst-case adversarial policies via mean-field MARL. The two problems are coupled together, making HAD-MFC difficult to solve. To handle this, we first decouple the hierarchical process by Fenchel-Rockafellar transform, resulting a regularized mean-field Bellman operator for upper level that enables independent learning at each level, thus reducing computational complexity. We next reformulate the upper-level NP-hard problem as an MDP with dense rewards, allowing sequential identification of vulnerable agents via greedy and RL algorithms. This decomposition provably preserves the optimal solution. Experiments show our method effectively identifies more vulnerable agents in large-scale MARL and the rule-based system, fooling system into worse failures, and reveals the vulnerability of each agent in large systems. Code available at https://github.com/Waken-dream/VAI

cs.MA