SearcharxivSearch

arXiv subjects

Siyao Li

Publications and source records attributed to Siyao Li.

At least 19 recordsLinked to original sources

Diversity in Coded TE-QKD Channels: Achieving Infinite Diversity out of Finite System Resources

We establish conditions and give proofs on how an error-correcting code can attain infinite diversity in a time-entanglement quantum key distribution (TE-QKD) reconciliation. The shocking result, never encountered in the literature on coding and communication theory, is that a decoder exhibits an infinite diversity order while the channel has finite diversity and the code has a relatively short finite length. This paper studies the diversity order of coded TE-QKD reconciliation, defined by the asymptotic slope of the error probability at high signal-to-noise ratio. For bounded-distance algebraic decoding, we derive a necessary and sufficient condition in terms of the number of photons per codeword and the decoding radius. For soft-decision decoding, we introduce the maximal finite diversity (MFD) property and prove that infinite diversity is achieved if and only if the code is MFD deficient. The infinite diversity in TE-QKD has no counterpart in classical fading channels, where decoding can only multiply a finite diversity order by a finite factor. Examples of short codes based on Golay, Reed-Solomon, Bose-Chaudhuri-Hocquenghem (BCH), and Reed-Muller codes validate the analysis and illustrate how the TE-QKD system parameters and the relatively short code parameters affect the achievable diversity for both hard and soft information reconciliation.

cs.IT

Optimization of Collaborative Semantic Communication Network Performance with Channel and Content Preference Feedback

Existing semantic communication frameworks treat and transmit all image regions with equal importance, which is not practical for real-world applications which may prioritize different content in an image. To address this issue, we propose a novel semantic communication framework that enables a transmitter to use limited channel and content feedback to prioritize the transmission of important image regions. In particular, in the proposed framework, a base station (BS) divides each image into sub-images, extracts their semantic information, and transmits them to users according to their preferences. The users will reconstruct the image based on the received sub-images and cooperatively decide when to send channel state information (CSI) or content-preference feedback under dynamic channels and limited resources. We formulate an optimization problem to minimize the semantic-weighted mean square error between the original image and the regenerated image by optimizing sub-channel allocation, users' power allocation, and feedback selection. To address this problem, a value decomposition actor- critic (AC) with dynamic neighborhood construction (VDAC-DNC) scheme is proposed. The proposed method combines AC with value decomposition networks to allow the BS to approximate discrete actions by a continuous action distribution, thus reducing the output dimension and improving training efficiency. The introduced DNC method further improves training efficiency by constructing a small discrete neighboring action space to search for an action with the maximum Q value, thus avoiding traversing the large discrete action space. Simulation results show that the proposed VDAC-DNC scheme can improve the performance by up to 5.04% and 18.55% compared to the standard multi-agent QAC method and the proposed method without feedback transmission.

cs.NI

Oscillon decay via parametric resonance: the case of three-point scalar interactions

We investigate the decay dynamics of oscillons through interactions with an external scalar field. To examine how robust the decay dynamics of oscillons via parametric resonance we previously found in Li et al. 2025 are to the specific form of the coupling, we extend the analysis to include a three-point interaction $g_3\phi\chi^2$. We compute the Floquet exponents of the external field $\chi$ under an oscillating oscillon background and analyze how the instability bands depend on the coupling constants and the oscillon shapes. Numerical simulations of the two-field system show that, similar to the four-point case, the parametric resonance may cease before the oscillon is destroyed, leaving a smaller oscillon that decays only perturbatively. This indicates that the partial decay of oscillons through parametric resonance is a generic phenomenon of oscillon-scalar couplings, qualitatively insensitive to the specific interaction form, while the shape of instability bands, parameter dependence, and the precise critical oscillon energies depend on the specific coupling. Our findings provide further insights into the decay dynamics of oscillons and their potential role in the post-inflationary reheating process.

hep-ph

Blind Source Separation-Enabled Joint Communication and Sensing in IBFD MIMO Systems

This paper addresses the challenge of joint communication and sensing (JCAS) in next-generation wireless networks, with an emphasis on in-band full-duplex (IBFD) multiple-input multiple-output (MIMO) systems. Traditionally, self-interference (SI) in IBFD systems is a major obstacle to recovering the signal of interest (SOI). Under the JCAS paradigm, however, this high-power SI signal presents an opportunity for efficient sensing. Since each transceiver node has access to the original SI signal, its environmental reflections can be exploited to estimate channel conditions and detect changes, without requiring dedicated radar waveforms. We propose a blind source separation (BSS)-based framework to simultaneously perform self-interference cancellation (SIC) and extract sensing information in IBFD MIMO settings. The approach applies the Fast Independent Component Analysis (FastICA) algorithm to separate the SI and SOI signals while enabling simultaneous signal recovery and channel estimation. Simulation results confirm the framework's effectiveness, showing improved sensing and communication performance as signal frame size increases.

cs.ET

On Secrecy Capacity of Binary Beampointing Channels with Block Memory and Feedback

This paper investigates the secrecy capacity of the binary beampointing (BBP) channel with block memory and feedback, a simplified yet insightful model for millimeter-wave (mmWave) systems with beamformed transmissions and backscatter feedback. We consider a system where a legitimate receiver and a passive eavesdropper experience independent and uniformly distributed angular directions over transmission blocks, with the base station receiving noiseless, unit-delayed feedback from both, under the per-symbol input cost constraints. We establish a closed-form upper bound on the secrecy capacity, which is based on the main channel between the base station and the legitimate receiver. Moreover, we propose a joint communication and adaptive sensing (JCAS) scheme and derive its achievable secrecy rate. Simulation results show that the gap between the inner and outer bounds narrows as the number of block length increases. This reveals the efficiency of this JCAS scheme, which strategically leverages feedback to balance the demands of sensing the legitimate user and preventing information leakage to the eavesdropper.

cs.IT

On the Sensing Capacity of Gaussian "Beam-Pointing" Channels with Block Memory and Feedback

Driven by the demands of high-frequency wireless communications in 5G and 6G systems (e.g., mmWave, sub-THz), we explore a state-dependent {\em Gaussian beam-pointing} (GBP) channel. In this model, the channel state defines an unknown angle of departure (AoD), which remains constant within each coherence block of $Q$ time slots but changes independently across blocks. The transmitter receives strictly causal feedback which may originate from a radar detection system or explicit feedback from the receiver at the end of each slot and estimates the AoD at the end of each block. To enhance transmission efficiency, we propose a joint communication and sensing scheme. While the communication capacity of the GBP channel has been previously analyzed by the authors, this work focuses on sensing capacity, characterized by the mutual information between the channel state and the feedback conditioned on the transmitted signal. We derive an upper bound using dynamic programming and propose an achievable inner bound on the sensing capacity, both formulated as optimization problems. For the special case of $Q=1$, the proposed transmission scheme achieves the optimal sensing rate and highlights the inherent trade-off between sensing and communication performance.

cs.IT

VER-Bench: Evaluating MLLMs on Reasoning with Fine-Grained Visual Evidence

With the rapid development of MLLMs, evaluating their visual capabilities has become increasingly crucial. Current benchmarks primarily fall into two main types: basic perception benchmarks, which focus on local details but lack deep reasoning (e.g., "what is in the image?"), and mainstream reasoning benchmarks, which concentrate on prominent image elements but may fail to assess subtle clues requiring intricate analysis. However, profound visual understanding and complex reasoning depend more on interpreting subtle, inconspicuous local details than on perceiving salient, macro-level objects. These details, though occupying minimal image area, often contain richer, more critical information for robust analysis. To bridge this gap, we introduce the VER-Bench, a novel framework to evaluate MLLMs' ability to: 1) identify fine-grained visual clues, often occupying on average just 0.25% of the image area; 2) integrate these clues with world knowledge for complex reasoning. Comprising 374 carefully designed questions across Geospatial, Temporal, Situational, Intent, System State, and Symbolic reasoning, each question in VER-Bench is accompanied by structured evidence: visual clues and question-related reasoning derived from them. VER-Bench reveals current models' limitations in extracting subtle visual evidence and constructing evidence-based arguments, highlighting the need to enhance models's capabilities in fine-grained visual evidence extraction, integration, and reasoning for genuine visual understanding and human-like analysis. Dataset and additional materials are available https://github.com/verbta/ACMMM-25-Materials.

cs.CV

Decay and lifetime of oscillons coupled to an external scalar field: Insights from instability band analysis

Oscillons are long-lived, spherically symmetric solitons that can arise in real scalar field theories with potentials shallower than quadratic ones. They are considered to form via parametric resonance during the preheating stage after inflation and have extended lifetimes. However, the estimation of their lifespan becomes complicated when taking into account the interactions between the inflaton field and other fields, as naturally expected in realistic reheating scenarios. In this study, we investigate how the lifetime of a single oscillon is affected by the coupling to the external real scalar field. By numerically computing the instability bands of the external field with the inhomogeneous oscillon profile as background, we show that the resonance behavior depends intricately on the coupling strength and shape of the oscillon. We analyze distinct instability mechanisms that dominate across different regimes of the coupling strength and oscillon shapes. Especially, we show that the parametric resonance fails to occur when the oscillon size is too limited to drive enhancement of the external field. Furthermore, our simulations show that as the oscillon loses energy, the exponential growth of the external field can terminate before the oscillon reaches its critical energy for collapse, which indicates that the external field does not necessarily lead to rapid destruction of oscillons even in the presence of strong coupling or with large amplitudes. These results suggest that oscillons can remain long-lived across a wide range of coupling strengths, with potential implications for their role in cosmological evolution.

hep-ph

CMD-HAR: Cross-Modal Disentanglement for Wearable Human Activity Recognition

Human Activity Recognition (HAR) is a fundamental technology for numerous human - centered intelligent applications. Although deep learning methods have been utilized to accelerate feature extraction, issues such as multimodal data mixing, activity heterogeneity, and complex model deployment remain largely unresolved. The aim of this paper is to address issues such as multimodal data mixing, activity heterogeneity, and complex model deployment in sensor-based human activity recognition. We propose a spatiotemporal attention modal decomposition alignment fusion strategy to tackle the problem of the mixed distribution of sensor data. Key discriminative features of activities are captured through cross-modal spatio-temporal disentangled representation, and gradient modulation is combined to alleviate data heterogeneity. In addition, a wearable deployment simulation system is constructed. We conducted experiments on a large number of public datasets, demonstrating the effectiveness of the model.

cs.CV

Process Optimization and Deployment for Sensor-Based Human Activity Recognition Based on Deep Learning

Sensor-based human activity recognition is a key technology for many human-centered intelligent applications. However, this research is still in its infancy and faces many unresolved challenges. To address these, we propose a comprehensive optimization process approach centered on multi-attention interaction. We first utilize unsupervised statistical feature-guided diffusion models for highly adaptive data enhancement, and introduce a novel network architecture-Multi-branch Spatiotemporal Interaction Network, which uses multi-branch features at different levels to effectively Sequential ), which uses multi-branch features at different levels to effectively Sequential spatio-temporal interaction to enhance the ability to mine advanced latent features. In addition, we adopt a multi-loss function fusion strategy in the training phase to dynamically adjust the fusion weights between batches to optimize the training results. Finally, we also conducted actual deployment on embedded devices to extensively test the practical feasibility of the proposed method in existing work. We conduct extensive testing on three public datasets, including ablation studies, comparisons of related work, and embedded deployments.

eess.SP

FactCG: Enhancing Fact Checkers with Graph-Based Multi-Hop Data

Prior research on training grounded factuality classification models to detect hallucinations in large language models (LLMs) has relied on public natural language inference (NLI) data and synthetic data. However, conventional NLI datasets are not well-suited for document-level reasoning, which is critical for detecting LLM hallucinations. Recent approaches to document-level synthetic data generation involve iteratively removing sentences from documents and annotating factuality using LLM-based prompts. While effective, this method is computationally expensive for long documents and limited by the LLM's capabilities. In this work, we analyze the differences between existing synthetic training data used in state-of-the-art models and real LLM output claims. Based on our findings, we propose a novel approach for synthetic data generation, CG2C, that leverages multi-hop reasoning on context graphs extracted from documents. Our fact checker model, FactCG, demonstrates improved performance with more connected reasoning, using the same backbone models. Experiments show it even outperforms GPT-4-o on the LLM-Aggrefact benchmark with much smaller model size.

cs.CL

Compressed Sensing Inspired User Acquisition for Downlink Integrated Sensing and Communication Transmissions

This paper investigates radar-assisted user acquisition for downlink multi-user multiple-input multiple-output (MIMO) transmission using Orthogonal Frequency Division Multiplexing (OFDM) signals. Specifically, we formulate a concise mathematical model for the user acquisition problem, where each user is characterized by its delay and beamspace response. Therefore, we propose a two-stage method for user acquisition, where the Multiple Signal Classification (MUSIC) algorithm is adopted for delay estimation, and then a least absolute shrinkage and selection operator (LASSO) is applied for estimating the user response in the beamspace. Furthermore, we also provide a comprehensive performance analysis of the considered problem based on the pair-wise error probability (PEP). Particularly, we show that the rank and the geometric mean of non-zero eigenvalues of the squared beamspace difference matrix determines the user acquisition performance. More importantly, we reveal that simultaneously probing multiple beams outperforms concentrating power on a specific beam direction in each time slot under the power constraint, when only limited OFDM symbols are transmitted. Our numerical results confirm our conclusions and also demonstrate a promising acquisition performance of the proposed two-stage method.

cs.IT

Using Large Language Models for Humanitarian Frontline Negotiation: Opportunities and Considerations

Humanitarian negotiations in conflict zones, called \emph{frontline negotiation}, are often highly adversarial, complex, and high-risk. Several best-practices have emerged over the years that help negotiators extract insights from large datasets to navigate nuanced and rapidly evolving scenarios. Recent advances in large language models (LLMs) have sparked interest in the potential for AI to aid decision making in frontline negotiation. Through in-depth interviews with 13 experienced frontline negotiators, we identified their needs for AI-assisted case analysis and creativity support, as well as concerns surrounding confidentiality and model bias. We further explored the potential for AI augmentation of three standard tools used in frontline negotiation planning. We evaluated the quality and stability of our ChatGPT-based negotiation tools in the context of two real cases. Our findings highlight the potential for LLMs to enhance humanitarian negotiations and underscore the need for careful ethical and practical considerations.

cs.HC

Joint Fronthaul Load Balancing and Computation Resource Allocation in Cell-Free User-Centric Massive MIMO Networks

We consider scalable cell-free massive multiple-input multiple-output networks under an open radio access network paradigm comprising user equipments (UEs), radio units (RUs), and decentralized processing units (DUs). UEs are served by dynamically allocated user-centric clusters of RUs. The corresponding cluster processors (implementing the physical layer for each user) are hosted by the DUs as software-defined virtual network functions. Unlike the current literature, mainly focused on the characterization of the user rates under unrestricted fronthaul communication and computation, in this work we explicitly take into account the fronthaul topology, the limited fronthaul communication capacity, and computation constraints at the DUs. In particular, we systematically address the new problem of joint fronthaul load balancing and allocation of the computation resource. As a consequence of our new optimization framework, we present representative numerical results highlighting the existence of an optimal number of quantization bits in the analog-to-digital conversion at the RUs.

cs.IT

Interactions between several types of cosmic strings

We study the interaction of several types of static straight cosmic strings, including local strings, global strings, and bosonic superconducting strings with and without magnetic currents. First, we evaluate the interaction energy of two widely separated cosmic strings using the point source formalism and show that the most dominant contribution to the interaction energy comes from the excitation of the lightest mediator particles in a underlying theory. The interaction energy at arbitrary separation distances is then analyzed numerically by the gradient flow method. It turns out that an additional scalar field introduced in the bosonic superconducting string becomes an additional source of attraction. For such a bosonic superconducting string, we find that a string with two winding numbers is energetically favorable compared to two strings with a single winding number in a certain parameter region. Our analysis reveals that a phase structure of bosonic superconducting strings is richer than that of local and global strings and that the formation of bound states at intersections of bosonic superconducting strings is favored.

hep-ph

On the State Estimation Error of "Beam-Pointing'' Channels: The Binary Case

Sensing capabilities as an integral part of the network have been identified as a novel feature of sixth-generation (6G) wireless networks. As a key driver, millimeterwave (mmWave) communication largely boosts speed, capacities, and connectivity. In order to maximize the potential of mmWave communication, precise and fast beam acquisition (BA) is crucial, since it compensates for a high pathloss and provides a large beamforming gain. Practically, the angle-of-departure (AoD) remains almost constant over numerous consecutive time slots, the backscatter signal experiences some delay, and the hardware is restricted under the peak power constraint. This work captures these main features by a simple binary beam-pointing (BBP) channel model with in-block memory (iBM) [1], peak cost constraint, and one unit-delayed feedback. In particular, we focus on the sensing capabilities of such a model and characterize the performance of the BA process in terms of the Hamming distortion of the estimated channel state. We encode the position of the AoD and derive the minimum distortion of the BBP channel under the peak cost constraint with no communication constraint. Our previous work [2] proposed a joint communication and sensing (JCAS) algorithm, which achieves the capacity of the same channel model. Herein, we show that by employing this JCAS transmission strategy, optimal data communication and channel estimation can be accomplished simultaneously. This yields the complete characterization of the capacity-distortion tradeoff for this model.

cs.IT

ChartReader: A Unified Framework for Chart Derendering and Comprehension without Heuristic Rules

Charts are a powerful tool for visually conveying complex data, but their comprehension poses a challenge due to the diverse chart types and intricate components. Existing chart comprehension methods suffer from either heuristic rules or an over-reliance on OCR systems, resulting in suboptimal performance. To address these issues, we present ChartReader, a unified framework that seamlessly integrates chart derendering and comprehension tasks. Our approach includes a transformer-based chart component detection module and an extended pre-trained vision-language model for chart-to-X tasks. By learning the rules of charts automatically from annotated datasets, our approach eliminates the need for manual rule-making, reducing effort and enhancing accuracy.~We also introduce a data variable replacement technique and extend the input and position embeddings of the pre-trained model for cross-task training. We evaluate ChartReader on Chart-to-Table, ChartQA, and Chart-to-Text tasks, demonstrating its superiority over existing methods. Our proposed framework can significantly reduce the manual effort involved in chart analysis, providing a step towards a universal chart understanding model. Moreover, our approach offers opportunities for plug-and-play integration with mainstream LLMs such as T5 and TaPas, extending their capability to chart comprehension tasks. The code is available at https://github.com/zhiqic/ChartReader.

cs.CV

GSRFormer: Grounded Situation Recognition Transformer with Alternate Semantic Attention Refinement

Grounded Situation Recognition (GSR) aims to generate structured semantic summaries of images for "human-like" event understanding. Specifically, GSR task not only detects the salient activity verb (e.g. buying), but also predicts all corresponding semantic roles (e.g. agent and goods). Inspired by object detection and image captioning tasks, existing methods typically employ a two-stage framework: 1) detect the activity verb, and then 2) predict semantic roles based on the detected verb. Obviously, this illogical framework constitutes a huge obstacle to semantic understanding. First, pre-detecting verbs solely without semantic roles inevitably fails to distinguish many similar daily activities (e.g., offering and giving, buying and selling). Second, predicting semantic roles in a closed auto-regressive manner can hardly exploit the semantic relations among the verb and roles. To this end, in this paper we propose a novel two-stage framework that focuses on utilizing such bidirectional relations within verbs and roles. In the first stage, instead of pre-detecting the verb, we postpone the detection step and assume a pseudo label, where an intermediate representation for each corresponding semantic role is learned from images. In the second stage, we exploit transformer layers to unearth the potential semantic relations within both verbs and semantic roles. With the help of a set of support images, an alternate learning scheme is designed to simultaneously optimize the results: update the verb using nouns corresponding to the image, and update nouns using verbs from support images. Extensive experimental results on challenging SWiG benchmarks show that our renovated framework outperforms other state-of-the-art methods under various metrics.

cs.CV