SearcharxivSearch

arXiv subjects

Liang Jin

Publications and source records attributed to Liang Jin.

At least 19 recordsLinked to original sources

Multi-Stream Spatiotemporal Channel Coding for MIMO Systems: Transmission Scheme Design and Achievable Rate Optimization

Spatiotemporal channel coding (STCC) can improve the achievable rate over traditional temporal channel coding (TCC) by leveraging spatial degrees of freedom to extend the codeword length. Although several information-theoretic foundations on STCC have been established, the investigation of transmission schemes from a communication-theoretic perspective remains in its early stages. This paper proposes a multi-stream over multi-subchannel STCC (STCC-MSC) under full channel state information assumption and optimizes its achievable rate in the finite blocklength regime. We first formulate the transmission architecture of STCC-MSC in a point-to-point MIMO system, which introduces a stream-subchannel matching mechanism. We then maximize the achievable rate of STCC-MSC by jointly optimizing the subchannel assignment and power allocation strategies, which is formulated as a mixed-integer-nonlinear-programming problem. Next, a penalized alternating convex approximation (PACA) algorithm is proposed to solve this problem. Subsequently, we extend the point-to-point STCC-MSC designs to the more general multi-user MIMO systems, including both uplink and downlink scenarios. Finally, simulation results indicate that the PACA algorithm achieves a 9.85% rate improvement over the benchmark algorithm within the STCC-MSC scheme. Furthermore, the joint STCC-MSC-PACA scheme improves the achievable rate by 28.68% over TCC scheme.

eess.SP

Secure Cooperative THz ISAC via Mamba Empowered Graph Neural Network Precoding

The terahertz (THz) band offers abundant spectrum resources for high-throughput communication and ultra high-precision localization. This paper investigates secure communication in cooperative THz orthogonal frequency-division multiplexing (OFDM) bistatic integrated sensing and communications (ISAC) systems, where multiple base stations (BSs) equipped with extremely large-scale antenna arrays (ELAAs) collaboratively serve downlink users while concurrently locating multiple targets. Malicious targets are assumed to act as potential eavesdroppers attempting to intercept confidential information intended for legitimate users. To mitigate these threats, we formulate a joint optimization problem for analog beamforming, digital precoding, true-time delayers (TTDs), and sensing signal covariance matrix design. The objective is to maximize the minimum secrecy rate subject to Cramer-Rao bound (CRB) constraints that ensure localization accuracy. This problem is highly challenging due to the non-convex CRB constraint, strongly coupled variables, high computational complexity from ELAA, and near-field channel modeling. To address these challenges, we propose a novel data-driven framework that integrates graph neural networks (GNNs) with the Mamba architecture. Our proposed framework first encodes the interactions among users, targets, and BSs into a heterogeneous graph and then employs message passing to optimize vertex features. The Mamba blocks further enhance this process through their selection mechanism and state space modeling capabilities, enabling dynamic and context-aware optimization of beamforming, TTD configurations, and sensing parameters. Numerical simulations validate that the proposed method outperforms both conventional and learning-based baselines, while offering high computational efficiency and strong generalization across different network conditions.

cs.IT

Intelligent Wiretap Code Design: Exploiting Wireless Endogenous Security via Information Theory and Deep Learning Integration

Recent advancements in wireless endogenous security have explored leveraging the inherent randomness of wireless channels to enhance communication security, providing an effective alternative to traditional encryption methods. This paper proposes a wiretap coding scheme within the semantic communication framework, which leverages discrete semantic representations compatible with conventional digital modulation to jointly enhance communication security and reliability. We investigate two eavesdropping scenarios: (i) the eavesdropper employs a maximum a posteriori (MAP) decoder, and (ii) the eavesdropper has access to a decoder identical to that of the legitimate receiver. In the first scenario, we exploit mutual information as a metric to guide the design of an optimized coding strategy, minimizing information leakage while enhancing communication reliability. In the second scenario, considering the limitations of the eavesdropper's decoding capability, we employ generalized mutual information (GMI) to characterize recoverability under the prescribed decoding rule and guide reliability-aware code optimization.

cs.IT

On the inhomogeneous discounted Hamilton-Jacobi equations

In this paper, we study the family of inhomogeneous discounted Hamilton-Jacobi equations \begin{equation}\label{hjs1} \lambda(x)u+h(x,d_x u)=c \quad \tag{$\ast$} \end{equation} on a closed manifold $M$ with a non-identically vanishing discount factor $\lambda(x)$. There is a critical value $c_0\in[-\infty,\infty)$ such that \eqref{hjs1} admits a viscosity solution if $c>c_0$ and no solution if $c c_0$. In this case, we determine the basin of the stable solution and investigate the long time behavior of the solution semigroup associated to \eqref{hjs1}. In particular, we relate the lowest convergence rate to the integral of $\lambda$ over Mather measures, which leads to an asymptotic behavior of Mather measures when $c$ goes to infinity. Assume $c\geqslant c_0$ and the equation admits a solution, we classify ergodic Mather measures and locate their distribution in the phase space.

math.AP

High-Order Perfect Absorption in the Absence of Exceptional Point

High-order perfect absorption of coherent input has recently attracted significant attention due to its broadband absorption capacity. However, the realization of a high-order perfect absorber relies on the exceptional point (EP) to coalesce the scattering zeros. Here, we present a general scattering framework and achieve the high-order perfect absorber in the absence of EP. We consider the asynchronous coherent input, where a spatial delay introduces a momentum-dependent phase factor beyond the amplitude and phase control in synchronous coherent input. This new degree of freedom enables active control of the momentum dependent output, effectively reshaping the absorption line shape necessary for the high-order perfect absorber. Remarkably, despite the absence of EP, the proposed high-order perfect absorber exhibits significant response to the perturbations in the delay length. Our findings provide insights for the delay induced momentum-sensitive interference phenomenon and offer a new route for wave control.

physics.optics

Impurity-induced topological decomposition

Controlling topological phases is a central goal in quantum materials and related fields, enabling applications such as robust transport and programmable edge states. Here we uncover a mechanism in which local on-site impurities act as knobs to decompose global topological properties in discrete steps. In non-Hermitian lattices with spectral winding topology, we show that each impurity sequentially reduces the winding number by one, which is directly manifested as a stepwise decomposition of quantized plateaus in the steady-state response. Based on this principle, we further develop a scheme that sequentially induces topological edge states under impurity control, in a class of Hermitian topological systems constructed by doubling the non-Hermitian ones. Our findings reveal a general scheme to tune global topological properties with local perturbations, establishing a universal framework for impurity-controlled topological phases and offering a foundation for future exploration of reconfigurable topological phenomena across diverse physical platforms.

cond-mat.mes-hall

Droplet3D: Commonsense Priors from Videos Facilitate 3D Generation

Scaling laws have validated the success and promise of large-data-trained models in creative generation across text, image, and video domains. However, this paradigm faces data scarcity in the 3D domain, as there is far less of it available on the internet compared to the aforementioned modalities. Fortunately, there exist adequate videos that inherently contain commonsense priors, offering an alternative supervisory signal to mitigate the generalization bottleneck caused by limited native 3D data. On the one hand, videos capturing multiple views of an object or scene provide a spatial consistency prior for 3D generation. On the other hand, the rich semantic information contained within the videos enables the generated content to be more faithful to the text prompts and semantically plausible. This paper explores how to apply the video modality in 3D asset generation, spanning datasets to models. We introduce Droplet3D-4M, the first large-scale video dataset with multi-view level annotations, and train Droplet3D, a generative model supporting both image and dense text input. Extensive experiments validate the effectiveness of our approach, demonstrating its ability to produce spatially consistent and semantically plausible content. Moreover, in contrast to the prevailing 3D solutions, our approach exhibits the potential for extension to scene-level applications. This indicates that the commonsense priors from the videos significantly facilitate 3D creation. We have open-sourced all resources including the dataset, code, technical framework, and model weights: https://dropletx.github.io/.

cs.CV

Conditional Information Bottleneck for Multimodal Fusion: Overcoming Shortcut Learning in Sarcasm Detection

Multimodal sarcasm detection is a complex task that requires distinguishing subtle complementary signals across modalities while filtering out irrelevant information. Many advanced methods rely on learning shortcuts from datasets rather than extracting intended sarcasm-related features. However, our experiments show that shortcut learning impairs the model's generalization in real-world scenarios. Furthermore, we reveal the weaknesses of current modality fusion strategies for multimodal sarcasm detection through systematic experiments, highlighting the necessity of focusing on effective modality fusion for complex emotion recognition. To address these challenges, we construct MUStARD++$^{R}$ by removing shortcut signals from MUStARD++. Then, a Multimodal Conditional Information Bottleneck (MCIB) model is introduced to enable efficient multimodal fusion for sarcasm detection. Experimental results show that the MCIB achieves the best performance without relying on shortcut learning.

cs.LG

Dynamic Agile Reconfigurable Intelligent Surface Antenna (DARISA) MIMO: DoF Analysis and Effective DoF Optimization

In this paper, we propose a dynamic agile reconfigurable intelligent surface antenna (DARISA) array integrated into multi-input multi-output (MIMO) transceivers. Each DARISA comprises a number of metasurface elements activated simultaneously via a parallel feed network. The proposed system enables rapid and intelligent phase response adjustments for each metasurface element within a single symbol duration, facilitating a dynamic agile adjustment of phase response (DAAPR) strategy. By analyzing the theoretical degrees of freedom (DoF) of the DARISA MIMO system under the DAAPR framework, we derive an explicit relationship between DoF and critical system parameters, including agility frequentness (i.e., the number of phase adjustments of metasurface elements during one symbol period), cluster angular spread of wireless channels, DARISA array size, and the number of transmit/receive DARISAs. The DoF result reveals a significant conclusion: when the number of receive DARISAs is smaller than that of transmit DARISAs, the DAAPR strategy of the DARISA MIMO enhances the overall system DoF. Furthermore, relying on DoF alone to measure channel capacity is insufficient, so we analyze the effective DoF (EDoF) that reflects the impacts of the DoF and channel matrix singular value distribution on capacity. We show channel capacity monotonically increases with EDoF, and optimize the agile phase responses of metasurface elements by using fractional programming (FP) and semidefinite relaxation (SDR) algorithms to maximize the EDoF. Simulations validate the theoretical DoF gains and reveal that increasing agility frequentness, metasurface element density, and phase quantization accuracy can enhance the EDoF. Additionally, densely deployed elements can compensate for the loss in communication performance caused by lower phase quantization accuracy.

cs.IT

Large Language Model Empowered Design of Fluid Antenna Systems: Challenges, Frameworks, and Case Studies for 6G

The Fluid Antenna System (FAS), which enables flexible Multiple-Input Multiple-Output (MIMO) communications, introduces new spatial degrees of freedom for next-generation wireless networks. Unlike traditional MIMO, FAS involves joint port selection and precoder design, a combinatorial NP-hard optimization problem. Moreover, fully leveraging FAS requires acquiring Channel State Information (CSI) across its ports, a challenge exacerbated by the system's near-continuous reconfigurability. These factors make traditional system design methods impractical for FAS due to nonconvexity and prohibitive computational complexity. While deep learning (DL)-based approaches have been proposed for MIMO optimization, their limited generalization and fitting capabilities render them suboptimal for FAS. In contrast, Large Language Models (LLMs) extend DL's capabilities by offering general-purpose adaptability, reasoning, and few-shot learning, thereby overcoming the limitations of task-specific, data-intensive models. This article presents a vision for LLM-driven FAS design, proposing a novel flexible communication framework. To demonstrate the potential, we examine LLM-enhanced FAS in multiuser scenarios, showcasing how LLMs can revolutionize FAS optimization.

cs.IT

Synchronized in-gap edge states and robust copropagation in topological insulators without magnetic flux

Copropagation of antichiral edge states in the metallic phase requires the bulk states as counterpropagating modes. Without the band gap protection, the copropagation along the boundaries is easily scattered into the bulk and the counterpropagation in the bulk is not robust against disorder and defects. Here, we propose a novel time-reversal symmetric topological insulator holding the synchronized in-gap edge states. To prevent the participation of bulk states, we introduce the time-reversal symmetry to detach the edge states from the bulk band, then the robust copropagation is realized through widening the band gap across a metal-insulator transition with anisotropic next-nearest-neighbor couplings. The time-reversal symmetry ensures the in-gap edge states with opposite momenta as the counterpropagating modes. The inversion symmetry protects the quantized polarization as a topological invariant. The synchronized in-gap edge states not only enrich the family of topological edge states, but also provide additional flexibility in the design of reconfigurable topological optical devices. Our findings open a new avenue for the tailored robust wave transport using the in-gap edge states for future acousto-optic topological metamaterials.

cond-mat.mes-hall

Probing Bulk Band Topology from Time Boundary Effect in Synthetic Dimension

An incident wave at a temporal interface, created by an abrupt change in system parameters, generates time-refracted and time-reflected waves. We find topological characteristics associated with the temporal interface that separates distinct spatial topologies and report a novel bulk-boundary correspondence for the temporal interface. The vanishing of either time refraction or time reflection records a topological phase transition across the temporal interface, and the difference of bulk band topology predicts nontrivial braiding hidden in the time refraction and time reflection coefficients. These findings, which are insensitive to spatial boundary conditions and robust against disorder, are demonstrated in a synthetic frequency lattice with rich topological phases engendered by long-range couplings. Our work reveals the topological aspect of temporal interface and paves the way for using the time boundary effect to probe topological phase transitions and topological invariants.

physics.optics

DropletVideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation

Spatio-temporal consistency is a critical research topic in video generation. A qualified generated video segment must ensure plot plausibility and coherence while maintaining visual consistency of objects and scenes across varying viewpoints. Prior research, especially in open-source projects, primarily focuses on either temporal or spatial consistency, or their basic combination, such as appending a description of a camera movement after a prompt without constraining the outcomes of this movement. However, camera movement may introduce new objects to the scene or eliminate existing ones, thereby overlaying and affecting the preceding narrative. Especially in videos with numerous camera movements, the interplay between multiple plots becomes increasingly complex. This paper introduces and examines integral spatio-temporal consistency, considering the synergy between plot progression and camera techniques, and the long-term impact of prior content on subsequent generation. Our research encompasses dataset construction through to the development of the model. Initially, we constructed a DropletVideo-10M dataset, which comprises 10 million videos featuring dynamic camera motion and object actions. Each video is annotated with an average caption of 206 words, detailing various camera movements and plot developments. Following this, we developed and trained the DropletVideo model, which excels in preserving spatio-temporal coherence during video generation. The DropletVideo dataset and model are accessible at https://dropletx.github.io.

cs.CV

On the dynamics of contact Hamiltonian systems II: Variational construction of asymptotic orbits

This paper is a continuation of our study of the dynamics of contact Hamiltonian systems in \cite{JY}, but without monotonicity assumption. Due to the complexity of general cases, we focus on the behavior of action minimizing orbits. We pick out certain action minimizing invariant sets $\{\widetilde{\mathcal{N}}_u\}$ in the phase space naturally stratified by solutions $u$ to the corresponding Hamilton-Jacobi equation. Using an extension of characteristic method, we establish the existence of semi-infinite orbits that is asymptotic to some $\widetilde{\mathcal{N}}_u$ and heteroclinic orbits between $\widetilde{\mathcal{N}}_u$ and $\widetilde{\mathcal{N}}_v$ for two different solutions $u$ and $v$.

math.DS

TDGCN-Based Mobile Multiuser Physical-Layer Authentication for EI-Enabled IIoT

Physical-Layer Authentication (PLA) offers endogenous security, lightweight implementation, and high reliability, making it a promising complement to upper-layer security methods in Edge Intelligence (EI)-empowered Industrial Internet of Things (IIoT). However, state-of-the-art Channel State Information (CSI)-based PLA schemes face challenges in recognizing mobile multi-users due to the constantly shifting CSI distributions with user movements. To address this issue, we propose a Temporal Dynamic Graph Convolutional Network (TDGCN)-based PLA scheme, which employs Graph Neural Networks (GNNs) to capture the spatio-temporal dynamics induced by user movements. Firstly, we partition CSI fingerprints into multivariate time series and utilize dynamic GNNs to capture their associations. Secondly, Temporal Convolutional Networks (TCNs) handle temporal dependencies within each CSI fingerprint dimension. Additionally, Dynamic Graph Isomorphism Networks (GINs) and cascade node clustering pooling further enable efficient information aggregation and reduced computational complexity. Simulations demonstrate the proposed scheme's superior authentication accuracy compared to seven baseline schemes.

eess.SP

Infer Induced Sentiment of Comment Response to Video: A New Task, Dataset and Baseline

Existing video multi-modal sentiment analysis mainly focuses on the sentiment expression of people within the video, yet often neglects the induced sentiment of viewers while watching the videos. Induced sentiment of viewers is essential for inferring the public response to videos, has broad application in analyzing public societal sentiment, effectiveness of advertising and other areas. The micro videos and the related comments provide a rich application scenario for viewers induced sentiment analysis. In light of this, we introduces a novel research task, Multi-modal Sentiment Analysis for Comment Response of Video Induced(MSA-CRVI), aims to inferring opinions and emotions according to the comments response to micro video. Meanwhile, we manually annotate a dataset named Comment Sentiment toward to Micro Video (CSMV) to support this research. It is the largest video multi-modal sentiment dataset in terms of scale and video duration to our knowledge, containing 107,267 comments and 8,210 micro videos with a video duration of 68.83 hours. To infer the induced sentiment of comment should leverage the video content, so we propose the Video Content-aware Comment Sentiment Analysis (VC-CSA) method as baseline to address the challenges inherent in this new task. Extensive experiments demonstrate that our method is showing significant improvements over other established baselines.

cs.CV

On the causal discontinuity of Morse spacetimes

Morse spacetime is a model of singular Lorentzian manifold, built upon a Morse function which serves as a global time function outside its critical points. The Borde-Sorkin conjecture states that a Morse spacetime is causally continuous if and only if the index and coindex of critical points of the corresponding Morse function are both different from 1. The conjecture has recently been confirmed by Garcia Heveling for the case of small anisotropy and Euclidean background metric. Here, we provide a complementary counterexample: a four dimensional Morse spacetime whose critical point has index 2 and large enough anisotropy is causally discontinuous and thus the Borde-Sorkin conjecture does not hold. The proof features a low regularity causal structure and causal bubbling.

math.DG

Deep Rib Fracture Instance Segmentation and Classification from CT on the RibFrac Challenge

Rib fractures are a common and potentially severe injury that can be challenging and labor-intensive to detect in CT scans. While there have been efforts to address this field, the lack of large-scale annotated datasets and evaluation benchmarks has hindered the development and validation of deep learning algorithms. To address this issue, the RibFrac Challenge was introduced, providing a benchmark dataset of over 5,000 rib fractures from 660 CT scans, with voxel-level instance mask annotations and diagnosis labels for four clinical categories (buckle, nondisplaced, displaced, or segmental). The challenge includes two tracks: a detection (instance segmentation) track evaluated by an FROC-style metric and a classification track evaluated by an F1-style metric. During the MICCAI 2020 challenge period, 243 results were evaluated, and seven teams were invited to participate in the challenge summary. The analysis revealed that several top rib fracture detection solutions achieved performance comparable or even better than human experts. Nevertheless, the current rib fracture classification solutions are hardly clinically applicable, which can be an interesting area in the future. As an active benchmark and research resource, the data and online evaluation of the RibFrac Challenge are available at the challenge website. As an independent contribution, we have also extended our previous internal baseline by incorporating recent advancements in large-scale pretrained networks and point-based rib segmentation techniques. The resulting FracNet+ demonstrates competitive performance in rib fracture detection, which lays a foundation for further research and development in AI-assisted rib fracture detection and diagnosis.

eess.IV