SearcharxivSearch

arXiv subjects

Jingwei Liu

Publications and source records attributed to Jingwei Liu.

At least 19 recordsLinked to original sources

An AoI-oriented Time-Frequency Distributed Access Mechanism in Wireless Sensor Networks with Spectrum Division

The increasing adoption of spectrum-division techniques enables concurrent uplink transmissions over multiple orthogonal resources, yet low-overhead access design with effective information freshness remains insufficiently studied for large-scale randomly activated sensor networks. In this paper, we apply the age of information (AoI) to measure information freshness and propose an AoI-efficient deterministic time-frequency distributed access (D-TFDA) mechanism. D-TFDA combines centralized configuration and distributed operation through a periodic token-based time-frequency structure, which provides sensors with collision-free and predictable transmission opportunities without considerable run-time overhead. We develop an analytical framework to characterize the long-term average AoI (AAoI) by exploiting the periodicity of the token assignment pattern and modeling the steady local state of each sensor with a one-dimensional discrete-time Markov chain (DTMC). We further reveal structural properties of the token assignment pattern and identify AoI-equivalent token clusters, which substantially reduce the search space of the AAoI-optimal token allocation problem. Based on this structure, we formulate the reduced problem as a linear programming (LP) problem and develop an AAoI-optimal search algorithm, together with an auction-inspired heuristic algorithm of lower complexity. Simulation results validate the proposed AAoI analysis, demonstrate the effectiveness of the token allocation algorithms, and show that D-TFDA achieves substantially lower AAoI than optimized random access baselines by avoiding collisions and exploiting heterogeneous sensor--resource transmission reliability.

cs.IT

CausalMoE: A Billion-Scale Multimodal Foundation Model for Granger Causal Discovery with Pattern-Routed Heterogeneous Experts

Granger Causal Discovery (GCD) is fundamental for analyzing temporal dependencies in complex systems. However, existing neural GCD methods predominantly rely on a "one-size-fits-all" paradigm, struggling to capture distribution shifts and dynamic regime changes inherent in real-world time series. This often leads to entangled representations and spurious causal graphs. In this paper, we propose CausalMoE, a billion-scale multimodal Granger causal foundation model that explicitly models patch-level heterogeneity. CausalMoE introduces a Pattern-Routed Mixture of Heterogeneous Experts, which dynamically identifies latent temporal patterns and routes patches to specialized domain experts, effectively decoupling regime-specific mechanisms from shared dynamics. To ensure interpretable graph recovery, we design a Causality-Aware Self-Attention mechanism operating across variables, yielding sparse Granger causal graphs via proximal optimization. Furthermore, CausalMoE is the first to integrate LLMs and VLMs to align numerical signals with textual and visual priors, regularizing causal estimation in complex scenarios. Extensive experiments demonstrate that CausalMoE establishes a new state-of-the-art on fully supervised benchmarks, while effectively generalizing to few-shot settings where traditional methods fail.

cs.LG

Preacher: Paper-to-Video Agentic System

The paper-to-video task converts a research paper into a structured video abstract, distilling key concepts, methods, and conclusions into an accessible, well-organized format. While state-of-the-art video generation models demonstrate potential, they are constrained by limited context windows, rigid video duration constraints, limited stylistic diversity, and an inability to represent domain-specific knowledge. To address these limitations, we introduce Preacher, the first paper-to-video agentic system. Preacher employs a topdown approach to decompose, summarize, and reformulate the paper, followed by bottom-up video generation, synthesizing diverse video segments into a coherent abstract. To align cross-modal representations, we define key scenes and introduce a Progressive Chain of Thought (P-CoT) for granular, iterative planning. Preacher successfully generates high-quality video abstracts across five research fields, demonstrating expertise beyond current video generation models. Code will be released at: https://github.com/Gen-Verse/Paper2Video

cs.CV

A Truncated Primordial Power Spectrum and its Impact on CMB Polarization

We investigate the impact of a hypothesized delayed initiation of inflation, characterized by a cutoff k_min to the primordial power spectrum in the cosmic microwave background (CMB). This cutoff affects both the scalar and tensor spectra, which therefore impacts several measurements of the temperature and polarization distributions. We calculate the angular power spectrum and correlation function with and without k_min in the context of Planck-LCDM, and demonstrate that a non-zero k_min significantly improves the alignment between theory and the observations, including the temperature, E-mode polarization, TE cross-correlation, Q+U polarization and Q-U polarization. It creates an observable signature in both the angular power spectrum and correlation function for all cases. We thus also explore the B-mode polarization, for which current data are not yet precise enough to determine k_min, but whose impact should be detectable with high-precision measurements using future missions, such as LiteBIRD, if the tensor-to-scalar ratio, r, is not much smaller than its current upper limit. We find that the introduction of k_min not only addresses large-angle anomalies in the CMB but also provides a more consistent framework for understanding the early Universe's inflationary phase. These findings highlight the importance of future high-precision CMB observations in validating the existence and implications of k_min.

astro-ph.CO

Expressive Music Data Processing and Generation

Musical expressivity and coherence are indispensable in music composition and performance, while often neglected in modern AI generative models. In this work, we introduce a listening-based data-processing technique that captures the expressivity in musical performance. This technique derived from Weber's law reflects the human perceptual truth of listening and preserves musical subtlety and expressivity in the training input. To facilitate musical coherence, we model the output interdependencies among multiple arguments in the music data such as pitch, duration, velocity, etc. in the neural networks based on the probabilistic chain rule. In practice, we decompose the multi-output sequential model into single-output submodels and condition previously sampled outputs on the subsequent submodels to induce conditional distributions. Finally, to select eligible sequences from all generations, a tentative measure based on the output entropy was proposed. The entropy sequence is set as a criterion to select predictable and stable generations, which is further studied under the context of informational aesthetic measures to quantify musical pleasure and information gain along the music tendency.

cs.SD

Optimizing Information Freshness of IEEE 802.11ax Uplink OFDMA-Based Random Access

The latest WiFi standard, IEEE 802.11ax (WiFi 6), introduces a novel uplink random access mechanism called uplink orthogonal frequency division multiple access-based random access (UORA). While existing work has evaluated the performance of UORA using conventional performance metrics, such as throughput and delay, its information freshness performance has not been thoroughly investigated in the literature. This is of practical significance as WiFi 6 and beyond are expected to support real-time applications. This paper presents the first attempt to fill this gap by investigating the information freshness, quantified by the Age of Information (AoI) metric, in UORA networks. We establish an analytical framework comprising two discrete-time Markov chains (DTMCs) to characterize the transmission states of stations (STAs) in UORA networks. Building on the formulated DTMCs, we derive an analytical expression for the long-term average AoI (AAoI), facilitating the optimization of UORA parameters for enhanced AoI performance through exhaustive search. To gain deeper design insights and improve the effectiveness of UORA parameter optimization, we derive a closed-form expression for the AAoI and its approximated lower bound for a simplified scenario characterized by a fixed backoff contention window and generate-at-will status updates. By analyzing the approximated lower bound of the AAoI, we propose efficient UORA parameter optimization algorithms that can be realized with only a few comparisons of different possible values of the parameters to be optimized. Simulation results validate our analysis and demonstrate that the AAoI achieved through our proposed parameter optimization algorithm closely approximates the optimal AoI performance obtained via exhaustive search, outperforming the round-robin and max-AoI policies in large and low-traffic networks.

cs.IT

Unfaithful Probability Distributions in Binary Triple of Causality Directed Acyclic Graph

Faithfulness is the foundation of probability distribution and graph in causal discovery and causal inference. In this paper, several unfaithful probability distribution examples are constructed in three--vertices binary causality directed acyclic graph (DAG) structure, which are not faithful to causal DAGs described in J.M.,Robins,et al. Uniform consistency in causal inference. Biometrika (2003),90(3): 491--515. And the general unfaithful probability distribution with multiple independence and conditional independence in binary triple causal DAG is given.

stat.ML

Optimizing AoI at Query in Multiuser Wireless Uplink Networks: A Whittle Index Approach

In this paper, we explore how to schedule multiple users to optimize information freshness in a pull-based wireless network, where the status updates from users are requested by randomly arriving queries at the destination. We use the age of information at query (QAoI) to characterize the performance of information freshness. Such a decision-making problem is naturally modeled as a Markov decision process (MDP), which, however, is prohibitively high to be solved optimally by the standard method due to the curse of dimensionality. To address this issue, we employ Whittle index approach, which allows us to decouple the original MDP into multiple sub-MDPs by relaxing the scheduling constraints. However, the binary Markovian query arrival process results in a bi-dimensional state and complex state transitions within each sub-MDP, making it challenging to verify Whittle indexability using conventional methods. After a thorough analysis of the sub-MDP's structure, we show that it is unichain and its optimal policy follows a threshold-type structure. This facilitates the verification of Whittle indexability of the sub-MDP by employing an easy-to-verify condition. Subsequently, the steady-state probability distributions of the sub-MDP under different threshold-type policies are analyzed, constituting the analytical expressions of different Whittle indices in terms of the expected average QAoI and scheduling time of the sub-MDP. Building on these, we devise an efficient algorithm to calculate Whittle indices for the formulated sub-MDPs. The simulation results validate our analyses and show the proposed Whittle index policy outperforms baseline policies and achieves near-optimal performance.

cs.IT

Retrieval-Augmented Diffusion Models for Time Series Forecasting

While time series diffusion models have received considerable focus from many recent works, the performance of existing models remains highly unstable. Factors limiting time series diffusion models include insufficient time series datasets and the absence of guidance. To address these limitations, we propose a Retrieval- Augmented Time series Diffusion model (RATD). The framework of RATD consists of two parts: an embedding-based retrieval process and a reference-guided diffusion model. In the first part, RATD retrieves the time series that are most relevant to historical time series from the database as references. The references are utilized to guide the denoising process in the second part. Our approach allows leveraging meaningful samples within the database to aid in sampling, thus maximizing the utilization of datasets. Meanwhile, this reference-guided mechanism also compensates for the deficiencies of existing time series diffusion models in terms of guidance. Experiments and visualizations on multiple datasets demonstrate the effectiveness of our approach, particularly in complicated prediction tasks.

cs.LG

Expressive MIDI-format Piano Performance Generation

This work presents a generative neural network that's able to generate expressive piano performance in MIDI format. The musical expressivity is reflected by vivid micro-timing, rich polyphonic texture, varied dynamics, and the sustain pedal effects. This model is innovative from many aspects of data processing to neural network design. We claim that this symbolic music generation model overcame the common critics of symbolic music and is able to generate expressive music flows as good as, if not better than generations with raw audio. One drawback is that, due to the limited time for submission, the model is not fine-tuned and sufficiently trained, thus the generation may sound incoherent and random at certain points. Despite that, this model shows its powerful generative ability to generate expressive piano pieces.

cs.SD

Challenges to Inflation in the Post-Planck Era

Space-based missions studying the cosmic microwave background (CMB) have progressively refined the parameter space in conventional models of inflation shortly ($\sim 10^{-37}$ seconds) after the big bang. While most inflationary scenarios proposed thus far in the context of GR have since been ruled out, the basic idea of inflation may still be tenable, albeit with several unresolved conundrums, such as conflicting initial conditions and inconsistencies with the measured CMB power spectrum. In the new slow-roll inflationary picture, inflation arising in plateau-like potentials requires an initiation beyond the Planck time. This delay may be consistent with the cutoff, $k_{\rm min}$, measured recently in the primordial power spectrum. However, the actual value of $k_{\rm min}$ would imply an initiation time too far beyond the big bang for inflation to solve the horizon problem. In this paper, we also describe several other undesirable consequences of this delay, including an absence of well motivated initial conditions and a significant difficulty providing a viable mechanism for properly quantizing the primordial fluctuations. Nevertheless, many of these inconsistencies may still be avoided if one introduces non-conventional modifications to inflation, such as a brief departure from slow-roll dynamics, possibly due to a dramatic change in the inflationary potential, inflation driven by multiple fields, or a non-minimal coupling to gravity. In addition, some of these difficulties could be mitigated via the use of alternative cosmologies based, e.g., on loop quantum gravity, which replaces the initial big-bang singularity with finite conditions at a bounce-like beginning.

astro-ph.CO

A Truncated Primordial Power Spectrum and Its Impact on B-Mode Polarization

The absence of large-angle correlations in the temperature of the cosmic microwave background (CMB), confirmed by three independent satellite missions, creates significant tension with the standard model of cosmology. Previous work has shown, however, that a truncation, $k_{min}$, of the primordial power spectrum comprehensively resolves the anomaly and the missing power at $\ell\lesssim 5$ (the low multipoles). Since this cutoff is consistent with the hypothesized delay of inflation well beyond the Planck time, we are strongly motivated to consider its possible impact on other observational signatures. In this Letter, we analyze and predict its influence on the most revealing probe awaiting measurement by upcoming missions -- the B-mode polarization of the CMB, whose accurate determination should greatly impact the inflationary picture. We highlight the quantitative power of this discriminant by specifically considering the LiteBIRD mission, predicting the effect of $k_{min}$ on both the angular power spectrum and the angular correlation function of the B-mode, for a range of tensor-to-scalar ratios, $r$. While its impact on the latter appears to be negligible, $k_{\rm min}$ should have a very pronounced effect on the former. We show that for $r=0.036$, $k_{min}$'s impact on $C_{\ell}^{BB}$ at low $\ell$'s should be easily detectable by LiteBIRD, but will be largely hidden by the total uncertainty of the measurement if $r\lesssim 0.02$.

astro-ph.CO

Structure-Guided Adversarial Training of Diffusion Models

Diffusion models have demonstrated exceptional efficacy in various generative applications. While existing models focus on minimizing a weighted sum of denoising score matching losses for data distribution modeling, their training primarily emphasizes instance-level optimization, overlooking valuable structural information within each mini-batch, indicative of pair-wise relationships among samples. To address this limitation, we introduce Structure-guided Adversarial training of Diffusion Models (SADM). In this pioneering approach, we compel the model to learn manifold structures between samples in each training batch. To ensure the model captures authentic manifold structures in the data distribution, we advocate adversarial training of the diffusion generator against a novel structure discriminator in a minimax game, distinguishing real manifold structures from the generated ones. SADM substantially improves existing diffusion transformers (DiT) and outperforms existing methods in image generation and cross-domain fine-tuning tasks across 12 datasets, establishing a new state-of-the-art FID of 1.58 and 2.11 on ImageNet for class-conditional image generation at resolutions of 256x256 and 512x512, respectively.

cs.CV

Contextualized Diffusion Models for Text-Guided Image and Video Generation

Conditional diffusion models have exhibited superior performance in high-fidelity text-guided visual generation and editing. Nevertheless, prevailing text-guided visual diffusion models primarily focus on incorporating text-visual relationships exclusively into the reverse process, often disregarding their relevance in the forward process. This inconsistency between forward and reverse processes may limit the precise conveyance of textual semantics in visual synthesis results. To address this issue, we propose a novel and general contextualized diffusion model (ContextDiff) by incorporating the cross-modal context encompassing interactions and alignments between text condition and visual sample into forward and reverse processes. We propagate this context to all timesteps in the two processes to adapt their trajectories, thereby facilitating cross-modal conditional modeling. We generalize our contextualized diffusion to both DDPMs and DDIMs with theoretical derivations, and demonstrate the effectiveness of our model in evaluations with two challenging tasks: text-to-image generation, and text-to-video editing. In each task, our ContextDiff achieves new state-of-the-art performance, significantly enhancing the semantic alignment between text condition and generated samples, as evidenced by quantitative and qualitative evaluations. Our code is available at https://github.com/YangLing0818/ContextDiff

cs.CV

Improving Diffusion-Based Image Synthesis with Context Prediction

Diffusion models are a new class of generative models, and have dramatically promoted image generation with unprecedented quality and diversity. Existing diffusion models mainly try to reconstruct input image from a corrupted one with a pixel-wise or feature-wise constraint along spatial axes. However, such point-based reconstruction may fail to make each predicted pixel/feature fully preserve its neighborhood context, impairing diffusion-based image synthesis. As a powerful source of automatic supervisory signal, context has been well studied for learning representations. Inspired by this, we for the first time propose ConPreDiff to improve diffusion-based image synthesis with context prediction. We explicitly reinforce each point to predict its neighborhood context (i.e., multi-stride features/tokens/pixels) with a context decoder at the end of diffusion denoising blocks in training stage, and remove the decoder for inference. In this way, each point can better reconstruct itself by preserving its semantic connections with neighborhood context. This new paradigm of ConPreDiff can generalize to arbitrary discrete and continuous diffusion backbones without introducing extra parameters in sampling procedure. Extensive experiments are conducted on unconditional image generation, text-to-image generation and image inpainting tasks. Our ConPreDiff consistently outperforms previous methods and achieves a new SOTA text-to-image generation results on MS-COCO, with a zero-shot FID score of 6.21.

cs.CV

Optimizing Information Freshness in Uplink Multiuser MIMO Networks with Partial Observations

This paper investigates a multiuser scheduling problem within an uplink multiple-input multi-output (MIMO) status update network, consisting of a multi-antenna base station (BS) and multiple single-antenna devices. The presence of multiple antennas at the BS introduces spatial degrees-of-freedom, enabling concurrent transmission of status updates from multiple devices in each time slot. Our objective is to optimize network-wide information freshness, quantified by the age of information (AoI) metric, by determining how the BS can best schedule device transmissions, while taking into account the random arrival of status updates at the device side.To address this decision-making problem, we model it as a partially observable Markov decision process (POMDP) and establish that the evolution of belief states for different devices is independent.We also prove that feasible belief states can be described by finite-dimensional vectors. Building on these observations, we develop a dynamic scheduling (DS) policy to solve the POMDP, and then derive an upper bound of its AoI performance, which is used to optimize the parameter configuration. To gain more design insights, we investigate a symmetric network, and put forth a fixed scheduling (FS) policy with lower computational complexity. An action space reduction strategy is applied to further reduce the computational complexity of both DS and FS policies. Our numerical results validate our analyses and indicate that the DS policy with the reduced action space performs almost identically to the original DS policy, and both outperform the baseline policies.

cs.IT

Local-Global Information Interaction Debiasing for Dynamic Scene Graph Generation

The task of dynamic scene graph generation (DynSGG) aims to generate scene graphs for given videos, which involves modeling the spatial-temporal information in the video. However, due to the long-tailed distribution of samples in the dataset, previous DynSGG models fail to predict the tail predicates. We argue that this phenomenon is due to previous methods that only pay attention to the local spatial-temporal information and neglect the consistency of multiple frames. To solve this problem, we propose a novel DynSGG model based on multi-task learning, DynSGG-MTL, which introduces the local interaction information and global human-action interaction information. The interaction between objects and frame features makes the model more fully understand the visual context of the single image. Long-temporal human actions supervise the model to generate multiple scene graphs that conform to the global constraints and avoid the model being unable to learn the tail predicates. Extensive experiments on Action Genome dataset demonstrate the efficacy of our proposed framework, which not only improves the dynamic scene graph generation but also alleviates the long-tail problem.

cs.CV

A Neural Network Implementation for Free Energy Principle

The free energy principle (FEP), as an encompassing framework and a unified brain theory, has been widely applied to account for various problems in fields such as cognitive science, neuroscience, social interaction, and hermeneutics. As a computational model deeply rooted in math and statistics, FEP posits an optimization problem based on variational Bayes, which is solved either by dynamic programming or expectation maximization in practice. However, there seems to be a bottleneck in extending the FEP to machine learning and implementing such models with neural networks. This paper gives a preliminary attempt at bridging FEP and machine learning, via a classical neural network model, the Helmholtz machine. As a variational machine learning model, the Helmholtz machine is optimized by minimizing its free energy, the same objective as FEP. Although the Helmholtz machine is not temporal, it gives an ideal parallel to the vanilla FEP and the hierarchical model of the brain, under which the active inference and predictive coding could be formulated coherently. Besides a detailed theoretical discussion, the paper also presents a preliminary experiment to validate the hypothesis. By fine-tuning the trained neural network through active inference, the model performance is promoted to accuracy above 99\%. In the meantime, the data distribution is continuously deformed to a salience that conforms to the model representation, as a result of active sampling.

cs.NE