SearcharxivSearch

arXiv subjects

Miao Jiang

Publications and source records attributed to Miao Jiang.

14 recordsLinked to original sources

MovieTeller: Tool-augmented Movie Synopsis with ID Consistent Progressive Abstraction

With the explosive growth of digital entertainment, automated video summarization has become indispensable for applications such as content indexing, personalized recommendation, and efficient media archiving. Automatic synopsis generation for long-form videos, such as movies and TV series, presents a significant challenge for existing Vision-Language Models (VLMs). While proficient at single-image captioning, these general-purpose models often exhibit critical failures in long-duration contexts, primarily a lack of ID-consistent character identification and a fractured narrative coherence. To overcome these limitations, we propose MovieTeller, a novel framework for generating movie synopses via tool-augmented progressive abstraction. Our core contribution is a training-free, tool-augmented, fact-grounded generation process. Instead of requiring costly model fine-tuning, our framework directly leverages off-the-shelf models in a plug-and-play manner. We first invoke a specialized face recognition model as an external "tool" to establish Factual Groundings--precise character identities and their corresponding bounding boxes. These groundings are then injected into the prompt to steer the VLM's reasoning, ensuring the generated scene descriptions are anchored to verifiable facts. Furthermore, our progressive abstraction pipeline decomposes the summarization of a full-length movie into a multi-stage process, effectively mitigating the context length limitations of current VLMs. Experiments demonstrate that our approach yields significant improvements in factual accuracy, character consistency, and overall narrative coherence compared to end-to-end baselines.

cs.CV

GridCodex: A RAG-Driven AI Framework for Power Grid Code Reasoning and Compliance

The global shift towards renewable energy presents unprecedented challenges for the electricity industry, making regulatory reasoning and compliance increasingly vital. Grid codes, the regulations governing grid operations, are complex and often lack automated interpretation solutions, which hinders industry expansion and undermines profitability for electricity companies. We introduce GridCodex, an end to end framework for grid code reasoning and compliance that leverages large language models and retrieval-augmented generation (RAG). Our framework advances conventional RAG workflows through multi stage query refinement and enhanced retrieval with RAPTOR. We validate the effectiveness of GridCodex with comprehensive benchmarks, including automated answer assessment across multiple dimensions and regulatory agencies. Experimental results showcase a 26.4% improvement in answer quality and more than a 10 fold increase in recall rate. An ablation study further examines the impact of base model selection.

cs.AI

Double Splay Nematic Order in Confined Polar Fluids

In this study, we demonstrate that when a ferroelectric nematic is confined between two glass plates coated with ionic polymers, a modulated phase emerges in a narrow temperature range between the nematic and ferroelectric nematic phases. This modulated phase emerges from the nematic phase in a continuous manner and then transforms into the ferroelectric nematic phase via a first-order transition upon cooling. Using optical microscopy, we provide compelling evidence that this modulated phase corresponds to the theoretically predicted double splay nematic phase. In this phase, splay deformations alternate in two orthogonal directions oriented at 45{\deg} to the substrate surfaces, creating a modulation wavelength that is twice the thickness of the cell. Our experiments with different ionic coatings reveal that only polymeric cationic coatings effectively promote the formation of this phase, highlighting the critical role of electrical screening. These findings not only confirm the existence of the double splay nematic phase but also provide insights into the distinctive topological defects of this phase in confined geometries.

cond-mat.soft

Chiral {\pi} Domain Walls Composed of Twin Half-Integer Surface Disclinations in Ferroelectric Nematic Liquid Crystals

Ferroelectric nematic liquid crystals are polar fluids characterized by microscopic orientational ordering and macroscopic spontaneous polarizations. Within these fluids, walls that separate domains of different polarizations are ubiquitous. We demonstrate that the {\pi} walls in films of polar fluids consist of twin half-integer surface disclinations spaced horizontally, enclosing a subdomain where the polarization exhibits left- or right-handed {\pi} twists across the film. The degenerate geometric configurations of these twin disclinations give rise to kinks and antikinks, effectively partitioning subdomains of opposite chirality like Ising chains. The hierarchical topological structures dictate that field-driven polar switching entails a two-step annihilation process of the disclinations. These findings serve as a cornerstone for comprehending other walls in ferroelectric and ferromagnetic materials, thereby laying the base for domain engineering crucial for advancing their nonlinear and optoelectronic applications.

cond-mat.soft

Half-integer Topological Defects Paired via String Micelles in Polar Liquids

Ferroelectric nematic (NF) liquid crystals present a compelling platform for exploring topological defects in polar fields, while their structural properties can be significantly altered by ionic doping. In this study, we demonstrate that doping the ferroelectric nematic material RM734 with cationic polymers enable the formation of polymeric micelles that connect pairs of half-integer topological defects. Polarizing optical microscopy reveals that these string defects exhibit butterfly textures, featured with a two-dimensional polarization field divided by N\'eel-type kink-walls into domains exhibiting either uniform polarization or negative splay and bend deformations. Through analysis of electrophoretic motion and direct measurements of polarization divergences, we show that the string micelles are positively charged and their side regions exhibit positive bound charges. To elucidate these observations, we propose a charge double layer model for the string defects: the positive charged cationic polymer chains and densely packed RM734 molecules form a Stern charge layer, while small anionic ions and positive bound charges constitute the charge diffusion layer. Notably, our experiments indicate that only cationic polymer doping effectively induces the formation of these unique string defects. These findings enhance our understanding of ionic doping effects and provide valuable insights for engineering polar topologies in liquid crystal systems.

cond-mat.soft

The enhancement of pair production in oscillated overlapped fields

The influence of potential well width on electron-positron pair production has been examined through theoretical and numerical approaches by employing the computational quantum field theory. Quantum interference effects in pair production is investigated in the two overlapped potential wells with varied widths and frequencies. Several dominant processes, involving the absorption of an integer number of photons, significantly impact on pair production. Notably, specific multiphoton absorption processes exhibit distinct changes as the potential well width expands, with the absorption of four photons process displaying noteworthy effects. Additionally, the influence of the smaller frequency to the yield of the pair production can not be ignored and the most optimized frequencies in our overlapped fields has been studied and exhibited.

hep-ph

Line defects in nematic liquid crystals as charged superelastic rods with negative twist--stretch coupling

Topological defects are a ubiquitous phenomenon in diverse physical systems. In nematic liquid crystals (LCs), they are dynamic, physicochemically distinct, sensitive to stimuli, and are thereby promising for a range of applications. However, our current understanding of the mechanics and dynamics of defects in nematic LCs remain limited and are often overwhelmed by the intricate details of the specific systems. Here, we unify singular and nonsingular line defects as superelastic rods and combine theory, simulation, and experiment to quantitatively measure their effective elastic moduli, including line tension, torsional rigidity, and twist--stretch coefficient. Interestingly, we found that line defects exhibit a negative twist--stretch coupling, meaning that twisted line defects tend to unwind under stretching, which is reminiscent of DNA molecules. A patterned nematic cell experiment further confirmed the above findings. Taken together, we have established an effective elasticity theory for nematic defects, paving the way towards understanding and engineering their deformation and transformation in driven and active nematic materials.

cond-mat.soft

Multi-user Co-inference with Batch Processing Capable Edge Server

Graphics processing units (GPUs) can improve deep neural network inference throughput via batch processing, where multiple tasks are concurrently processed. We focus on novel scenarios that the energy-constrained mobile devices offload inference tasks to an edge server with GPU. The inference task is partitioned into sub-tasks for a finer granularity of offloading and scheduling, and the user energy consumption minimization problem under inference latency constraints is investigated. To deal with the coupled offloading and scheduling introduced by concurrent batch processing, we first consider an offline problem with a constant edge inference latency and the same latency constraint. It is proven that optimizing the offloading policy of each user independently and aggregating all the same sub-tasks in one batch is optimal, and thus the independent partitioning and same sub-task aggregating (IP-SSA) algorithm is inspired. Further, the optimal grouping (OG) algorithm is proposed to optimally group tasks when the latency constraints are different. Finally, when future task arrivals cannot be precisely predicted, a deep deterministic policy gradient (DDPG) agent is trained to call OG. Experiments show that IP-SSA reduces up to 94.9\% user energy consumption in the offline setting, while DDPG-OG outperforms DDPG-IP-SSA by up to 8.92\% in the online setting.

cs.DC

Hierarchical Graph Convolutional Skeleton Transformer for Action Recognition

Graph convolutional networks (GCNs) have emerged as dominant methods for skeleton-based action recognition. However, they still suffer from two problems, namely, neighborhood constraints and entangled spatiotemporal feature representations. Most studies have focused on improving the design of graph topology to solve the first problem but they have yet to fully explore the latter. In this work, we design a disentangled spatiotemporal transformer (DSTT) block to overcome the above limitations of GCNs in three steps: (i) feature disentanglement for spatiotemporal decomposition;(ii) global spatiotemporal attention for capturing correlations in the global context; and (iii) local information enhancement for utilizing more local information. Thereon, we propose a novel architecture, named Hierarchical Graph Convolutional skeleton Transformer (HGCT), to employ the complementary advantages of GCN (i.e., local topology, temporal dynamics and hierarchy) and Transformer (i.e., global context and dynamic attention). HGCT is lightweight and computationally efficient. Quantitative analysis demonstrates the superiority and good interpretability of HGCT.

cs.CV

Joint Device Scheduling and Resource Allocation for Latency Constrained Wireless Federated Learning

In federated learning (FL), devices contribute to the global training by uploading their local model updates via wireless channels. Due to limited computation and communication resources, device scheduling is crucial to the convergence rate of FL. In this paper, we propose a joint device scheduling and resource allocation policy to maximize the model accuracy within a given total training time budget for latency constrained wireless FL. A lower bound on the reciprocal of the training performance loss, in terms of the number of training rounds and the number of scheduled devices per round, is derived. Based on the bound, the accuracy maximization problem is solved by decoupling it into two sub-problems. First, given the scheduled devices, the optimal bandwidth allocation suggests allocating more bandwidth to the devices with worse channel conditions or weaker computation capabilities. Then, a greedy device scheduling algorithm is introduced, which in each step selects the device consuming the least updating time obtained by the optimal bandwidth allocation, until the lower bound begins to increase, meaning that scheduling more devices will degrade the model accuracy. Experiments show that the proposed policy outperforms state-of-the-art scheduling policies under extensive settings of data distributions and cell radius.

cs.IT

Joint Beamforming Design in Multi-Cluster MISO NOMA Intelligent Reflecting Surface-Aided Downlink Communication Networks

Considering intelligent reflecting surface (IRS), we study a multi-cluster multiple-input-single-output (MISO) non-orthogonal multiple access (NOMA) downlink communication network. In the network, an IRS assists the communication from the base station (BS) to all users by passive beamforming. Our goal is to minimize the total transmit power by jointly optimizing the transmit beamforming vectors at the BS and the reflection coefficient vector at the IRS. Because of the restrictions on the IRS reflection amplitudes and phase shifts, the formulated quadratically constrained quadratic problem is highly non-convex. For the aforementioned problem, the conventional semidefinite programming (SDP) based algorithm has prohibitively high computational complexity and deteriorating performance. Here, we propose an effective second-order cone programming (SOCP)-alternating direction method of multipliers (ADMM) based algorithm to obtain the locally optimal solution. To reduce the computational complexity, we also propose a low-complexity zero-forcing (ZF) based suboptimal algorithm. It is shown through simulation results that our proposed SOCP-ADMM based algorithm achieves significant performance gain over the conventional SDP based algorithm. Furthermore, when the number of passive reflection elements is relatively high, our proposed ZF-based suboptimal algorithm also outperforms the SDP based algorithm.

cs.IT

Beamforming Design in Multiple-Input-Multiple-Output Symbiotic Radio Backscatter Systems

Symbiotic radio (SR) backscatter systems are possible techniques for the future low-power wireless communications for Internet of Things devices. In this paper, we propose a multiple-input-multiple-output (MIMO) SR backscatter system, where the secondary multi-antenna transmission from the backscatter device (BD) to the receiver is riding on the primary multi-antenna transmission from the transmitter to the receiver. We investigate the beamforming design optimization problem which maximizes the achievable rate of secondary transmission under the achievable rate constraint of primary transmission. In the MIMO SR backscatter system, each antenna of the SR BD reflects its received ambient radio frequency signals from all the transmitting antennas of the transmitter, which causes the globally optimal solution is difficult to obtain. In this paper, we propose a method to obtain the achievable rate upper bound. Furthermore, considering both primary and secondary transmissions, we propose an exact penalty method based locally optimal solution. Simulation results illustrate that our proposed exact penalty method based locally optimal solution performs close to the upper bound.

cs.IT

Secure Beamforming in MISO NOMA Backscatter Device Aided Symbiotic Radio Networks

Symbiotic radio (SR) networks are possible solutions to the future low-power wireless communications for massive Internet of Things devices. In this paper, we investigate a multiple-input-single-output non-orthogonal multiple access (NOMA) backscatter device (BD) aided SR network with a potential eavesdropper. In the network, a base station (BS) broadcasts signals to a central user and a cell-edge user using the NOMA protocol. With ambient backscatter modulation, the BD transmits its own messages to the central user over incident signals from the BS. We propose a constrained concave convex procedure-based algorithm which maximizes the $ε$-outage secrecy rate from the BD to the central user under the achievable secrecy rate constraints from the BS to the central and cell-edge users. Simulation results illustrate that our proposed network achieves a much larger secrecy rate region than the orthogonal multiple access (OMA) network.

cs.IT

Achieving Fairness in Determining Medicaid Eligibility through Fairgroup Construction

Effective complements to human judgment, artificial intelligence techniques have started to aid human decisions in complicated social problems across the world. In the context of United States for instance, automated ML/DL classification models offer complements to human decisions in determining Medicaid eligibility. However, given the limitations in ML/DL model design, these algorithms may fail to leverage various factors for decision making, resulting in improper decisions that allocate resources to individuals who may not be in the most need. In view of such an issue, we propose in this paper the method of \textit{fairgroup construction}, based on the legal doctrine of \textit{disparate impact}, to improve the fairness of regressive classifiers. Experiments on American Community Survey dataset demonstrate that our method could be easily adapted to a variety of regressive classification models to boost their fairness in deciding Medicaid Eligibility, while maintaining high levels of classification accuracy.

cs.LG