SearcharxivSearch

arXiv subjects

Peng Zhang

Publications and source records attributed to Peng Zhang.

At least 19 recordsLinked to original sources

FedHUR: Learning Hierarchical Utility-Guided Client Relations for Personalized Federated Recommendation

Federated recommendation enables collaborative model training while keeping user interaction data on local clients. A central problem in federated recommendation is how to aggregate useful information across clients for personalized recommendation. Existing personalized aggregation methods usually construct client relations from predefined parameter-based assumptions, such as parameter similarity or complementarity, and use these relations to determine aggregation weights. However, such methods construct a single global relation, which is insufficient to capture the hierarchical and multi-granularity nature of user relations in recommendation. Moreover, these predefined relations cannot directly reflect whether the related clients can improve prediction performance after aggregation. To address these limitations, we propose FedHUR, a federated recommendation framework for learning hierarchical utility-guided client relations. FedHUR takes item-item filters as the object for relation construction and aggregation. Specifically, it first aggregates and clusters each client's local information to obtain global hierarchical information. Each client computes hierarchical utility signals based on its local information and the global hierarchical information, indicating which collaborative information is useful for improving its prediction. The server uses these utility signals to retrieve clients that are useful to that client for further personalized aggregation. Extensive experiments on five real-world datasets show that FedHUR consistently outperforms existing federated recommendation baselines, demonstrating the effectiveness of hierarchical utility-guided client relation learning. Code is available at https://github.com/Mingzhe-Han/FedHUR.

cs.IR

Search for neutrinoless quadruple beta decay of $^{136}$Xe in PandaX-4T detector

The observation of neutrinoless quadruple beta decay (0$\nu$4$\beta$) in the absence of neutrinoless double beta decay (0$\nu$2$\beta$) has been argued to provide a strong indication that neutrinos are Dirac particles. We report a search for 0$\nu$4$\beta$ decay of $^{136}\text{Xe}$ using a total $^{136}\text{Xe}$ exposure of 148.4 kg$\cdot$yr, collected during the commissioning and the first science runs of the PandaX-4T experiment. No significant excess of events over the background is observed. A lower limit on the 0$\nu$4$\beta$ decay half-life of $^{136}\text{Xe}$ is set at 6.01 x $10^{24}$ yr at the 90% confidence level. This result establishes the most stringent constraint on this process in xenon, demonstrating the unique capability of the PandaX-4T detector in probing lepton number violation and shedding light on the fundamental nature of neutrinos.

nucl-ex

Spatial-Code-Domain Grouped Index Modulation: Fluid-Antenna-Assisted System Design and BER Performance Analysis

Fluid antenna systems (FASs) provide reconfigurable spatial resources within compact apertures. In this paper, we introduce code-domain grouped index modulation (CGIM) and its spatial-code-domain extension, termed SCGIM, for FA-assisted transceivers. CGIM partitions the available orthogonal spreading codes into multiple subsets and jointly maps information onto their in-phase and quadrature indices and constellation symbols. In an Rx-FAS-assisted single-input multiple-output (SIMO) link, group-wise despreading separates the orthogonal code groups for parallel detection, while receive-port selection provides spatial diversity. SCGIM further associates interleaved Tx-FA port subsets with the code subsets, with the Tx-FAS conveying spatial-index information and the Rx-FAS providing selection diversity in a multiple-input multiple-output (MIMO) link. For SCGIM, we develop maximum-likelihood (ML), staged greedy (GD), and cross-domain index message-passing (CD-IMPD) detectors. CD-IMPD exchanges soft information over a cycle-free factor graph to account for the coupling between the spatial and code indices, requiring only one inward and one outward message pass. For CGIM, the BER is derived from the joint decision regions of the despread-domain observations and averaged over the Rx-FAS selected-gain distribution under Rayleigh, Nakagami-m, and additive white Gaussian noise channels. For SCGIM, an average-BER approximation is derived from a full-pair union bound using the selected-gain density ratio and exponentially tilted quadratic-form Laplace transforms. Simulation results validate the BER analysis and show that the proposed schemes achieve lower BER and higher throughput than the considered IM schemes, while CD-IMPD achieves near-ML BER performance with lower detection complexity.

cs.IT

ElderBench: Benchmarking Autonomous Mobile Agents for Older Adults

While autonomous mobile agents hold great potential for assisting older adults with smartphone usage, existing GUI benchmarks mainly rely on explicit, goal-oriented instructions and rarely capture the naturally occurring language patterns of older users, such as indirect speech, referential ambiguity, and under-specified requests. This mismatch between benchmark instructions and real-world elderly interactions may hinder reliable agent deployment. To address this gap, we present ElderBench, the first benchmark for evaluating mobile GUI agents in authentic elderly-oriented scenarios. ElderBench is constructed from 249 naturally elicited smartphone tasks collected from older adults across 20 applications. We first characterize the linguistic divergence between elderly instructions and existing GUI benchmark instructions from syntactic, semantic, and pragmatic perspectives. We then evaluate mainstream GUI agents and Vision-Language Models under both online and offline settings, revealing substantial performance degradation when handling elderly-oriented instructions. Through controlled instruction normalization, failure analysis, and fine-grained linguistic feature analysis, we further identify how elderly-specific language patterns contribute to agent failures. Our findings provide actionable design insights toward more adaptive, interpretable, and age-inclusive GUI agents for older adults.

cs.AI

Embedding Single-Phase Grid-Forming Inverters in Three-Phase Unbalanced Power Flow

This letter introduces Grid-Forming Three-Phase Power Flow (GFM-3PF), an augmented Newton power flow framework for modeling single-phase grid-forming inverters (GFMs) in three-phase unbalanced networks. The contributions of this work are threefold: 1) developing a two-stage solution strategy based on positive-sequence initialization followed by three-phase phase-domain lifting for robust large-scale computation; 2) incorporating phase-domain unbalanced network modeling with controlled mutual coupling together with phase-dependent, frequency-sensitive load representation; and 3) embedding a single-phase GFM bus model directly into the three-phase unbalanced Newton power flow equations through a common droop-governed frequency state. GFM-3PF is validated on 3-bus, 33-bus, and 118-bus test systems, demonstrating its accuracy, robustness, and scalability.

cs.ET

QSVT-Based Three-Phase Unbalanced Power Flow

This letter introduces QSVT-3PF, a quantum singular value transformation (QSVT) based solver for three-phase unbalanced power flow with embedded single-phase grid-forming inverter (GFM) operation. The contributions are threefold: 1) reformulating the Newton correction as a QSVT-compatible inverse problem using a normalized block-encoded phase-domain Jacobian; 2) introducing a regularized singular-value filter to improve robustness under ill-conditioned and stressed operating conditions; and 3) validating the proposed solver on an IEEE 5-bus system and the IEEE 123-node test feeder with single-phase GFM integration. Test results show that QSVT-3PF matches the classical Newton benchmark in residual convergence, voltage profile, and final operating point, demonstrating the feasibility of QSVT-3 for large, unbalanced distribution systems.

cs.ET

Stimulated Oscillations in Renewable Energy Integrated Power Systems - Part II: Methodology of Oscillation Mitigation

As presented in Part I of this series, closely located poles can produce high-amplitude oscillations even under small perturba-tions, which are referred to as stimulated oscillations. As the second installment of this series, this paper develops a method-ology for stimulated oscillation mitigation. Firstly, the logical relationship between system stability and oscillation risk is in-vestigated, clarifying that stability is neither a necessary nor a sufficient condition for oscillation risk. Secondly, the dominant factors affecting the pole and zero position on the complex plane, as well as their effectiveness and limitations are analyzed. For feedback control systems in particular, the influence of feedback paths on the pole-zero distribution of closed-loop systems is ana-lyzed. On this basis, a methodology for stimulated oscillation mitigation is proposed. Combined with the reduced-order trans-fer function of systems with closely separated poles, the order relation and configuration of poles and zeros for the feedback path transfer function, together with the criteria for gain selec-tion, are elaborated. Finally, a parameter tuning scheme for the poles, zeros and gain of the feedback path transfer function is presented. The proposed methodology for stimulated oscillation mitigation based on feedback control is not restricted to specific devices and can serve as a methodological reference for the design of diverse oscillation suppression schemes.

eess.SY

DTD-VAE: Disentangled Temporal Dependencies VAE for Credit Risk Prediction

Evaluating customer creditworthiness is crucial for retail banking operations, as it impacts marketing strategies, customer relationship management, and credit risk control. Traditional methods often struggle to capture complex temporal dependencies and extract pertinent information from customer data, crucial for accurate risk assessment. Specifically, they fail to differentiate between temporal patterns indicative of credit risk and those reflecting general customer behavior or preferences, leading to suboptimal risk predictions. In this study, we introduce the Disentangled Temporal Dependencies Variational Autoencoder (DTD-VAE), an advancement over conventional VAE, designed to disentangle temporal dependencies and distinguish credit risk-related features from past customer preferences. The feature inference module of the DTD-VAE incorporates an autoregressive temporal dependency learning mechanism that adeptly captures the temporal dependencies among latent variables, enriching the model's comprehension of the inherent data structure. Furthermore, the feature generative module utilizes an element-wise gating mechanism that assigns independent weights to each dimension of the expert models, enabling a finer-grained disentanglement of latent variables, particularly those relevant to credit risk prediction. Extensive experiments on six real-world datasets demonstrate that the proposed framework consistently outperforms existing methods, achieving performance gains of 3.2%-4.86% in ROC-AUC and 6.41%-9.71% in Accuracy Ratio.

q-fin.RM

Tether the Subject, Release the Scene: Query-Aware Memory Routing for Long-Horizon Autoregressive Video Generation

Streaming autoregressive video models generate long videos chunk by chunk, using historical memory to maintain consistency. Existing methods typically expose subject and scene queries to history through similar policies. This stabilizes the subject, but can also lock backgrounds, viewpoints, and scene structure to previously generated states even when local motion continues. We call this failure memory-anchored scene under-progression; consistency and motion metrics alone can miss it. We introduce TetherMem, a training-free, query-aware spatiotemporal memory router for frozen video generators. TetherMem separates subject and scene queries and modulates historical access with region- and age-conditioned priors: subject queries retain identity-bearing history, while scene queries reduce reliance on subject history and stale backgrounds. Across 2,400 blinded pairwise judgments from 10 annotators, TetherMem achieves the highest estimated expected preference among eight streaming long-video baselines for overall quality (0.780) and scene progression (0.769). On complete 30-second videos, it sustains changes in background, viewpoint, and scene state while preserving subject recognizability and temporal continuity.

cs.CV

Visually-Guided Spatial Audio Generation for $360^\circ$ In-the-Wild Speech Scenes

Spatial audio is a key component of immersive $360^\circ$ media, yet high-quality spatial capture remains limited in real-world speech-dominant scenes. We study visually guided First-Order Ambisonics (FOA) speech spatialization in the wild: given aligned $360^\circ$ video and an omnidirectional audio track, we recover the missing directional FOA components. To support this task, we introduce YT-SPEECH, a speech-oriented $360^\circ$ video-FOA dataset curated from YouTube. We propose a two-stage Localizer-Renderer framework, where an audio-visual segmentation backbone provides frame-wise spatial heatmaps and a conditional complex-domain U-Net reconstructs directional FOA signals from the omnidirectional channel. A confidence-based gating strategy stabilizes conditioning under ambiguous acoustic conditions. Experiments show improved reconstruction fidelity, spatial accuracy, and perceptual speech quality relative to ablated variants and prior approaches.

eess.AS

Code-Domain Grouped Index Modulation for Spectrally Efficient Spread-Spectrum Communications

Efficient exploitation of spreading resources is important for improving the transmission efficiency of spread-spectrum communications. Code index modulation (CIM) conveys additional information through spreading-code indices without requiring extra radio-frequency chains. The fixed number of code-index branches in existing CIM schemes limits information allocation between the code-index and modulation-symbol domains as the target rate increases. This paper proposes code-domain grouped index modulation (CGIM), which partitions an orthogonal spreading-code bank into multiple groups and maps the in-phase and quadrature components of one quadrature amplitude modulation (QAM) symbol in each group onto two independent code indices. Group-wise despreading enables parallel detection. Group-wise maximum-likelihood (ML) and low-complexity greedy detection (GD) are developed, with both achieving joint ML decisions under the stated conditions. A pairwise error probability (PEP)-based bit error rate (BER) analysis is derived for Rayleigh and Nakagami-$m$ fading, with additive white Gaussian noise (AWGN) as a benchmark. Numerical results show that CGIM achieves better BER performance than traditional CIM benchmark schemes at the same transmission rate, owing to a more balanced distribution of information bits between the code-index and modulation-symbol domains.

cs.IT

InfinityEdit: Infinite Video Editing with a Lightweight Edit-Ignition Adapter

With large pretrained models, existing methods have effectively improved instruction-based video editing. However, most of them rely on an in-place editing assumption. They align the edited video with the given source clip frame by frame over a fixed time span. This pattern fails for open-ended streams, e.g., restyling a live game or applying a camera move to an ongoing shot. In such cases, edits must extend to future frames as they arrive, rather than be applied to a static input clip. In this paper, we study this setting and name it infinite video editing: given a preceding segment and an edit request, a model must generate the next segment that continues the stream while applying the requested edit. This process repeats as an unbounded sequence of edit instructions arrives. This task brings two challenges: the edit must be a faithful continuation rather than a frame-wise rewrite, and generation quality must remain stable as edits accumulate. To address them, we first design a data-collection pipeline for infinite video editing. Based on the collected data, we propose InfinityEdit, a lightweight edit adapter that equips a streaming video generator with unbounded editing ability. The adapter contains three attention modules. History cross-attention guides the denoising frames using the input frames. Temporal causal self-attention keeps temporal cues flowing only from earlier frames to later ones. Edit cross-attention injects the edit request into generation. During inference, the adapter is activated only in the chunk where an edit request arrives. Subsequent chunks are generated by the original model with a reset anchor frame. This scheme applies the edit while preserving the original model's infinite generation ability. Extensive experiments show that InfinityEdit faithfully continues the stream under each edit, and stays stable over unbounded edit sequences.

cs.CV

GEAR: Generative Expansion and Real Anchoring for Two-Stage Distillation of Tabular Foundation Models

Tabular foundation models (TFMs) achieve strong performance through in-context learning, but context-dependent inference imposes substantial latency and memory costs, hindering large-scale deployment. We propose GEAR (\emph{Generative Expansion and Real Anchoring}), a modular two-stage framework that distills TFMs into lightweight MLP or tree-based predictors that can be deployed on commodity CPUs. Stage 1 uses synthetic covariates solely as teacher-query locations and trains the student on soft TFM targets, expanding coverage beyond observed rows. Stage 2 re-anchors the student to the target distribution using real labels and out-of-fold teacher predictions, whitch avoids self-labeling leakage. We further derive a risk certificate characterizing the trade-off between generated-query volume and generator fidelity. Experiments on TALENT and TabArena demonstrate the broad applicability of GEAR. Two-stage MLPs outperform supervised MLPs by 1.81--2.00 AUC points on binary tasks and 1.19--1.35 points on multiclass tasks, with additional gains over real-data-only distillation of 1.76--2.19 and 2.09--2.40 points, respectively. On binary tasks, the gains also transfer to LightGBM and XGBoost, and all three student families outperform CatBoost, the strongest non-TFM baseline, in mean AUC. Ablations show gains beyond longer training or alternative warm starts, greater stability from staged than mixed optimization, and generator-dependent diminishing returns as query volume increases. Finally, GEAR reduces median inference time by 57--2866 times and peak prediction memory by 1.9--3.3 times, while retaining higher AUC than matched supervised baselines.

cs.LG

The Snake Algorithm: A Rejection-Free Sampler for Binary Matrices with Fixed Margins

We study uniform sampling of binary matrices with fixed row and column sums, a recurring problem in ecological null models, Rasch-model testing, network analysis, and combinatorics. We propose the Snake algorithm, a rejection-free Markov chain Monte Carlo sampler that grows an alternating path until its first self-intersection and flips the resulting loop. The chain is reversible and irreducible on the fixed-margin state space, hence has the uniform stationary distribution. We prove that one step flips on the order of $\sqrt{n}$ entries in sparse and balanced square regimes, give upper bounds on the per-step path length, and show that the resulting work per flipped entry is rate optimal in sparse and balanced regimes and near-optimal up to a polylogarithmic factor under a one-sided half-balanced condition. A Markov-chain comparison, combined with the recently established universal spectral-gap bound for the swap chain, proves that the lazy Snake chain is rapidly mixing for every feasible pair of margins; in the permutation-matrix case, the raw chain has the sharp total-variation mixing time $\Theta(n \log n)$. We also describe a directed-graph extension and an equal-margin label-shuffling variant. Numerical experiments against Swap, Rectangle Loop, Curveball, sequential importance sampling, and a directed edge-swap algorithm show consistent gains in move size, wall-clock convergence, and sampling efficiency.

stat.CO

Evidence of self-organized criticality in the prompt emission of a bright gamma-ray burst

Gamma-ray bursts (GRBs) are the most energetic explosive events in the Universe, yet the physical mechanism of their prompt emission remains a mystery. Especially, it is unclear whether the energy dissipation mechanism in the GRB jet is dominated by kinetic energy or magnetic energy. Here, we studied the pulses in the prompt emission of the second brightest GRB to date, GRB 230307A, which was accurately measured by the Gravitational wave high-energy electromagnetic counterpart all-sky monitor (GECAM), with focus on the cumulative distributions of peak counts and duration of pulses as well as the waiting time between pulses. We find that these cumulative distributions show scale-invariant behavior, well consistent with the prediction of the self-organized criticality (SOC) theory. This is the first robust evidence of an SOC feature in the prompt emission of a single GRB. Moreover, the statistical properties of pulses in the prompt emission of GRB 230307A are very similar to those of solar flares. Our findings suggest that the prompt emission of GRB is powered by the dissipation of magnetic energy in the ultra-relativistic jet, supporting the Poynting-flux-dominated prompt models.

astro-ph.HE

FeatureHospital: A Skill-Driven Multi-Agent Framework for Automated Algorithm Customization in Multi-View Multi-Label Feature Selection

Multi-view multi-label feature selection aims to identify a compact and informative feature subset from heterogeneous views while preserving discriminative information for multiple labels. Existing methods are generally developed from specific modeling perspectives and incorporate mechanisms tailored to particular data characteristics. Designing suitable feature selection algorithms across datasets with diverse and heterogeneous characteristics still relies heavily on expert knowledge and substantial manual effort, imposing considerable time and labor costs that severely hinder the practical adoption of feature selection. To address this problem, we propose FeatureHospital, a Skill-driven multi-agent framework for automated multi-view multi-label feature selection algorithm design. FeatureHospital first diagnoses the target dataset to identify its feature selection issues. Based on the diagnosis, specialist agents equipped with domain Skills then prescribe corresponding optimization strategies and Loss terms for different issues. After that, the resulting prescriptions are reconciled to remove overlaps and resolve conflicts before being integrated into a compact dataset-specific objective. Finally, the constructed objective is optimized to select the final feature subset. Experimental results demonstrate that FeatureHospital can construct effective feature selection algorithms for different datasets based on their individual characteristics.

cs.AI

Stimulated Oscillations in Renewable Energy Integrated Power Systems - Part I : Mechanism and Analysis Methods

Oscillation is a critical issue that power systems have long faced. Especially over the past two decades, with the large-scale inte-gration of renewable energy into the grid, oscillation problems have posed a serious threat to the secure operation of power systems. However, the current literature has not fully explained the oscillation mechanism of renewable energy integrated power systems (REIPSs). In this paper, the underlying mechanism of stimulated oscillations is explored, with novel analytical methods proposed. Firstly, it is explained from both mathematical formu-las and physical interpretations that for an oscillation mode characterized by a pair of complex conjugate poles, the oscilla-tion risk under disturbance depends on the relative positional relationship between the corresponding poles and all other poles and zeros on the complex plane, rather than their standalone locations, i.e., the stability perceived by classical theory. Then the underlying mechanism of high amplitude oscillations induced by closely-located poles under even slight disturbance is clarified. On this basis, a theoretical framework for stimulated oscilla-tions applicable to REIPSs, covering its definition, mechanism, and methods, is proposed. Finally, this paper discusses the rela-tionship between the stimulated oscillation theory proposed herein and the classical stability-based theory, revealing that the research findings surpass rather than negate the classical theo-ries.

eess.SY

Inferring 1-Minimal Trigger Configurations for Assessing Linux Kernel CVE Triggerability

Vendors assessing Linux kernel CVEs need to know whether a bug is triggerable under production-tailored configurations, not merely whether a version is affected, yet upstream reproducers and vulnerability databases rarely provide configuration-level context. We study minimal trigger-configuration inference: given a CVE entry and a target kernel version (optionally a baseline .config), we synthesize a Kconfig-satisfiable option set that remains effective after make olddefconfig and, when a reproducer is available, still triggers under a specified evaluation protocol; we then prune it to a 1-minimal (subset-minimal) boundary for evaluation. Our framework FCC links vulnerability cues to build-system symbols, completes implicit prerequisites under olddefconfig feedback to avoid silent rollback, and performs runtime-validated minimization guided by dependency topology. We evaluate on KernJC and KernelCTF, totaling 88 CVEs across multiple kernel versions. On the 88-CVE set, FCC improves the post-make olddefconfig configuration success rate from 62.5% (55/88) to 96.6% (85/88) over an olddef-only injection baseline; on the KernJC set, FCC reduces the average candidate set size by 78.7% compared to KernJC (Avg. 14.72 vs. 69.00 options per CVE). A stage-wise analysis of time and token costs shows that Stage I dominates overhead, while CVE-focused evidence selection substantially reduces this cost. By returning an effective and auditable 1-minimal configuration boundary, FCC helps vendors scope triggerability against their deployment configurations with a clear, tool-supported decision line.

cs.CR