Searcharxiv⌕ Search

arXiv subjects

Mingzhi Wang

Publications and source records attributed to Mingzhi Wang.

At least 19 recordsLinked to original sources

Geodesics and shadows of the spindle-deformed Kerr black hole

Recently, a new exact Ricci-flat rotating black-hole solution was constructed in four-dimensional general relativity, in which an additional parameter $B$ characterizes a spindle deformation of the Kerr geometry. We study geodesic motion and black-hole shadows in this spacetime. The Hamilton-Jacobi equation is not exactly separable for either timelike or null geodesics. Remarkably, however, at the leading nontrivial order, ${\cal O}(B^2)$, null but not timelike geodesics become separable. In the timelike sector, the spindle deformation shifts the innermost stable circular orbit and can give rise to an outermost stable circular orbit. In the null sector, exploiting the perturbatively separated equations, we analytically determine the photon region, equatorial photon orbits, and black-hole shadow, and compare the resulting predictions with direct ray tracing in the exact spacetime. The numerical results validate the perturbative treatment and quantify the deviations from Kerr.

gr-qc↗

Rotation-Parameterized Graph Fractional Fourier Transform: Definition, Properties, and Optimal Filtering

Graph spectral representations are fundamental in graph signal processing, providing a rigorous frameworkforanalyzing graph-structured data. The graph fractional Fourier transform (GFRFT) extends the graph Fourier transform (GFT) through a fractional-order parameter, enabling flexible spectral analysis with mathematical consistency. The angular graph Fourier transform (AGFT) further introduces angular control by rotating GFT eigenvectors; however, existing constructions may fail to reduce exactly to the GFT at zero angle, weakening theoretical consistency and interpretability. To address these complementary limitations, namely the lack of rotation-based basis control in GFRFT and the defective zero-angle degeneracy of AGFT, this paper proposes the rotation-parameterized graph fractional Fourier transform (RP-GFRFT), which unifies fractional order and rotation-parameterized spectral analysis. A degeneracy preserving rotation matrix family is constructed to guarantee exact GFT reduction at zero angle. TwoRP-GFRFTvariants,I-RP-GFRFTandII-RP-GFRFT,arethenformulated, with theoretical analyses confirming their unitarity, invertibility, reduction behavior, and smooth parameter dependence. The fractional order and rotation angle are jointly optimized for adaptive graph spectral filtering. Experiments on real-world signals, images, and point clouds demonstrate that RP-GFRFT improves denoising accuracy, reconstruction quality, and feature preservation over GFRFT, AGFT, and representative filtering baselines.

stat.ML↗

FGFRFT: Fast Graph Fractional Fourier Transform via Exact Spectral Splitting and Fourier-Series Approximation

The graph fractional Fourier transform (GFRFT) for unitary graph Fourier transform (GFT) matrices can be interpreted through the scalar function $e^{jαθ}$ on the unit circle. Under the principal branch, its Fourier-series representation encounters an intrinsic obstruction at the spectral point $λ=-1$ for non-integer orders. To address this issue, we propose a fast graph fractional Fourier transform (FGFRFT) based on exact spectral splitting: the $λ=-1$ component is treated exactly, and the complementary component is approximated by a truncated Fourier series in integer powers of the GFT matrix. This construction yields an offline--online implementation that reduces the online complexity of repeated operator updates from $O(N^3)$ to $O(2LN^2)$ for truncation order $L$, while preserving differentiability with respect to the transform order. We further derive truncation-error bounds, approximate unitarity and additivity, and reconstruction-error bounds. Experiments on approximation accuracy, transform-order learning, image denoising, and point-cloud denoising show that FGFRFT provides substantial online acceleration while remaining close to the exact GFRFT under the tested settings.

eess.SP↗

TopoClaw: A Human-Centric and Topology-Aware Agent Operating System

Large language models (LLMs) have evolved AI assistants into autonomous reasoning engines that maintain context, invoke tools, and pursue long-horizon tasks. This has spurred Agent Operating Systems (Agent OS) as kernel-like layers for lifecycle management, memory, scheduling, and access control. Yet most designs remain agent-centric, treating the OS as a single-host runtime for internal reasoning and tool use, leaving open how autonomous actions integrate with distributed, collaborative, permission-sensitive workflows. TopoClaw is an open-source, human-centric, topology-aware Agent OS modeling the user's ecosystem as two coupled structures: a physical device topology of heterogeneous surfaces and a social relationship topology of shared spaces, teams, and delegated roles. It unifies device operation, messaging, and skills around accountable cross-boundary execution, with three core contributions: (1) cross-device action placement, decoupling intent from actuation and routing distributed actions across the device cluster based on hardware affordances and user context; (2) cross-user identity attribution, treating agents as socially situated "Digital Twins" that coordinate in multi-user spaces while preserving provenance, role-aware permissions, and human accountability; (3) cross-context authority governance, pairing broad capability with distributed, context-aware policy enforcement across physical and social trust boundaries to bound proactive autonomy at the OS layer. This report presents TopoClaw as an engineering-oriented reference architecture, covering its design principles, runtime, cross-device execution, collaboration mechanisms, security model, and deployment outlook.

cs.HC↗

A Unified Fractional Spectral Framework for Spatiotemporal Graph Signals: Bi-Fractional Transform and Geodesic Coupling

Graph signal processing extends spectral analysis to data supported on irregular domains. Existing fractional transforms for two-dimensional graph signals, including the two-dimensional graph fractional Fourier transform (GFRFT), typically impose a shared fractional order across dimensions, which limits adaptivity to heterogeneous spatiotemporal spectra. To address this limitation, we propose the two-dimensional graph bi-fractional Fourier transform, which assigns independent fractional orders to the factor graphs of a Cartesian product, enabling decoupled spectral control while preserving separability, unitarity, and invertibility. To further resolve the basis ambiguity in temporal fractional analysis, we develop a geodesic-coupled GFRFT by constructing a coupling path along the principal geodesic on the unitary manifold, thereby unifying graph-induced and discrete temporal bases with guaranteed unitarity and a closed-form inverse. Building on these transforms, we derive a differentiable Wiener-type filtering framework with a hybrid optimization strategy: the fractional orders are learned end-to-end from data, while the coupling parameter is fixed as a structural regularizer. Experiments on real-world time-varying graph datasets and dynamic image restoration tasks demonstrate consistent gains over state-of-the-art fractional transforms and competitive learning-based baselines.

eess.SP↗

Fusion-PSRO: Nash Policy Fusion for Policy Space Response Oracles

For solving zero-sum games involving non-transitivity, a useful approach is to maintain a policy population to approximate the Nash Equilibrium (NE). Previous studies have shown that the Policy Space Response Oracles (PSRO) algorithm is an effective framework for solving such games. However, current methods initialize a new policy from scratch or inherit a single historical policy in Best Response (BR), missing the opportunity to leverage past policies to generate a better BR. In this paper, we propose Fusion-PSRO, which employs Nash Policy Fusion to initialize a new policy for BR training. Nash Policy Fusion serves as an implicit guiding policy that starts exploration on the current Meta-NE, thus providing a closer approximation to BR. Moreover, it insightfully captures a weighted moving average of past policies, dynamically adjusting these weights based on the Meta-NE in each iteration. This cumulative process further enhances the policy population. Empirical results on classic benchmarks show that Fusion-PSRO achieves lower exploitability, thereby mitigating the shortcomings of previous research on policy initialization in BR.

cs.GT↗

Dynamic shadow of a black hole with a self-interacting massive complex scalar hair

We investigate the dynamic shadows of a black hole with a self-interacting massive complex scalar hair. The complex scalar field ψevolves with time t, and its magnitude on the apparent horizon |ψ_{h}| starts from zero, undergoes a sharp rise followed by rapid oscillations, and eventually converges to a constant value. The variation in the photon sphere radius r_{ps} is similar to that of the magnitude |ψ_{h}|. Owing to the emergence of the complex scalar hair ψ, the apparent horizon radius r_{h} starts increasing sharply and then smoothly approaches a stable value eventually. The shadow radius R_{sh} of the black hole with an accretion disk increases with time t_{o} at the observer's position. In the absence of an accretion disk, the shadow radius R_{sh} is larger and also increases as t_{o} increases. Furthermore, we slice the dynamical spacetime into spacelike hypersurfaces for all time points t. For the case with an accretion disk, the variation in R_{sh} is similar to that in the apparent horizon r_{h}, because the inner edge of the accretion disk extends to the apparent horizon. In the absence of an accretion disk, the variation in R_{sh} is similar to that in the photon sphere radius r_{ps}, because the black hole shadow boundary is determined by the photon sphere. As the variation in r_{ps} is induced by ψ, it can be stated that the variation in the size of the shadow is similarly caused by the change in ψ. Regardless of the presence or absence of the accretion disk, the emergence of the complex scalar hair ψcauses the radius R_{sh} of the shadow to start changing. Moreover, we investigate the time delay Δt of lights propagating from light sources to the observer. These findings not only enrich the theoretical models of dynamic black hole shadows but also provide a foundation for testing black hole spacetime dynamics.

gr-qc↗

Two-Dimensional Graph Bi-Fractional Fourier Transform

Graph signal processing (GSP) advances spectral analysis on irregular domains. However, existing two-dimensional graph fractional Fourier transform (2D-GFRFT) employs a single fractional order for both factor graphs, thereby limiting its adaptability to heterogeneous signals. We proposed the two-dimensional graph bi-fractional Fourier transform (2D-GBFRFT), which assigns independent fractional orders to the factor graphs of a Cartesian product while preserving separability. We established invertibility, unitarity, and index additivity, and developed two filtering schemes: a Wiener-style design through grid search and a differentiable framework that jointly optimizes transform orders and diagonal spectral filters. We further introduced a hybrid interpolation with the joint time-vertex fractional Fourier transform (JFRFT), controlled by a tunable parameter that balances the two methods. In the domains of synthetic Cartesian product graph signals, authentic temporal graph datasets, and dynamic image deblurring, 2D-GBFRFT consistently surpasses 2D-GFRFT and enhances JFRFT. Experimental results confirmed the versatility and superior performance of 2D-GBFRFT for filtering in GSP.

eess.SP↗

Roadmap on Incentive Compatibility for AI Alignment and Governance in Sociotechnical Systems

The burgeoning integration of artificial intelligence (AI) into human society brings forth significant implications for societal governance and safety. While considerable strides have been made in addressing AI alignment challenges, existing methodologies primarily focus on technical facets, often neglecting the intricate sociotechnical nature of AI systems, which can lead to a misalignment between the development and deployment contexts. To this end, we posit a new problem worth exploring: Incentive Compatibility Sociotechnical Alignment Problem (ICSAP). We hope this can call for more researchers to explore how to leverage the principles of Incentive Compatibility (IC) from game theory to bridge the gap between technical and societal components to maintain AI consensus with human societies in different contexts. We further discuss three classical game problems for achieving IC: mechanism design, contract theory, and Bayesian persuasion, in addressing the perspectives, potentials, and challenges of solving ICSAP, and provide preliminary implementation conceptions.

cs.AI↗

Amulet: ReAlignment During Test Time for Personalized Preference Adaptation of LLMs

How to align large language models (LLMs) with user preferences from a static general dataset has been frequently studied. However, user preferences are usually personalized, changing, and diverse regarding culture, values, or time. This leads to the problem that the actual user preferences often do not coincide with those trained by the model developers in the practical use of LLMs. Since we cannot collect enough data and retrain for every demand, researching efficient real-time preference adaptation methods based on the backbone LLMs during test time is important. To this end, we introduce Amulet, a novel, training-free framework that formulates the decoding process of every token as a separate online learning problem with the guidance of simple user-provided prompts, thus enabling real-time optimization to satisfy users' personalized preferences. To reduce the computational cost brought by this optimization process for each token, we additionally provide a closed-form solution for each iteration step of the optimization process, thereby reducing the computational time cost to a negligible level. The detailed experimental results demonstrate that Amulet can achieve significant performance improvements in rich settings with combinations of different LLMs, datasets, and user preferences, while maintaining acceptable computational efficiency.

cs.CL↗

Falcon: Fast Visuomotor Policies via Partial Denoising

Diffusion policies are widely adopted in complex visuomotor tasks for their ability to capture multimodal action distributions. However, the multiple sampling steps required for action generation significantly harm real-time inference efficiency, which limits their applicability in real-time decision-making scenarios. Existing acceleration techniques either require retraining or degrade performance under low sampling steps. Here we propose Falcon, which mitigates this speed-performance trade-off and achieves further acceleration. The core insight is that visuomotor tasks exhibit sequential dependencies between actions. Falcon leverages this by reusing partially denoised actions from historical information rather than sampling from Gaussian noise at each step. By integrating current observations, Falcon reduces sampling steps while preserving performance. Importantly, Falcon is a training-free algorithm that can be applied as a plug-in to further improve decision efficiency on top of existing acceleration techniques. We validated Falcon in 48 simulated environments and 2 real-world robot experiments. demonstrating a 2-7x speedup with negligible performance degradation, offering a promising direction for efficient visuomotor policy design.

cs.RO↗

Magnetic Preference Optimization: Achieving Last-iterate Convergence for Language Model Alignment

Self-play methods have demonstrated remarkable success in enhancing model capabilities across various domains. In the context of Reinforcement Learning from Human Feedback (RLHF), self-play not only boosts Large Language Model (LLM) performance but also overcomes the limitations of traditional Bradley-Terry (BT) model assumptions by finding the Nash equilibrium (NE) of a preference-based, two-player constant-sum game. However, existing methods either guarantee only average-iterate convergence, incurring high storage and inference costs, or converge to the NE of a regularized game, failing to accurately reflect true human preferences. In this paper, we introduce Magnetic Preference Optimization (MPO), a novel approach capable of achieving last-iterate convergence to the NE of the original game, effectively overcoming the limitations of existing methods. Building upon Magnetic Mirror Descent (MMD), MPO attains a linear convergence rate, making it particularly suitable for fine-tuning LLMs. To ensure our algorithm is both theoretically sound and practically viable, we present a simple yet effective implementation that adapts the theoretical insights to the RLHF setting. Empirical results demonstrate that MPO can significantly enhance the performance of LLMs, highlighting the potential of self-play methods in alignment.

cs.CL↗

Conflux-PSRO: Effectively Leveraging Collective Advantages in Policy Space Response Oracles

Policy Space Response Oracle (PSRO) with policy population construction has been demonstrated as an effective method for approximating Nash Equilibrium (NE) in zero-sum games. Existing studies have attempted to improve diversity in policy space, primarily by incorporating diversity regularization into the Best Response (BR). However, these methods cause the BR to deviate from maximizing rewards, easily resulting in a population that favors diversity over performance, even when diversity is not always necessary. Consequently, exploitability is difficult to reduce until policies are fully explored, especially in complex games. In this paper, we propose Conflux-PSRO, which fully exploits the diversity of the population by adaptively selecting and training policies at state-level. Specifically, Conflux-PSRO identifies useful policies from the existing population and employs a routing policy to select the most appropriate policies at each decision point, while simultaneously training them to enhance their effectiveness. Compared to the single-policy BR of traditional PSRO and its diversity-improved variants, the BR generated by Conflux-PSRO not only leverages the specialized expertise of diverse policies but also synergistically enhances overall performance. Our experiments on various environments demonstrate that Conflux-PSRO significantly improves the utility of BRs and reduces exploitability compared to existing methods.

cs.GT↗

Computing Ex Ante Equilibrium in Heterogeneous Zero-Sum Team Games

The ex ante equilibrium for two-team zero-sum games, where agents within each team collaborate to compete against the opposing team, is known to be the best a team can do for coordination. Many existing works on ex ante equilibrium solutions are aiming to extend the scope of ex ante equilibrium solving to large-scale team games based on Policy Space Response Oracle (PSRO). However, the joint team policy space constructed by the most prominent method, Team PSRO, cannot cover the entire team policy space in heterogeneous team games where teammates play distinct roles. Such insufficient policy expressiveness causes Team PSRO to be trapped into a sub-optimal ex ante equilibrium with significantly higher exploitability and never converges to the global ex ante equilibrium. To find the global ex ante equilibrium without introducing additional computational complexity, we first parameterize heterogeneous policies for teammates, and we prove that optimizing the heterogeneous teammates' policies sequentially can guarantee a monotonic improvement in team rewards. We further propose Heterogeneous-PSRO (H-PSRO), a novel framework for heterogeneous team games, which integrates the sequential correlation mechanism into the PSRO framework and serves as the first PSRO framework for heterogeneous team games. We prove that H-PSRO achieves lower exploitability than Team PSRO in heterogeneous team games. Empirically, H-PSRO achieves convergence in matrix heterogeneous games that are unsolvable by non-heterogeneous baselines. Further experiments reveal that H-PSRO outperforms non-heterogeneous baselines in both heterogeneous team games and homogeneous settings.

cs.GT↗

The ring-shaped shadow of rotating naked singularity with a complete photon sphere

We investigate the shadows of Konoplya-Zhidenko naked singularity. In the spacetime of Konoplya-Zhidenko naked singularity, not only can unstable retrograde light ring (LR) exist, but also unstable prograde LR, leading to the formation of a complete photon sphere (PS). Due to the absence of an event horizon, a dark disc-shaped shadow does not appear; instead, a ring-shaped shadow is observed. The ring-shaped shadow appears as an infinite number of relativistic Einstein rings in the image of the naked singularity. For some parameter values, only the unstable retrograde LR exists, resulting in an incomplete unstable PS and consequently giving rise to the arc-shaped shadow for Konoplya-Zhidenko naked singularity. The shadow of Konoplya-Zhidenko naked singularity gradually shifts to the right as the rotation parameter $a$ increases, and gradually becomes smaller as the deformation parameter $|η|$ increases. Moreover, the stable LRs and stable photon spherical orbits can also exist in Konoplya-Zhidenko naked singularity spacetime, but they have no effect on the image of the naked singularity. This study demonstrates that rotating naked singularity can exhibit not only an arc-shaped shadow but also a ring-shaped shadow.

gr-qc↗

Efficient Model-agnostic Alignment via Bayesian Persuasion

With recent advancements in large language models (LLMs), alignment has emerged as an effective technique for keeping LLMs consensus with human intent. Current methods primarily involve direct training through Supervised Fine-tuning (SFT) or Reinforcement Learning from Human Feedback (RLHF), both of which require substantial computational resources and extensive ground truth data. This paper explores an efficient method for aligning black-box large models using smaller models, introducing a model-agnostic and lightweight Bayesian Persuasion Alignment framework. We formalize this problem as an optimization of the signaling strategy from the small model's perspective. In the persuasion process, the small model (Advisor) observes the information item (i.e., state) and persuades large models (Receiver) to elicit improved responses. The Receiver then generates a response based on the input, the signal from the Advisor, and its updated belief about the information item. Through training using our framework, we demonstrate that the Advisor can significantly enhance the performance of various Receivers across a range of tasks. We theoretically analyze our persuasion framework and provide an upper bound on the Advisor's regret, confirming its effectiveness in learning the optimal signaling strategy. Our Empirical results demonstrates that GPT-2 can significantly improve the performance of various models, achieving an average enhancement of 16.1% in mathematical reasoning ability and 13.7% in code generation. We hope our work can provide an initial step toward rethinking the alignment framework from the Bayesian Persuasion perspective.

cs.CL↗

Leveraging Team Correlation for Approximating Equilibrium in Two-Team Zero-Sum Games

Two-team zero-sum games are one of the most important paradigms in game theory. In this paper, we focus on finding an unexploitable equilibrium in large team games. An unexploitable equilibrium is a worst-case policy, where members in the opponent team cannot increase their team reward by taking any policy, e.g., cooperatively changing to other joint policies. As an optimal unexploitable equilibrium in two-team zero-sum games, correlated-team maxmin equilibrium remains unexploitable even in the worst case where players in the opponent team can achieve arbitrary cooperation through a joint team policy. However, finding such an equilibrium in large games is challenging due to the impracticality of evaluating the exponentially large number of joint policies. To solve this problem, we first introduce a general solution concept called restricted correlated-team maxmin equilibrium, which solves the problem of being impossible to evaluate all joint policy by a sample factor while avoiding an exploitation problem under the incomplete joint policy evaluation. We then develop an efficient sequential correlation mechanism, and based on which we propose an algorithm for approximating the unexploitable equilibrium in large games. We show that our approach achieves lower exploitability than the state-of-the-art baseline when encountering opponent teams with different exploitation ability in large team games including Google Research Football.

cs.GT↗

Strain-Tunable Magnetic Compensation Temperature of Epitaxial Tb$_3$Fe$_5$O$_{12}$ Thin Films

High-quality rare-earth iron garnet (ReIG) Tb$_3$Fe$_5$O$_{12}$ (TbIG) thin films are epitaxially grown on a series of (111)-oriented garnet substrates with various lattice constants. The coherent growth induces a substrate-dependent in-plane tensile or compressive strain in the TbIG film. Measurements of the anomalous Hall-like effect (AHLE) in TbIG/Pt heterostructures show that the compensation temperature of TbIG films monotonically changes with the film strain. The strain results in a variation of the distances between magnetic atoms in the TbIG crystal and therefore the corresponding exchange interactions. The latter is explicitly calculated as a function of the lattice strain based on density functional theory, reproducing the observed experimental results. This work provides a versatile way to optimize ReIG-based spin-orbit torque devices.

cond-mat.mtrl-sci↗