SearcharxivSearch

arXiv subjects

Yutian Wang

Publications and source records attributed to Yutian Wang.

At least 19 recordsLinked to original sources

Revisiting Sustainability by Design in AI Protocol Governance: An Empirical Review of Comparative DAO and Corporate-Led Standards for the SDGs

As artificial intelligence (AI) agents enter production infrastructure, interoperability protocols shape its governance and sustainability. This paper revisits our comparative study of two AI-agent interoperability standards, Ethereum Request for Comments 8004 (ERC-8004), governed by a decentralized autonomous organization (DAO), and Google's Agent2Agent (A2A), governed by a corporate consortium, through a Sustainability by Design (SbD) lens. Using an LLM-powered pipeline combining automated annotation, neural topic modeling, and multi-layer network analysis, we identify contrasting governance and innovation architectures. ERC-8004 relies on permissionless participation, rough consensus, and decoupled deployment, while A2A assigns binding authority to an eight-seat Technology Steering Committee. The DAO concentrates on constitutive questions of trust and security, including what to build and why, whereas the consortium distributes attention across executive engineering questions of how to implement, document, and deliver the protocol. Both show high participation inequality, while corporate contributors span roughly twice as many themes as DAO contributors. We ask how these architectures produce distinct SDG-relevant signatures and what design principles they suggest for sustainable AI governance. We interpret institutional, discursive, and network patterns through SDGs 8, 9, 10, 11, 12, 16, and 17, identifying capacities for transparency, participation, contestability, and cross-protocol coordination. We argue that sustainable AI infrastructure requires a corrective feedback loop between designed charters and governance in practice, advancing SDG 16 on strong institutions. By integrating computational evidence, organizational research, and sustainable development, this review derives actionable design principles for sustainable AI governance.

cs.CY

Dynamic mean-variance portfolio selection with no-shorting constraints and unknown investment opportunity sets

We study continuous-time mean-variance portfolio selection with no-shorting constraints and unknown investment opportunity sets from a reinforcement learning (RL) perspective. The problem is a constrained stochastic linear -- quadratic control problem for which the entropy-regularized exploratory formulation of Wang et al. (2020) leads to difficulty in theoretical analysis, because enforcing the constraint on the support of randomized policies nullifies the tractable Gaussian exploration. To tackle this challenge, we introduce an auxiliary exploratory problem without entropy in which exploratory policies are still Gaussian whose samples may violate the no-shorting requirement but their means satisfy it. We then prove that, for a suitable choice of exploration variance, the mean of the optimal Gaussian policy of the auxiliary problem coincides with the optimal policy of the original problem. Motivated by this theoretical result, we develop a model-free RL algorithm that learns the optimal policy of the auxiliary (and hence the original) problem directly from trajectory data without estimating the investment opportunity set. A numerical example demonstrates the performance of the proposed algorithm.

math.OC

An Optimization-Based Framework for Solving Forward-Backward Stochastic Differential Equations: Convergence Analysis and Error Bounds

Forward-backward stochastic differential equations have recently become a key focus in the computational field, and their role in continuous-time stochastic optimal control and reinforcement learning has grown increasingly prominent. In this paper, we develop an optimization-based framework for solving coupled forward-backward stochastic differential equations, which naturally arise in stochastic optimal control through the stochastic maximum principle and related Hamiltonian systems. We introduce an integral-form objective function and prove its equivalence to the error between consecutive Picard iterates. Our convergence analysis establishes that minimizing this objective generates sequences that converge to the true solution. We provide explicit upper and lower bounds that relate the objective value to the error between trial and exact solutions. We validate the proposed objective and its theoretical interpretation using two analytical test cases, and further illustrate its numerical applicability on a nonlinear stochastic optimal control problem with up to 1000 dimensions.

math.OC

Agentic Analysis for Agentic Infrastructure: An LLM-Powered Pipeline for Comparative Governance of DAO and Corporate AI Protocols

As AI agent protocols proliferate, the governance structures shaping their interoperability standards remain empirically underexamined. We introduce an LLM-powered comparative pipeline for large-scale governance discourse analysis, integrating automated annotation, neural topic modeling, and multi-layer network analysis to study socio-technical power structures at scale. We validate it on two contrasting standards for agent interoperability: ERC-8004 (permissionless, on-chain) and Google A2A (corporate-led). Analyzing 4,323 governance participation records, we combine LLM-assisted coding, topic modeling, and multi-layer network analysis to examine how institutional design shapes thematic priorities and community structure. We find that while governance form influences substantive focus, both regimes exhibit comparable levels of participation inequality and community fragmentation. Discourse alignment is denser in the permissionless setting, suggesting that open governance may foster greater thematic convergence despite decentralized participation. These findings illustrate how LLM-assisted methods can advance the empirical study of technology governance, with implications for designing more equitable agentic AI standards. All data and code are openly available.

cs.AI

Anatomy of Spin Wave Polarization in Ferromagnets

Spin waves in ferromagnetic materials are predominantly characterized by right-handed circular polarization due to symmetry breaking induced by net magnetization. However, magnetic interactions, including the external magnetic field, Heisenberg exchange, Dzyaloshinskii-Moriya interaction, and dipole-dipole interaction, can modify this behavior, leading to elliptical polarization. This study provides a systematic analysis of these interactions and their influence on spin wave polarization, establishing principles to predict traits such as polarization degree and orientation based on equilibrium magnetization textures. The framework is applied to diverse magnetic configurations, including spin spirals, domain walls, and Skyrmions, offering a comprehensive yet simple approach to understanding polarization dynamics in ferromagnetic systems.

cond-mat.mes-hall

The Connection between Spin Wave Polarization and Dissipation

This study establishes a fundamental connection between the dissipation and polarization of spin waves, which are often treated as independent phenomena. Through theoretical analysis and numerical validation, we demonstrate that within the linearized spin wave regime, a spin wave mode's dissipation rate, defined as the ratio of linewidth to the resonance frequency, exceeds Gilbert damping by a factor given by its spatially averaged polarization. This average is governed by a non-positive definite weight, whose magnitude depends on the magnon density of the local excitation, while its sign is dictated by the local polarization handedness. Remarkably, this universal connection applies across diverse magnetic interactions and textures, offering crucial insights into spin wave dynamics and dissipation.

cond-mat.mes-hall

MeloTrans: A Text to Symbolic Music Generation Model Following Human Composition Habit

At present, neural network models show powerful sequence prediction ability and are used in many automatic composition models. In comparison, the way humans compose music is very different from it. Composers usually start by creating musical motifs and then develop them into music through a series of rules. This process ensures that the music has a specific structure and changing pattern. However, it is difficult for neural network models to learn these composition rules from training data, which results in a lack of musicality and diversity in the generated music. This paper posits that integrating the learning capabilities of neural networks with human-derived knowledge may lead to better results. To archive this, we develop the POP909$\_$M dataset, the first to include labels for musical motifs and their variants, providing a basis for mimicking human compositional habits. Building on this, we propose MeloTrans, a text-to-music composition model that employs principles of motif development rules. Our experiments demonstrate that MeloTrans excels beyond existing music generation models and even surpasses Large Language Models (LLMs) like ChatGPT-4. This highlights the importance of merging human insights with neural network capabilities to achieve superior symbolic music generation.

cs.SD

Mimicking the Mavens: Agent-based Opinion Synthesis and Emotion Prediction for Social Media Influencers

Predicting influencers' views and public sentiment on social media is crucial for anticipating societal trends and guiding strategic responses. This study introduces a novel computational framework to predict opinion leaders' perspectives and the emotive reactions of the populace, addressing the inherent challenges posed by the unstructured, context-sensitive, and heterogeneous nature of online communication. Our research introduces an innovative module that starts with the automatic 5W1H (Where, Who, When, What, Why, and How) questions formulation engine, tailored to emerging news stories and trending topics. We then build a total of 60 anonymous opinion leader agents in six domains and realize the views generation based on an enhanced large language model (LLM) coupled with retrieval-augmented generation (RAG). Subsequently, we synthesize the potential views of opinion leaders and predicted the emotional responses to different events. The efficacy of our automated 5W1H module is corroborated by an average GPT-4 score of 8.83/10, indicative of high fidelity. The influencer agents exhibit a consistent performance, achieving an average GPT-4 rating of 6.85/10 across evaluative metrics. Utilizing the 'Russia-Ukraine War' as a case study, our methodology accurately foresees key influencers' perspectives and aligns emotional predictions with real-world sentiment trends in various domains.

cs.AI

Mechanism for the Broadened Linewidth in Antiferromagnetic Resonance

The linewidth of antiferromagnetic resonance (AFMR) is found to be significantly broader than that of ferromagnetic resonance (FMR), even when the intrinsic Gilbert damping parameter is the same for both systems. We investigate the origin of this enhanced damping rate in AFMR by studying a bipartite magnet model. Through analytical calculations and numerical simulations, we present three perspectives on understanding this linewidth broadening in AFMR: i) The non-dissipative Heisenberg exchange interaction develops a damping-like component in the presence of Gilbert damping, ii) The transverse component of the exchange coupling reduces the AFMR frequency, thereby increasing the damping rate, and iii) The antiferromagnetic eigenmode exhibits characteristics of a two-mode squeezed state, which is inherently linked to an enhanced damping rate. Our findings provide a comprehensive understanding of the complex dynamics governing magnetic dissipation in antiferromagnet and offer insights into the experimentally observed broadened linewidths in AFMR spectra.

cond-mat.mes-hall

Solving Coupled Nonlinear Forward-backward Stochastic Differential Equations: An Optimization Perspective with Backward Measurability Loss

This paper aims to extend the BML method proposed in Wang et al. [22] to make it applicable to more general coupled nonlinear FBSDEs. We interpret BML from the fixed-point iteration perspective and show that optimizing BML is equivalent to minimizing the distance between two consecutive trial solutions in a fixed-point iteration. Thus, this paper provides a theoretical foundation for an optimization-based approach to solving FBSDEs. We also empirically evaluate the method through four numerical experiments.

math.OC

An Efficient Temporary Deepfake Location Approach Based Embeddings for Partially Spoofed Audio Detection

Partially spoofed audio detection is a challenging task, lying in the need to accurately locate the authenticity of audio at the frame level. To address this issue, we propose a fine-grained partially spoofed audio detection method, namely Temporal Deepfake Location (TDL), which can effectively capture information of both features and locations. Specifically, our approach involves two novel parts: embedding similarity module and temporal convolution operation. To enhance the identification between the real and fake features, the embedding similarity module is designed to generate an embedding space that can separate the real frames from fake frames. To effectively concentrate on the position information, temporal convolution operation is proposed to calculate the frame-specific similarities among neighboring frames, and dynamically select informative neighbors to convolution. Extensive experiments show that our method outperform baseline models in ASVspoof2019 Partial Spoof dataset and demonstrate superior performance even in the crossdataset scenario.

cs.SD

Clifford Algebra-Based Iterated Extended Kalman Filter with Application to Low-Cost INS/GNSS Navigation

The traditional GNSS-aided inertial navigation system (INS) usually exploits the extended Kalman filter (EKF) for state estimation, and the initial attitude accuracy is key to the filtering performance. To spare the reliance on the initial attitude, this work generalizes the previously proposed trident quaternion within the framework of Clifford algebra to represent the extended pose, IMU biases and lever arms on the Lie group. Consequently, a quasi-group-affine system is established for the low-cost INS/GNSS integrated navigation system, and the right-error Clifford algebra-based EKF (Clifford-RQEKF) is accordingly developed. The iterated filtering approach is further applied to significantly improve the performances of the Clifford-RQEKF and the previously proposed trident quaternion-based EKFs. Numerical simulations and experiments show that all iterated filtering approaches fulfill the fast and global convergence without the prior attitude information, whereas the iterated Clifford-RQEKF performs much better than the others under especially large IMU biases.

eess.SY

Probabilistic Framework of Howard's Policy Iteration: BML Evaluation and Robust Convergence Analysis

This paper aims to build a probabilistic framework for Howard's policy iteration algorithm using the language of forward-backward stochastic differential equations (FBSDEs). As opposed to conventional formulations based on partial differential equations, our FBSDE-based formulation can be easily implemented by optimizing criteria over sample data, and is therefore less sensitive to the state dimension. In particular, both on-policy and off-policy evaluation methods are discussed by constructing different FBSDEs. The backward-measurability-loss (BML) criterion is then proposed for solving these equations. By choosing specific weight functions in the proposed criterion, we can recover the popular Deep BSDE method or the martingale approach for BSDEs. The convergence results are established under both ideal and practical conditions, depending on whether the optimization criteria are decreased to zero. In the ideal case, we prove that the policy sequences produced by proposed FBSDE-based algorithms and the standard policy iteration have the same performance, and thus have the same convergence rate. In the practical case, the proposed algorithm is still proved to converge robustly under mild assumptions on optimization errors.

math.OC

DBT-Net: Dual-branch federative magnitude and phase estimation with attention-in-attention transformer for monaural speech enhancement

The decoupling-style concept begins to ignite in the speech enhancement area, which decouples the original complex spectrum estimation task into multiple easier sub-tasks i.e., magnitude-only recovery and the residual complex spectrum estimation)}, resulting in better performance and easier interpretability. In this paper, we propose a dual-branch federative magnitude and phase estimation framework, dubbed DBT-Net, for monaural speech enhancement, aiming at recovering the coarse- and fine-grained regions of the overall spectrum in parallel. From the complementary perspective, the magnitude estimation branch is designed to filter out dominant noise components in the magnitude domain, while the complex spectrum purification branch is elaborately designed to inpaint the missing spectral details and implicitly estimate the phase information in the complex-valued spectral domain. To facilitate the information flow between each branch, interaction modules are introduced to leverage features learned from one branch, so as to suppress the undesired parts and recover the missing components of the other branch. Instead of adopting the conventional RNNs and temporal convolutional networks for sequence modeling, we employ a novel attention-in-attention transformer-based network within each branch for better feature learning. More specially, it is composed of several adaptive spectro-temporal attention transformer-based modules and an adaptive hierarchical attention module, aiming to capture long-term time-frequency dependencies and further aggregate intermediate hierarchical contextual information. Comprehensive evaluations on the WSJ0-SI84 + DNS-Challenge and VoiceBank + DEMAND dataset demonstrate that the proposed approach consistently outperforms previous advanced systems and yields state-of-the-art performance in terms of speech quality and intelligibility.

cs.SD

Optimizing Shoulder to Shoulder: A Coordinated Sub-Band Fusion Model for Real-Time Full-Band Speech Enhancement

Due to the high computational complexity to model more frequency bands, it is still intractable to conduct real-time full-band speech enhancement based on deep neural networks. Recent studies typically utilize the compressed perceptually motivated features with relatively low frequency resolution to filter the full-band spectrum by one-stage networks, leading to limited speech quality improvements. In this paper, we propose a coordinated sub-band fusion network for full-band speech enhancement, which aims to recover the low- (0-8 kHz), middle- (8-16 kHz), and high-band (16-24 kHz) in a step-wise manner. Specifically, a dual-stream network is first pretrained to recover the low-band complex spectrum, and another two sub-networks are designed as the middle- and high-band noise suppressors in the magnitude-only domain. To fully capitalize on the information intercommunication, we employ a sub-band interaction module to provide external knowledge guidance across different frequency bands. Extensive experiments show that the proposed method yields consistent performance advantages over state-of-the-art full-band baselines.

cs.SD

Atomic-scale Deformation Process of Glasses Unveiled by Stress-induced Structural Anisotropy

Experimentally resolving atomic-scale structural changes of a deformed glass remains challenging owing to the disordered nature of glass structure. Here, we show that the structural anisotropy emerges as a general hallmark for different types of glasses (metallic glasses, oxide glass, amorphous selenium, and polymer glass) after thermo-mechanical deformation, and it is highly correlates with local nonaffine atomic displacements detected by the high-energy X-ray diffraction technique. By analyzing the anisotropic pair density function, we unveil the atomic-level mechanism responsible for the plastic flow, which notably differs between metallic glasses and covalent glasses. The structural rearrangements in metallic glasses are mediated through cutting and formation of atomic bonds, which occurs in some localized inelastic regions embedded in elastic matrix, whereas that of covalent glasses is mediated through the rotation of atomic bonds or chains without bond length change, which occurs in a less localized manner.

cond-mat.mtrl-sci

Unsupervised Quantized Prosody Representation for Controllable Speech Synthesis

In this paper, we propose a novel prosody disentangle method for prosodic Text-to-Speech (TTS) model, which introduces the vector quantization (VQ) method to the auxiliary prosody encoder to obtain the decomposed prosody representations in an unsupervised manner. Rely on its advantages, the speaking styles, such as pitch, speaking velocity, local pitch variance, etc., are decomposed automatically into the latent quantize vectors. We also investigate the internal mechanism of VQ disentangle process by means of a latent variables counter and find that higher value dimensions usually represent prosody information. Experiments show that our model can control the speaking styles of synthesis results by directly manipulating the latent variables. The objective and subjective evaluations illustrated that our model outperforms the popular models.

eess.AS

Joint magnitude estimation and phase recovery using Cycle-in-Cycle GAN for non-parallel speech enhancement

For the lack of adequate paired noisy-clean speech corpus in many real scenarios, non-parallel training is a promising task for DNN-based speech enhancement methods. However, because of the severe mismatch between input and target speeches, many previous studies only focus on the magnitude spectrum estimation and remain the phase unaltered, resulting in the degraded speech quality under low signal-to-noise ratio conditions. To tackle this problem, we decouple the difficult target w.r.t. original spectrum optimization into spectral magnitude and phase, and a novel Cycle-in-Cycle generative adversarial network (dubbed CinCGAN) is proposed to jointly estimate the spectral magnitude and phase information stage by stage under unpaired data. In the first stage, we pretrain a magnitude CycleGAN to coarsely estimate the spectral magnitude of clean speech. In the second stage, we incorporate the pretrained CycleGAN with a complex-valued CycleGAN as a cycle-in-cycle structure to simultaneously recover phase information and refine the overall spectrum. Experimental results demonstrate that the proposed approach significantly outperforms previous baselines under non-parallel training. The evaluation on training the models with standard paired data also shows that CinCGAN achieves remarkable performance especially in reducing background noise and speech distortion.

cs.SD