SearcharxivSearch

arXiv subjects

Siyi Zhou

Publications and source records attributed to Siyi Zhou.

At least 19 recordsLinked to original sources

SUSY meets pseudo-Hermiticity

In this work, we construct the simplest pseudo-Hermitian quantum field theory that is supersymmetric. This is the pseudo-Hermitian Wess-Zumino model in the sense that it contains a pair of symplectic fermions (anti-commuting scalar fields) that satisfy the Klein-Gordon equation and a spin-half boson that satisfies the Dirac equation. The conventional spin-statistics theorem is circumvented through the use of pseudo-Hermitian conjugation to define field adjoints. To make the supersymmetry manifest, we formulate the pseudo-Hermitian Wess-Zumino model using the superfield formalism. These superfields are Grassmann-odd so it is not possible to construct non-vanishing cubic interactions using only these superfields. We show that this problem can be resolved by coupling the pseudo-Hermitian Wess-Zumino model with the Hermitian Wess-Zumino model while preserving supersymmetry.

hep-th

The Traffickers' Pitch: Detecting Deceptive Recruitment in Online Job Boards

While substantial efforts in anti-trafficking research and practice have focused on identifying and assisting victims after exploitation occurs, comparatively less attention has been paid to preventing victimization at the recruitment stage. Although some platforms offer preventive tools, such as background checks triggered by in-person meeting detection, these measures primarily protect potential victims rather than directly limiting traffickers' recruitment activities. In this paper, we propose a computational framework to identify human trafficking recruiters through their linguistic features and to characterize their online recruitment patterns. We introduce a network-driven labeling method to construct large-scale ground truth for trafficking-at-risk job advertisements. Our results reveal significant linguistic differences between safe and risky advertisements and demonstrate that language models and embedding representations behave distinctly across these linguistic spaces. Building on these insights, we propose a multi-model ensemble classifier to improve the detection of trafficking-at-risk job ads. Finally, we analyze the geographic, gender, industry, and contact-method preferences of trafficking recruiters, revealing systematic patterns in recruitment strategies.

cs.CY

StepAudio 2.5 Technical Report

Unified audio-language modeling has emerged as a prominent trend in modern speech systems, promising to bring the reasoning capabilities of large language models to auditory tasks. However, existing unified foundations often struggle to match the depth of specialized systems across automatic speech recognition (ASR), text-to-speech synthesis (TTS), and realtime spoken interaction. Bridging this gap remains an open challenge. This report presents StepAudio 2.5, a unified audio-language foundation model that matches or exceeds specialized systems across all three capabilities. Rather than treating these tasks as architecturally distinct, we operate on the premise that once text and audio share a multimodal representational space, task specialization becomes a matter of operational regimes: data construction, optimization targets, and decoding constraints. Guided by this insight, we advance the post-training paradigm from standard supervised learning to task-tailored Reinforcement Learning from Human Feedback (RLHF), using it as the primary mechanism to define complex optimization targets. We leverage this RLHF-centric alignment, alongside specialized decoding, to shape a shared backbone into three distinct operational modes. Concretely, the ASR branch advances transcription efficiency via verifiable multi-token decoding; the TTS branch achieves controllable, expressive synthesis through preference-based RLHF and context-rich supervision; and the Realtime branch realizes low-latency, persona-consistent dialogue via generative reward modeling within an RLHF framework. On standard benchmarks, StepAudio 2.5 achieves state-of-the-art results across ASR, TTS, and Realtime, demonstrating that a singular audio-language foundation can successfully internalize the distinct deployment objectives of speech understanding, generation, and live interaction.

eess.AS

Who, Why, and How: Disentangling the Effects of Moderation Source, Context, and Language on Post-Removal Behavior

Content moderation is a central mechanism through which platforms attempt to balance user engagement with community governance. Yet existing research has largely treated moderation as a uniform intervention, overlooking how moderator source, violation context, and linguistic style jointly shape user behavior. Drawing on the Human--AI Interaction Theory of Interactive Media Effects (HAII-TIME), this study examines how these three dimensions produce divergent post-moderation behavioral trajectories in a large-scale observational dataset of 11,795,036 moderation events across 9,285,410 users and 61,261 subreddits on Reddit (2021--2025). Using probabilistic behavioral classification, ANOVA, and OLS regression with PCA-derived linguistic features, we find that bot moderation consistently produces higher compliance and lower self-censorship than human or modteam moderation, challenging the assumption that human agency cues are inherently advantageous. Modteam moderation produces the strongest self-censorship effects, suggesting that institutional depersonalization is a meaningful driver of behavioral withdrawal. Violation severity emerges as a critical contingency: linguistic strategies effective in routine contexts -- elaborated explanation, community-scale appeals, direct personal address -- can backfire for serious violations, whereas prosocially framed and emotionally emphatic messages become most effective when stakes are highest. Of 480 linguistic interactions tested, 33 survive FDR correction. These findings extend HAII-TIME by introducing violation salience as a moderator of cue-based processing, and offer empirical grounding for context-adaptive moderation design.

cs.CY

IndexTTS 2.5 Technical Report

In prior work, we introduced IndexTTS 2, a zero-shot neural text-to-speech foundation model comprising two core components: a transformer-based Text-to-Semantic (T2S) module and a non-autoregressive Semantic-to-Mel (S2M) module, which together enable faithful emotion replication and establish the first autoregressive duration-controllable generative paradigm. Building upon this, we present IndexTTS 2.5, which significantly enhances multilingual coverage, inference speed, and overall synthesis quality through four key improvements: 1) Semantic Codec Compression: we reduce the semantic codec frame rate from 50 Hz to 25 Hz, halving sequence length and substantially lowering both training and inference costs; 2) Architectural Upgrade: we replace the U-DiT-based backbone of the S2M module with a more efficient Zipformer-based modeling architecture, achieving notable parameter reduction and faster mel-spectrogram generation; 3) Multilingual Extension: We propose three explicit cross-lingual modeling strategies, boundary-aware alignment, token-level concatenation, and instruction-guided generation, establishing practical design principles for zero-shot multilingual emotional TTS that supports Chinese, English, Japanese, and Spanish, and enables robust emotion transfer even without target-language emotional training data; 4) Reinforcement Learning Optimization: we apply GRPO in post-training of the T2S module, improving pronunciation accuracy and natrualness. Experiments show that IndexTTS 2.5 not only supports broader language coverage but also replicates emotional prosody in unseen languages under the same zero-shot setting. IndexTTS 2.5 achieves a 2.28 times improvement in RTF while maintaining comparable WER and speaker similarity to IndexTTS 2.

cs.SD

Pseudo-Hermitian QFT: relativistic scattering and symmetry structure

Unitarity is a cornerstone of quantum theory, ensuring the conservation of probability and information. Although non-Hermitian Hamiltonians are typically associated with open or dissipative systems, pseudo-Hermitian quantum mechanics shows that real spectra and unitary evolution can still emerge through a suitably defined inner product. Motivated by this insight, we extend the pseudo-Hermitian framework to relativistic quantum field theory and construct a consistent formulation of scattering processes. A novel structural feature of this theory is the presence of distinct metric operators for the in and out sectors, connected through a nontrivial metric projector that guarantees global probability conservation under pseudo-unitary time evolution. We further develop a general symmetry formalism, showing that each symmetry generally corresponds to two pseudo-unitary operators associated with the in and out metrics, respectively. Within this framework, the scattering matrix admits a perturbative expansion through the Dyson series and remains Lorentz invariant and unitary, remarkably in complete agreement with the conventional Hermitian case. The fundamental CPT theorem is also shown to hold. Our results provide a rigorous foundation for interacting pseudo-Hermitian quantum field theories and open new directions for exploring their possible physical implications beyond the standard Hermitian paradigm.

hep-th

The anomalous spin-statistics connection arising from pseudo-Hermiticity

We establish a new spin-statistics theorem for a class of free pseudo-Hermitian quantum field theories whose particles furnish unitary irreducible representations of the Poincaré group. In this framework, free pseudo-Hermitian fields with integer spin exhibit fermionic statistics, whereas those with half-integer spin exhibit bosonic statistics, opposite to the conventional case. This reversal arises from defining canonical field operators using pseudo-Hermitian conjugation rather than Hermitian conjugation, thereby circumventing the conventional spin-statistics theorem. The free fields retain locality, Lorentz covariance, and unitary evolution. However, interactions may violate unitarity due to the intrinsically non-Hermitian nature of the full Hamiltonian. We discuss potential resolutions to restore unitarity in interacting theories.

hep-th

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Existing autoregressive large-scale text-to-speech (TTS) models have advantages in speech naturalness, but their token-by-token generation mechanism makes it difficult to precisely control the duration of synthesized speech. This becomes a significant limitation in applications requiring strict audio-visual synchronization, such as video dubbing. This paper introduces IndexTTS2, which proposes a novel, general, and autoregressive model-friendly method for speech duration control. The method supports two generation modes: one explicitly specifies the number of generated tokens to precisely control speech duration; the other freely generates speech in an autoregressive manner without specifying the number of tokens, while faithfully reproducing the prosodic features of the input prompt. Furthermore, IndexTTS2 achieves disentanglement between emotional expression and speaker identity, enabling independent control over timbre and emotion. In the zero-shot setting, the model can accurately reconstruct the target timbre (from the timbre prompt) while perfectly reproducing the specified emotional tone (from the style prompt). To enhance speech clarity in highly emotional expressions, we incorporate GPT latent representations and design a novel three-stage training paradigm to improve the stability of the generated speech. Additionally, to lower the barrier for emotional control, we designed a soft instruction mechanism based on text descriptions by fine-tuning Qwen3, effectively guiding the generation of speech with the desired emotional orientation. Finally, experimental results on multiple datasets show that IndexTTS2 outperforms state-of-the-art zero-shot TTS models in terms of word error rate, speaker similarity, and emotional fidelity. Audio samples are available at: https://index-tts.github.io/index-tts2.github.io/

cs.CL

Two-Dimensional Superconductivity at the CaZrO3/KTaO3 (001) Heterointerfaces

Two-dimensional superconductivity at KTaO3 (KTO) heterointerfaces has sparked intensive investigations since its discovery, yet whether the (001)-oriented KTO interface hosts superconductivity remains to be elucidated. Here, we provide unambiguous evidence of superconductivity in two-dimensional electron gases (2DEGs) at CaZrO3/KTO(001) heterointerfaces, with a superconducting transition TC up to ~0.25 K. Notably, TC increases linearly with carrier density nS over the range of 4.5*10^13~10.3*10^13 cm^-2. Furthermore, superconductivity exhibits a pronounced dependence on crystallographic orientation, with TC rising from 0.25 K for (001) to 1.04 K for (110) and 2.22 K for (111), underscoring the crucial role of interfacial symmetry in the CaZrO3/KTO system. The two-dimensional nature of the superconducting state is corroborated by the Berezinskii-Kosterlitz-Thouless (BKT) transition and the large anisotropy of the upper critical field. For the CaZrO3/KTO(001) sample with nS=7.7*10^13 cm^-2, the estimated Ginzburg-Landau coherence length {\xi}GL=146.4 nm is larger than the superconducting layer thickness dSC=10.1 nm by a factor of ~14.5, confirming significant two-dimensional confinement of the CaZrO3/KTO(001) superconductor. In addition, we demonstrate that the two-dimensional superconductivity at the CaZrO3/KTO(001) interface can be effectively tuned by applying a back gate voltage. Our findings reveal the existence of two-dimensional superconductivity at CaZrO3/KTO(001), providing a new platform for exploring two-dimensional superconductivity at oxide interfaces.

cond-mat.supr-con

Information Suppression in Large Language Models: Auditing, Quantifying, and Characterizing Censorship in DeepSeek

This study examines information suppression mechanisms in DeepSeek, an open-source large language model (LLM) developed in China. We propose an auditing framework and use it to analyze the model's responses to 646 politically sensitive prompts by comparing its final output with intermediate chain-of-thought (CoT) reasoning. Our audit unveils evidence of semantic-level information suppression in DeepSeek: sensitive content often appears within the model's internal reasoning but is omitted or rephrased in the final output. Specifically, DeepSeek suppresses references to transparency, government accountability, and civic mobilization, while occasionally amplifying language aligned with state propaganda. This study underscores the need for systematic auditing of alignment, content moderation, information suppression, and censorship practices implemented into widely-adopted AI models, to ensure transparency, accountability, and equitable access to unbiased information obtained by means of these systems.

cs.CY

Wigner multiplets in QFT: from Wigner degeneracy to Elko fields

We establish the theoretical foundation of the Wigner superposition field, a quantum field framework for spin-1/2 fermions that exhibit a Wigner doublet -- a discrete quantum number arising from nontrivial representations of the extended Poincaré group. In contrast to the previously developed doublet formalism, which treats the Wigner degeneracy as a superficial label, the superposition formalism encodes it directly into the structure of a unified field via a coherent superposition of degenerate spinor fields. By imposing the Lorentz covariance, causality, and canonical quantization, we derive nontrivial constraints on the field configuration, which uniquely identify the Elko field as the consistent realization of the Wigner superposition field. Our analysis further clarifies that although the Elko field is a spinor field, it possesses mass dimension one and obeys the Klein-Gordon rather than the Dirac kinematics. Moreover, we explore the general Elko representation through basis redefinitions, showing that certain traditional properties, such as being eigenspinors of charge conjugation, are artifacts of specific basis choices rather than intrinsic features. Finally, we discuss the physical implications of Elko as a dark matter (DM) candidate. This work lays the foundation for a systematic reformulation of Elko interactions and its phenomenology as a viable component of DM.

hep-th

Wigner multiplets in QFT: dark sector and CPT-violating scenarios

The classification of elementary particles based on unitary irreducible representations of the Poincare group has been a cornerstone of modern Quantum Field Theory (QFT). While the Standard Model (SM) does not inherently include Dark Matter (DM), any fundamental DM candidate should still conform to this classification or its extensions. Beyond the standard representations, Wigner introduced a class of nontrivial states characterized by an additional discrete degree of freedom, known as the Wigner degeneracy. We systematically investigate the QFT of such Wigner degenerate multiplets, particularly focusing on the massive spin-1/2 case. We construct a theoretical framework where the two-fold Wigner spinor fields, $ψ_{\pm\frac{1}{2}}(x)$, form a doublet representation. We analyze their transformation properties under discrete symmetries (C, P, and T), revealing novel mixing effects due to Wigner degeneracy and an emergent accidental U(2) global symmetry. Furthermore, we explore their Yukawa and gauge interactions, demonstrating that such interactions generally break the CPT symmetry. However, we derive conditions for the CPT conservation and discuss potential phenomenological consequences beyond the SM. These results provide new insights into the possible role of Wigner-degenerate states in fundamental physics, particularly in the dark sector.

hep-ph

Elko as an inflaton candidate

Elko is a spin-half fermion with a two-fold Wigner degeneracy and Klein-Gordon dynamics. In this paper, we show that in a spatially flat FLRW space-time, slow-roll inflation can be initiated by the homogeneous Elko fields. The inflaton is a composite scalar field obtained by contracting the spinor field with its dual. This is possible because the background evolution as described by the Friedmann equation is completely determined by the scalar field. This approach has the advantage that we do not need to specify the initial conditions for every component of the spinor fields. We derive the equation of motion for the inflaton and also show that this solution is an attractor. Finally, we examine the slow-roll parameters and the power-spectrum, showing that obtaining a behavior in agreement with observational requirements is hard to be obtained, unless one uses more complicated potentials, which may act a limitation of Elko inflation.

hep-th

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Recently, large language model (LLM) based text-to-speech (TTS) systems have gradually become the mainstream in the industry due to their high naturalness and powerful zero-shot voice cloning capabilities.Here, we introduce the IndexTTS system, which is mainly based on the XTTS and Tortoise model. We add some novel improvements. Specifically, in Chinese scenarios, we adopt a hybrid modeling method that combines characters and pinyin, making the pronunciations of polyphonic characters and long-tail characters controllable. We also performed a comparative analysis of the Vector Quantization (VQ) with Finite-Scalar Quantization (FSQ) for codebook utilization of acoustic speech tokens. To further enhance the effect and stability of voice cloning, we introduce a conformer-based speech conditional encoder and replace the speechcode decoder with BigVGAN2. Compared with XTTS, it has achieved significant improvements in naturalness, content consistency, and zero-shot voice cloning. As for the popular TTS systems in the open-source, such as Fish-Speech, CosyVoice2, FireRedTTS and F5-TTS, IndexTTS has a relatively simple training process, more controllable usage, and faster inference speed. Moreover, its performance surpasses that of these systems. Our demos are available at https://index-tts.github.io.

cs.SD

Correlators for pseudo Hermitian systems

Pseudo-Hermitian system is a class of non-Hermitian system with Hamiltonian satisfying the condition $η^{-1}H^\daggerη=H$. We develop the in-in and Schwinger Keldysh formalism to calculate cosmological correlators for pseudo-Hermitian systems. We study a model consists of massive symplectic fermions coupled to the primordial curvature perturbation. The three-point function for the primordial curvature perturbation is computed up to one-loop and compared to earlier work where the loop correction comes from a massive scalar boson. The two results differ by a minus sign. Therefore, the one loop correction to the three-point function cannot be used to distinguished scalar bosons and symplectic fermions. To conclude, we discuss possibilities where the scalar bosons and symplectic fermions may be distinguished.

hep-th

Mass dimension one fermions in FLRW space-time

Cosmelkology is the study of Elko in cosmology. Elko is a massive spin-half field of mass dimension one. Elko differs from the Dirac and Majorana fermions because it furnishes the irreducible representation of the extended Poincare group with a two-fold Wigner degeneracy where the particle and anti-particle states both have four degrees of freedom. Elko has a renormalizable quartic self interaction which makes it a candidate for self-interacting dark matter. We study Elko in the spatially flat FLRW space-time and find exact solutions in the de Sitter space. By choosing the appropriate solutions and phases, the fields satisfy the canonical anti-commutation relations and have the correct time evolutions in the flat space limit.

hep-th

The Machiavellian frontier of stable mechanisms

The impossibility theorem in Roth (1982) states that no stable mechanism satisfies strategy-proofness. This paper explores the Machiavellian frontier of stable mechanisms by weakening strategy-proofness. For a fixed mechanism $φ$ and a true preference profile $\succ$, a $(φ,\succ)$-boost mispresentation of agent i is a preference of i that is obtained by (i) raising the ranking of the truth-telling assignment $φ_i(\succ)$, and (ii) keeping rankings unchanged above the new position of this truth-telling assignment. We require a matching mechanism $φ$ neither punish nor reward any such misrepresentation, and define such axiom as $φ$-boost-invariance. This is strictly weaker than requiring strategy-proofness. We show that no stable mechanism $φ$ satisfies $φ$-boost-invariance. Our negative result strengthens the Roth Impossibility Theorem.

econ.TH

Tracing the Unseen: Uncovering Human Trafficking Patterns in Job Listings

In the shadow of the digital revolution, the insidious issue of human trafficking has found new breeding grounds within the realms of social media and online job boards. Previous research efforts have predominantly centered on identifying victims via the analysis of escort advertisements. However, our work shifts the focus towards enabling a proactive approach: pinpointing potential traffickers before they lure their preys through false job opportunities. In this study, we collect and analyze a vast dataset comprising over a quarter million job postings collected from eight relevant regions across the United States, spanning nearly two decades (2006-2024). The job boards we considered are specifically catered towards Chinese-speaking immigrants in the US. We classify the job posts into distinct groups based on the self-reported information of the posting user. Our investigation into the types of advertised opportunities, the modes of preferred contact, and the frequency of postings uncovers the patterns characterizing suspicious ads. Additionally, we highlight how external events such as health emergencies and conflicts appear to strongly correlate with increased volume of suspicious job posts: traffickers are more likely to prey upon vulnerable populations in times of crises. This research underscores the imperative for a deeper dive into how online job boards and communication platforms could be unwitting facilitators of human trafficking. More importantly, it calls for the urgent formulation of targeted strategies to dismantle these digital conduits of exploitation.

cs.SI