Searcharxiv⌕ Search

arXiv subjects

Peng Cheng

Publications and source records attributed to Peng Cheng.

At least 55 records · Page 3Linked to original sources

MixFake: Benchmarking and Enhancing Audio Deepfake Detection in Diverse Real-world Mixed Audio

Speech deepfake detection has achieved remarkable success in clean environments but faces significant challenges in complex, real-world scenarios where speech is often mixed with background music or noise. Current state-of-the-art methods rely on semantic features from self-supervised learning (SSL) models, which often fail when processing non-speech or mixed-source audio. In this paper, we first introduce MixFake, a large-scale benchmark dataset designed to simulate diverse acoustic environments with varying SNR levels and mixed authenticity components. To address the "semantic-centric" limitation, we propose a Multi-stream Prompt Tuning framework that injects signal-level priors into SSL backbones. By integrating base, frequency, and texture streams through deep prompt injection, our model effectively captures acoustic artifacts. Experimental results demonstrate that our method significantly outperforms existing baselines, achieving a 0.95% EER in foreground detection and a substantial 7.72% absolute improvement in complex background detection tasks. Our dataset and code are available at https://github.com/saltfish233/MixFake.

cs.SD↗

From Patches to Trajectories: Privileged Process Supervision for Software-Engineering Agents

Supervised fine-tuning (SFT) on long teacher trajectories is the dominant way to instill investigation and reasoning in open software-engineering (SWE) agents. Since every retained response becomes an imitation target, the student inherits the final outcome and intermediate flaws, including ungrounded leaps and redundant loops. High-quality training data must be effective(each step is grounded and narrows the agent's epistemic gap to the correct fix) and efficient(each step is information-bearing rather than redundant or looping). Existing recipes filter or relabel teacher rollouts using only a binary terminal verifier, which does not directly target these axes and provides no supervision on instances where the teacher fails. Most real issue includes a developer-authored reference patch, $p^\star$, revealing the file paths, runtime behaviors, and coding conventions presupposed by the correct fix, yet standard pipelines discard it. We propose Patches-to-Trajectories (P2T), which uses $p^\star$ as privileged information during curation and formulates trajectory construction as bi-objective optimization over per-step effectiveness and trajectory length. A reverse phase distills $p^\star$ into a latent process graph, $G^\star$, of contextual facts and solution milestones. A forward phase curates trajectories from blinded teacher continuations by scoring per-step progress against $G^\star$ under a leakage-blocking groundedness check and retaining the shortest effective segments. Using only 1.8k curated SWE-Gym instances, P2T improves effectiveness and efficiency over outcome-filtered SFT and its tool-error-masking variant. On SWE-bench Verified, it raises Pass@1 by up to 10.8 points while reducing per-instance inference cost by ~15%, with consistent gains on SWE-bench Lite. Size-matched ablations and qualitative analysis further isolate trajectory quality from data scale.

cs.SE↗

Beyond Sunk Costs: Boosting LLM Pre-training Efficiency via Orthogonal Growth of Mixture-of-Experts

As the computational demands for pre-training Large Language Models (LLMs) continue to surge, the need for efficient training paradigms becomes critical. Despite the vast resources already invested in existing pre-trained checkpoints, these assets often remain under-leveraged due to architectural limitations. We introduce an "orthogonal growth" strategy designed to "recycle" these checkpoints by strategically expanding their parameters prior to continued training. Our method focuses on optimizing converged Mixture-of-Experts (MoE) models through two dimensions: interpositional layer copying for increased depth and noisy expert duplication for expanded width. Through extensive scaling laws analysis, we demonstrate a strong positive correlation between the "sunk cost" (prior investment) and the final model accuracy. Empirical results on models up to 70B parameters and 1T tokens show that our recycling approach yields a 10.6% accuracy improvement compared to training from scratch under identical extra compute budgets. This work provides a cost-effective blueprint for sustainable large-scale LLM development.

cs.LG↗

Structure-BiEval: A Self-Supervised, Dual-Track Framework for Decoupling Structure and Content in LLM Evaluation for Web Information Systems

As Large Language Models (LLMs) evolve into the core of Web-based autonomous agents and complex Web Information Systems, their ability to faithfully translate natural language into rigorous structured formats has become paramount, as this capability is critical for Web API invocation and data exchange. However, evaluating this structural fidelity in Web-native payloads remains a challenge: traditional text metrics fail to capture topological consistency in semi-structured Web data, while manual evaluation is prohibitively costly. To address this, we propose Structure-BiEval, a novel self-supervised framework for quantitative, annotation-free assessment tailored for Web data engineering. By leveraging deterministic Intermediate Representations, our framework effectively decouples structure from content, utilizing Content Semantic Accuracy and Normalized Tree Edit Distance as precise metrics. We empirically benchmark 15 state-of-the-art LLMs across dual Web structural topologies, namely Hierarchical Data (Web backend payloads) and Tabular Data (Web frontend presentation). The results reveal substantial variability in structural performance, including cases where mid-sized models unexpectedly outperform larger counterparts in Web data formatting. Furthermore, our findings show that deep recursive nesting poses a consistent challenge for Web agents across varying parameter scales.

cs.CL↗

Beyond Content: A Comprehensive Speech Toxicity Dataset and Detection Framework Incorporating Paralinguistic Cues

Toxic speech detection has become a crucial challenge in maintaining safe online communication environments. However, existing approaches to toxic speech detection often neglect the contribution of paralinguistic cues, such as emotion, intonation, and speech rate, which are key to detecting speech toxicity. Moreover, current toxic speech datasets are predominantly text-based, limiting the development of models that can capture paralinguistic cues.To address these challenges, we present ToxiAlert-Bench, a large-scale audio dataset comprising over 30,000 audio clips annotated with seven major toxic categories and twenty fine-grained toxic labels. Uniquely, our dataset annotates toxicity sources -- distinguishing between textual content and paralinguistic origins -- for comprehensive toxic speech analysis.Furthermore, we propose a dual-head neural network with a multi-stage training strategy tailored for toxic speech detection. This architecture features two task-specific classification headers: one for identifying the source of sensitivity (textual or paralinguistic), and the other for categorizing the specific toxic type. The training process involves independent head training followed by joint fine-tuning to reduce task interference. To mitigate data class imbalance, we incorporate class-balanced sampling and weighted loss functions.Our experimental results show that leveraging paralinguistic features significantly improves detection performance. Our method consistently outperforms existing baselines across multiple evaluation metrics, with a 21.1% relative improvement in Macro-F1 score and a 13.0% relative gain in accuracy over the strongest baseline, highlighting its enhanced effectiveness and practical applicability.

cs.SD↗

Reducing the Costs of Proof Synthesis on Rust Systems by Scaling Up a Seed Training Set

Large Language Models (LLMs) are widely used for code generation. However, the correctness of code generated by LLMs remains a concern. A potential remedy to this concern is to have LLMs generate formal correctness proofs along with such code. However, compared with code generation, code-proof generation requires much higher reasoning capability and has much less existing data to learn from. In this paper, we present VeruSyn, a data synthesis pipeline for Verus, a state-of-the-art verification tool for system software written in Rust. Through self-synthesis and tutorial-based synthesis, VeruSyn achieves much larger scale and Verus-feature coverage than previous data-synthesis techniques designed for Verus; VeruSyn also supplements its dataset with long-chain-of-thought (CoT) data through agent trajectory synthesis. With VeruSyn, we synthesize the largest set of Verus verified programs: 6.9 million Rust programs, each with a formal specification and a proof that it meets that specification. This dataset lets us create a fine-tuned Qwen2.5-Coder-32B-Instruct model with appealing cost-proof tradeoff compared with state-of-the-art commercial models like Claude Sonnet 4.5. It also significantly outperforms models like o4-mini and previously proposed research models.

cs.SE↗

Noncollinear antiferromagnetic structure and physical properties of CrRhAs with distorted kagome lattice

CrRhAs was theoretically proposed to be a kagome metal with unusual magnetic ground states; however, little is known about its magnetic structure and physical properties experimentally. Here, we present an experimental investigation of CrRhAs with ZrNiAl-type structure and a distorted Cr kagome lattice. CrRhAs is an antiferromagnet with TN = 149 K. Powder neutron diffraction analysis reveals a noncollinear antiferromagnetic structure with propagation vector k = (1/3, 1/3, 1/2), which features a ferromagnetic second nearest neighbor coupling in the kagome plane that is different from the prediction in previous density functional theory calculations. Furthermore, CrRhAs exhibits anomalous electrical transport properties which are possibly related to multiband effects and strong spin fluctuations. For the temperature-dependent longitudinal resistivity \r{ho}xx, it is semiconductinglike above TN and becomes metallic below TN . The Hall coefficients exhibit two sign changes near 70 and 300 K. Combined with the results of heat capacity measurements, a large Kadowaki-Woods ratio α = 33.9 μΩ cm mol2 K2/J2 is obtained. The above results suggest CrRhAs is a strongly correlated kagome metal with multiband and noncollinear magnetic structure features.

cond-mat.str-el↗

TSGuard: Automated User-Centric Incident Diagnosis for AI Workloads in the Cloud

AI workloads incur frequent failures and incidents from the underlying infrastructure. The current incident management workflow follows a provider-centric paradigm, where users report incidents to the infrastructure provider who then conducts troubleshooting. Due to the large number of incidents and the manual nature of the troubleshooting process, the provider often takes several days to resolve an incident, resulting in operational delays and productivity loss. To address these challenges, we present TSGuard, a user-centric multi-agent system that delivers immediate incident diagnosis to users who deploy the workloads. The core innovation of TSGuard is twofold: (1) constructing domain-specific knowledge bases by mining historical on-call experiences in the offline phase, and (2) mimicking human expert diagnosis via structured reasoning and iterative trial-and-error in the online phase. Evaluation using production incident records from Microsoft Azure demonstrates that TSGuard significantly outperforms state-of-the-art baselines, improving diagnostic accuracy by 19.8%. Furthermore, TSGuard reduces the average verification time by 63.4% compared to the sequential execution baseline.

cs.SE↗

Evidence for monopole-like topological magnetoelectric effect in image potential states

Magnetic monopoles, hypothetical particles behaving as isolated magnetic charges, have long been predicted by theories beyond the standard model but remain elusive in experimental detection. Subsequently, Xiaoliang Qi et al. proposed that magnetic monopoles can be constructed in real space by introducing an active electric field at the interface between a topological insulator and vacuum [Science 323, 1184 (2009)]. Here we use scanning tunneling microscopy in the field-emission regime to realize an active electric-field geometry at the surface of the higher-order topological insulator Bi(111), and observe an anomalous splitting of the image potential states (IPSs). By tuning the dielectric properties of the substrate and film thickness, we considered that the peak-splitting of IPSs is related to the radially active electric field and topological surface states. Combined with phenomenological analysis, this peaksplitting can be attributed to the equivalent magnetic field of monopole-like topological magnetoelectric response. This work establishes field-emission IPS spectroscopy as a sensitive platform for generating active electric fields and probing the resulting image magnetic monopoles at topological surfaces.

cond-mat.mtrl-sci↗

Dai-Freed anomalies and level matching in heterotic asymmetric orbifolds

We study asymmetric orbifolds of the $E_8\times E_8$ heterotic string from the perspective of worldsheet Dai-Freed anomalies. Focusing on cyclic symmetries $G = \mathbb{Z}_m$ that act chirally on the fermions and symmetrically on the bosons, we compute the corresponding spin-bordism invariants and derive the conditions for the vanishing of global anomalies from this perspective. In the fermionic description, these conditions are exactly the familiar level-matching constraints, together with the additional mod-2 conditions that appear for even $m$. We then discuss the same conditions from the transformation properties of higher-genus fermion partition functions and explain how the anomaly is matched under bosonization for a large class of inner automorphisms of the $E_8\times E_8$ lattice theory. This gives an interpretation of the standard consistency conditions for asymmetric heterotic orbifolds from the Dai-Freed anomaly perspective.

hep-th↗

On the Cuspy Structure of Rotating Wormhole Shadows

We investigate the shadow cast by a rotating traversable wormhole in the Teo class endowed with a general redshift function, with particular emphasis on the emergence of cuspy structures. The shadow boundary is the common envelope of two critical orbit families: unstable circular orbits outside the throat and orbits at the throat itself. The formation of cusps, marking the transition between smooth and cuspy shadow boundaries, only becomes possible when the redshift parameter $λ$ is allowed to vary. Moreover, we uncover a universal critical value $λ_c$ that signals the onset of the cusp. A phase diagram characterized by the spin and redshift parameters reveals four distinct morphologies: smooth, cuspy, ears touching, and throat drowning. The morphology of the wormhole shadow may provide observational diagnostics for the different compact objects in future high-resolution imaging observations.

gr-qc↗

Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-training

Continual pre-training on small-scale task-specific data is an effective method for improving large language models in new target fields, yet it risks catastrophic forgetting of their original capabilities. A common solution is to re-weight training data mixtures from source and target fields on a domain space to achieve balanced performance. Previous domain reweighting strategies rely on manual designation with certain heuristics based on human intuition or empirical results. In this work, we prove that more general heuristics can be parameterized by proposing Data Mixing Agent, the first model-based, end-to-end framework that learns to re-weight domains. The agent learns generalizable heuristics through reinforcement learning on large quantities of data mixing trajectories with corresponding feedback from an evaluation environment. Experiments in continual pre-training on math reasoning show that Data Mixing Agent outperforms strong baselines in achieving balanced performance across source and target field benchmarks. Furthermore, it generalizes well across unseen source fields, target models, and domain spaces without retraining. Direct application to the code generation field also indicates its adaptability across target domains. Further analysis showcases the agents' well-aligned heuristics with human intuitions and their efficiency in achieving superior model performance with less source-field data.

cs.LG↗

Supercell-size scaling of moiré band flatness

In moiré superlattices, the band flatness governs the degree of wave localization, which is central to harnessing emergent phenomena and designing functional meta-devices. While research has focused on the magic conditions such as magic angle and magic distance for optimal flatness, a fundamental understanding of how flatness changes with the supercell size has remained elusive. Here, we establish a universal scaling between band flatness and supercell size. Theoretically, by recognizing the statistical equivalence between structural perturbations in moiré superlattices and disordered systems, we introduce the Thouless number to evaluate the strength of moiré localization. This approach allows us to establish a scaling theory for the evolution of band flatness with the supercell size, from which an analytical expression is derived. Our full-wave simulations with one-dimensional and two-dimensional moiré superlattices show excellent agreement with the theoretical prediction. Our work reveals a general scaling law for moiré band flatness, offering a new perspective for understanding and designing moiré-based resonant systems.

physics.optics↗

LUMINA: LLM-Guided GPU Architecture Exploration via Bottleneck Analysis

GPU design space exploration (DSE) for modern AI workloads, such as Large-Language Model (LLM) inference, is challenging because of GPUs' vast, multi-modal design spaces, high simulation costs, and complex design optimization objectives (e.g. performance, power and area trade-offs). Existing automated DSE methods are often prohibitively expensive, either requiring an excessive number of exploration samples or depending on intricate, manually crafted analyses of interdependent critical paths guided by human heuristics. We present LUMINA, an LLM-driven GPU architecture exploration framework that leverage AI to enhance the DSE efficiency and efficacy for GPUs. LUMINA extracts architectural knowledge from simulator code and performs sensitivity studies to automatically compose DSE rules,which are auto-corrected during exploration. A core component of LUMINA is a DSE Benchmark that comprehensively evaluates and enhances LLMs' capabilities across three fundamental skills required for architecture optimization, which provides a principled and reproducible basis for model selection and ensuring consistent architectural reasoning. In the design space with 4.7 million possible samples, LUMINA identifies 6 designs of better performance and area than an A100 GPU efficiently, using only 20 steps via LLM-assisted bottleneck analysis. In comparison, LUMINA achieves 17.5x higher than design space exploration efficiency, and 32.9% better designs (i.e. Pareto Hypervolume) than Machine-Learning baselines, showcasing its ability to deliver high-quality design guidance with minimal search cost.

cs.AR↗

STEP: Detecting Audio Backdoor Attacks via Stability-based Trigger Exposure Profiling

With the widespread deployment of deep-learning-based speech models in security-critical applications, backdoor attacks have emerged as a serious threat: an adversary who poisons a small fraction of training data can implant a hidden trigger that controls the model's output while preserving normal behavior on clean inputs. Existing inference-time defenses are not well suited to the audio domain, as they either rely on trigger over-robustness assumptions that fail on transformation-based and semantic triggers, or depend on properties specific to image or text modalities. In this paper, we propose STEP (Stability-based Trigger Exposure Profiling), a black-box, retraining-free backdoor detector that operates under hard-label-only access. Its core idea is to exploit a characteristic dual anomaly of backdoor triggers: anomalous label stability under semantic-breaking perturbations, and anomalous label fragility under semantic-preserving perturbations. STEP profiles each test sample with two complementary perturbation branches that target these two properties respectively, scores the resulting stability features with one-class anomaly detectors trained on benign references, and fuses the two scores via unsupervised weighting. Extensive experiments across seven backdoor attacks show that STEP achieves an average AUROC of 97.92% and EER of 4.54%, substantially outperforming state-of-the-art baselines, and generalizes across model architectures, speech tasks, an open-set verification scenario, and over-the-air physical-world settings.

cs.CR↗

The stringy geometry of integral cohomology in mirror symmetry

We examine the physical significance of torsion co-cycles in the cohomology of a projective Calabi-Yau three-fold for the (2,2) superconformal field theory (SCFT) associated to the non-linear sigma model with such a manifold as a target space. There are two independent torsion subgroups in the cohomology. While one is associated to an orbifold construction of the SCFT, the other encodes the possibility of turning on a topologically non-trivial flat gerbe for the NS-NS B-field. Inclusion of these data enriches mirror symmetry by providing a refinement of the familiar structures and points to a generalization of the duality symmetry, where the topology of the flat gerbe enters on the same footing as the topology of the underlying manifold.

hep-th↗

Observation of Iso-Symmetric Structural and Lifshitz Transitions in Quasi-one-dimensional CrNbSe$_5$

Chalcogenides-rich transition metal compounds host a rich landscape of emergent quantum phenomena that are intimately governed by their quasi-one-dimensional chemical-bonding frameworks and their response to external perturbations such as pressure. Here, we report a pressure-induced iso-symmetric structural transition in the quasi-one-dimensional compound CrNbSe$_5$, in which the electronic ground state is controlled not by symmetry breaking but by a continuous reorganization of local bonding interactions. Applied pressure reversibly tunes CrNbSe$_5$ between semiconducting and semimetallic states, enabling access to low- and high-carrier electronic regimes through direct modulation of metal-chalcogen bonding. High-pressure single-crystal X-ray diffraction directly resolves the evolution of Cr-Se and Nb-Se bond distances, coordination polyhedra, and connectivity, revealing a fully reversible semimetal-semiconductor-semimetal transition driven by gradual yet cooperative bond rearrangements within a preserved crystallographic symmetry. In contrast to chemical substitution, which irreversibly alters composition and introduces disorder, pressure acts as a clean, continuous control parameter that reshapes the bonding landscape without disrupting structural symmetry. These results establish CrNbSe$_5$ as a model system for electronically driven phase switching via tunable chemical bonding, highlighting iso-symmetric bond reorganization as a powerful design principle for pressure-controlled electronic and spintronic functionalities.

cond-mat.mtrl-sci↗

TrasMuon: Trust-Region Adaptive Scaling for Orthogonalized Momentum Optimizers

Muon-style optimizers leverage Newton-Schulz (NS) iterations to orthogonalize updates, yielding update geometries that often outperform Adam-series methods. However, this orthogonalization discards magnitude information, rendering training sensitive to step-size hyperparameters and vulnerable to high-energy bursts. To mitigate this, we introduce TrasMuon (\textbf{T}rust \textbf{R}egion \textbf{A}daptive \textbf{S}caling \textbf{Muon}). TrasMuon preserves the near-isometric geometry of Muon while stabilizing magnitudes through (i) global RMS calibration and (ii) energy-based trust-region clipping. We demonstrate that while reintroducing adaptive scaling improves optimization efficiency, it typically exacerbates instability due to high-energy outliers. TrasMuon addresses this by defining a trust region based on relative energy ratios, confining updates to a stable zone. Empirical experiments on vision and language models demonstrate that TrasMuon converges faster than baselines. Furthermore, experiments without warmup stages confirm TrasMuon's superior stability and robustness.

cs.LG↗