SearcharxivSearch

arXiv subjects

Runkun Chen

Publications and source records attributed to Runkun Chen.

5 recordsLinked to original sources

Audio Language Model for Deepfake Detection Grounded in Acoustic Chain-of-Thought

Deepfake speech detection systems are often limited to binary classification tasks and struggle to generate interpretable reasoning or provide context-rich explanations for their decisions. These models primarily extract latent embeddings for authenticity detection but fail to leverage structured acoustic evidence such as prosodic, spectral, and physiological attributes in a meaningful manner. This paper introduces CoLMbo-DF, a Feature-Guided Audio Language Model that addresses these limitations by integrating robust deepfake detection with explicit acoustic chain-of-thought reasoning. By injecting structured textual representations of low-level acoustic features directly into the model prompt, our approach grounds the model's reasoning in interpretable evidence and improves detection accuracy. To support this framework, we introduce a novel dataset of audio pairs paired with chain-of-thought annotations. Experiments show that our method, trained on a lightweight open-source language model, significantly outperforms existing audio language model baselines despite its smaller scale, marking a significant advancement in explainable deepfake speech detection.

cs.SD

GameDevBench: Evaluating Agentic Capabilities Through Game Development

Despite rapid progress on coding agents, progress on their multimodal counterparts has lagged behind. A key challenge is the scarcity of evaluation testbeds that combine the complexity of software development with the need for deep multimodal understanding. In game development, agents must navigate large, dense codebases while manipulating intrinsically multimodal assets such as shaders, sprites, and animations within a visual game scene. We present GameDevBench, the first benchmark for evaluating agents on game development tasks. GameDevBench consists of 333 tasks derived from web and video tutorials. Tasks require significant multimodal understanding and are complex: the average solution requires over three times the lines of code and file changes compared to prior software development benchmarks. Agents struggle with game development, with the best agent and method solving only 53.8% of tasks. We find a strong correlation between perceived task difficulty and multimodal complexity, with average success rate dropping from 51.4% on gameplay-oriented tasks to 33.0% on 2D graphics tasks. To improve multimodal capability, we introduce two simple image- and video-based feedback mechanisms for agents. Despite their simplicity, these methods consistently improve performance, increasing GPT-5.4's performance from 41.1% to 52.0% when given visual feedback.

cs.AI

ML-Master: Towards AI-for-AI via Integration of Exploration and Reasoning

As AI capabilities advance toward and potentially beyond human-level performance, a natural transition emerges where AI-driven development becomes more efficient than human-centric approaches. A promising pathway toward this transition lies in AI-for-AI (AI4AI), which leverages AI techniques to automate and optimize the design, training, and deployment of AI systems themselves. While LLM-based agents have shown the potential to realize AI4AI, they are often unable to fully leverage the experience accumulated by agents during the exploration of solutions in the reasoning process, leading to inefficiencies and suboptimal performance. To address this limitation, we propose ML-Master, a novel AI4AI agent that seamlessly integrates exploration and reasoning by employing a selectively scoped memory mechanism. This approach allows ML-Master to efficiently combine diverse insights from parallel solution trajectories with analytical reasoning, guiding further exploration without overwhelming the agent with excessive context. We evaluate ML-Master on the MLE-Bench, where it achieves a 29.3% average medal rate, significantly surpassing existing methods, particularly in medium-complexity tasks, while accomplishing this superior performance within a strict 12-hour time constraint-half the 24-hour limit used by previous baselines. These results demonstrate ML-Master's potential as a powerful tool for advancing AI4AI.

cs.AI

Canalization acoustic phonon polaritons in metal-MoO3-metal sandwiched structures for nano-light guiding and manipulation

We theoretically propose and study in-plane anisotropic acoustic phonon polaritons (APhPs) based on a layered structure consisting of a monolayer (or few layers) α-phase molybdenum trioxide (α-MoO3) sandwiched between two metal layers. We find that the APhPs in the proposed sandwiched structures are a canalization (highly directional) electromagnetic mode propagating along with the layers and at the same time exhibit extreme electromagnetic-field confinement surpassing any other type of phonon-polariton modes. When a double layer of α-MoO3 is sandwiched by two Au layers, twisting the two α-MoO3 layers can adjust the interlayer polaritonic coupling and thus manipulate the in-plane propagation of the highly confined APhPs. Our results illustrate that the metal-MoO3-metal sandwiched structures are a promising platform for light guiding and manipulation at ultimate scale.

physics.optics

Investigation of plasmonic evolution of atomically size-selected Au clusters by electron energy loss spectrum--from solid state to molecular scale

Versatile quantum modes emerge for plasmon describing the collective oscillations of free electrons in metallic nanoparticles when the particle sizes are greatly reduced. Rather than traditional nanoscale study, the understanding of quantum plasmon desires extremal atomic control of the nanoparticles, calling for size dependent plasmon measurement over a series of nanoparticles with atomically adjustable atom number over several orders of magnitude. Here we report the N dependent plasmonic evolution of atomically size selected gold particles with N= 100 70000 using electron energy loss (EEL) spectroscopy in a scanning transmission electron microscope. The EEL mapping assigns a feature at 2.7 eV as the bulk plasmon and another at 2.4 eV as surface plasmon, which evolution reveals three regimes. When N decreases from 70000 to 887, the bulk plasmon stays unchanged while the surface plasmon exhibits a slight red shift from 2.4 to 2.3 eV. It can be understood by the dominance of classical plasmon physics and electron boundary scattering induced retardation. When N further decreases from 887 to 300, the bulk plasmon disappears totally and the surface plasmon shows a steady blueshift, which indicates that the quantum confinement emerges and modifies the intraband transition. When N 100 300, the plasmon is split to three fine features, which is attributed to superimposed single electron transitions between the quantized molecular like energy level by the time dependent density functional theory calculations. The surface plasmon's excitation ratio has a scaling law with an exponential dependence on N ( N^0.669), essentially the square of the radius. A unified evolution picture from the classical to quantum, molecular plasmon is thus demonstrated.

physics.atm-clus