SearcharxivSearch

arXiv subjects

Yifan Zhang

Publications and source records attributed to Yifan Zhang.

At least 19 recordsLinked to original sources

oMeBench: Towards Robust Benchmarking of LLMs in Organic Mechanism Elucidation and Reasoning

Organic reaction mechanisms describe the step-wise elementary processes by which reactants transform into intermediates and products, and are fundamental to understanding chemical reactivity and guiding molecular and reaction de-sign. While large language models (LLMs) have shown promise on chemical tasks such as synthesis design, it remains unclear to what extent this reflects genuine chemical reasoning capabilities: the ability to generate chemically valid intermediates, maintain consistency across reaction steps, and follow logically coherent multi-step pathways. To investigate this, we introduce oMeBench, the first large-scale, expert-curated benchmark for organic mechanism reasoning, comprising over 10,000 annotated mechanistic steps with reaction type labels, intermediate structures, and difficulty ratings. To enable fine-grained evaluation, we further propose oMeS, a dynamic scoring framework that jointly assesses step-level logical consistency and chemical structural similarity. Systematic evaluation of state-of-the-art LLMs reveals that while current models exhibit promising chemical intuition, they struggle to produce correct and consistent reasoning across multi-step mechanisms. Notably, combining prompting strategies with fine-tuning enables smaller-scale models to achieve performance comparable to closed-source frontier models. We hope oMeBench will serve as a rigorous foundation for advancing AI systems toward genuine chemical reasoning.

cs.AI

Agora: Git as Shared Memory for Collective AutoResearch

Autonomous research loops such as AutoResearch show that one coding agent can improve a training setup unattended. Run several of them and each session starts from scratch, so more agents tend to mean more duplicated search rather than more discovery. Agora is a shared memory for such agents: research is recorded as an append-only directed acyclic graph (DAG) stored in Git, so that every claim is a commit anyone can check out and rerun. Each result, insight, hypothesis, verification, and report is an immutable commit whose parent edges say what it builds on; a derived index exposes the frontier, the neglected branches, and the verification status of each claim, and a diversity-aware selection rule keeps the community from collapsing onto one leader. We describe the system and report its first sustained use: a run of nearly 12 days in which 13 language-model workers, with no assigned tasks and no central planner, worked on a weight-transfer problem. Given 141 pretrained donor models and a frozen 119.6M-parameter attention-SSM hybrid whose dimensions match no donor, the workers had to initialize the target without training data or gradient updates. They published 1,703 contributions and drove the evaluator from 3.39 to 1.899 bits per byte, closing 62% of the gap to a trained GPT-2 124M. The winning recipe compresses donor next-token statistics into the target's embedding and output head, then adds a short-range context signal through sparse edits to attention, feed-forward, and state-space blocks. Its 145-commit ancestry spans 15 accounts, and 165 independent reproductions were posted, none of which failed. We describe the single mid-run human intervention that pulled the community out of a monoculture, what the trace does and does not establish, and the controlled comparison that would settle whether shared research state improves discovery per unit of compute.

cs.LG

A Lattice Boltzmann Method with Adaptive Relaxation Parameter and Lax--Friedrichs-Type Equilibrium for Hyperbolic Systems

In this paper, a lattice Boltzmann method with an adaptive relaxation parameter and a Lax--Friedrichs-type equilibrium is proposed for hyperbolic systems with source terms. The equilibrium distribution recovers the conservative variables and physical fluxes while incorporating dissipation determined by characteristic-speed bounds. To balance accuracy and robustness, the relaxation parameter is selected from a local smoothness indicator based on characteristic projections. In smooth regions, the parameter approaches the low-dissipation limit, retaining second-order accuracy; near discontinuities, it is automatically reduced to introduce localized dissipation and suppress nonphysical oscillations. The stabilization acts directly through the local collision step and preserves the standard collide-and-stream structure, without a posteriori recomputation or interface-based limiting. Maxwell iteration establishes second-order consistency in smooth regions, and a weighted $L^2$-stability estimate is proved for linear hyperbolic systems with periodic boundary conditions under a standard CFL condition. Numerical experiments for scalar advection, Euler, shallow-water, and reactive Euler equations demonstrate the expected accuracy for smooth solutions and robust resolution of challenging one- and two-dimensional discontinuous problems, including wet--dry fronts and reactive discontinuities. A large-scale simulation of a 2D cellular detonation further demonstrates the capability of the method to resolve long-time multidimensional shock--reaction interactions.

math.NA

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning

Agentic reinforcement learning requires infrastructure that researchers can modify without sacrificing model scale or control over agent execution. We present Molt, a lightweight PyTorch-native framework that combines trillion-parameter training with standard agent interfaces. Molt integrates four capabilities: a compact training implementation built on composable model parallelism; unified OpenAI and Anthropic interfaces with automatic trajectory segmentation after context compaction; fully asynchronous rollout and optimization; and distributed experience storage for long, multimodal trajectories. Existing agents retain their execution and context-management logic while a shared capture layer records generated tokens and behavior probabilities. Rollout workers place heavy experience payloads in Ray's object store, and trainer ranks retrieve their assigned experiences by reference, avoiding a centralized gather of the full rollout batch. The framework-owned RL implementation comprises approximately 9.2K Python code lines, and its rollout, weight-refit, and training-update path has executed end to end on a one-trillion-parameter policy. On a 35B multimodal mixture-of-experts workload, speculative decoding accelerates the generation stage by 5.14x, and optimizer offload reduces peak actor memory by 18.3 GB. Together, these results establish a compact training framework for agentic RL research at trillion-parameter scale.

cs.LG

Enabling Creative Exploration for Vibe Design Agents

Vibe design agents turn natural-language briefs into rendered interfaces and frontend code. Yet a useful design agent should do more than produce one valid page: it should help users explore coherent alternatives. Increasing token-level temperature is a blunt solution because it varies aesthetic decisions and syntax-sensitive code at the same time. We instead separate exploration from implementation through an inference architecture that makes design direction an explicit intermediate decision. Inspired by Verbalized Sampling, a pre-pass proposes structured design specifications with typicality scores, an external selector samples one, and the downstream generator realizes the selected specification together with the original request under fixed settings. We apply this approach to UI themes and visual-asset prompts. Across 168 prompts, with 1,255 paired comparisons per temperature for each intervention, theme sampling broadens observed selection coverage and screenshot variation, while LLM-judge preferences vary across interventions, prompt complexity, and viewport. In an online experiment with more than 300,000 tasks, the observed code-export increase remains statistically uncertain, while fewer negative feedback events coexist with more correction interactions and modest operational costs. Together, these findings identify structured design specifications as a practical control point for exploring alternative UI concepts while keeping downstream generation settings fixed.

cs.AI

Discrete q-Hermitian Clifford analysis

We develop a discrete $q$-Hermitian Clifford calculus based on coordinatewise Jackson differences on multiplicative $q$-lattices. The resulting Hermitian Jackson--Dirac operators are nilpotent, factor a $q$-Laplacian, and have occupancy-dependent Euler anticommutators. These give explicit Fischer projectors and homotopies, but the local calculus is not controlled by total degree. We prove a two-face Cauchy--Kovalevskaya theorem for polynomials and extend it to a $q$-analytic Jackson--Fischer class. A divided-power conjugation recovers scalar Euler relations and transfers the classical joint Fischer decomposition and dimension formulas. The Jackson--Fischer completion is a vector-valued $q$-Fock space whose polarized nullspaces have projected reproducing kernels; the joint kernel is obtained by alternating projections. We also determine the linear symmetry group of the multiplicative lattice.

math.CV

Topological Surface Charge Detection via Terahertz Time-domain Spectroscopy

The topological magnetoelectric effect (TME) is a condensed-matter realization of the four-dimensional quantum Hall effect (4D-QHE), manifesting as quantized surface charge accumulation proportional to an applied magnetic field. To date, however, no optical technique has been developed to directly probe this charge accumulation. Here, we demonstrate a terahertz time-domain spectroscopy method for direct detection of surface charge accumulation---a physical quantity relevant to both the 4D-QHE and 2D-QHE, in sharp contrast to the previous optical measurements, which focused on the Hall conductivity $σ_{xy}$ of the 2D-QHE. Using a chromium-doped (Bi,Sb)$_2$Te$_3$ thin film, we achieve sub-milliradian Faraday rotation precision. Extending this scheme to axion insulators, we predict that the TME gives rise to an imaginary Faraday rotation linear in frequency, whose slope directly reflects the single-surface charge density. With further improvements in sample thickness and precision, this approach offers a viable pathway toward direct verification of the TME and 4D-QHE.

cond-mat.mes-hall

LLaTSA: Large Language Model-Aligned General-Purpose Transient Stability Analysis

Dynamic trajectory prediction has become an important paradigm for data-driven transient stability analysis (TSA), yet most existing predictors remain system-specific and require substantial retraining when network configurations, generation mixes, or state-variable sets change. Uni-TSA introduced a general-purpose TSA framework that combines channel-independent modeling with a pretrained large language model (LLM) predictor. Nevertheless, its application to heterogeneous systems is limited by ambiguity in short observations, a mismatch between numerical trajectories and LLM embeddings, neglected coupling among state variables, and the high inference cost of dense backbones. This paper proposes LLaTSA, an LLM-aligned framework for general-purpose trajectory-based TSA. LLaTSA first incorporates operating conditions, disturbance attributes, and state-variable identity through a structured textual prefix. It then aligns normalized temporal patches with a TSA-related vocabulary before processing them with a pretrained sparse decoder-only mixture-of-experts (MoE) backbone. A state-variable coupling module captures coordinated post-fault evolution, while teacher forcing and rollout-based training support iterative long-horizon prediction. Case studies on multiple test systems demonstrate accurate trajectory prediction, reliable stability discrimination, and effective adaptation across unseen scenarios.

eess.SY

Optimal location of small favourable regions for Robin eigenvalues with indefinite weights

We consider the positive principal eigenvalue of an elliptic problem with a Robin boundary condition and bang--bang indefinite weight $κ\mathbf 1_D-\mathbf 1_{Ω\setminus D}$, and ask where a favourable region $D$ of prescribed small volume $|D|=δ$ should be located. Put $\varepsilon=δ^{1/N}$ and $τ_δ=α_δ\varepsilon$. We prove that there is a finite threshold $τ_*=τ_*(N,κ)$, independent of the ambient domain, such that $δ^{2/N}Λ_δ(α_δ)\toΛ_{\mathbb H}(τ)$ whenever $τ_δ\toτ<\infty$. If $τ<τ_*$, optimal small regions concentrate at the boundary. If $τ>τ_*$, their concentration centres move to distances much larger than $\varepsilon$ from the boundary, and after recentring and rescaling the favourable sets converge in measure to the whole-space optimal ball. At $τ=τ_*$, the half-space problem admits a compact boundary optimiser as well as minimising sequences escaping to infinity. In the finer regime $τ_δ=τ_*+σ\varepsilon+o(\varepsilon)$, we determine the first-order competition between the boundary and interior configurations. Tangential symmetry of compact threshold optimisers reduces the geometric correction to mean curvature. In particular, every fixed finite Robin coefficient is asymptotically in the boundary regime. Numerical experiments illustrate the boundary--interior transition, the three cases in the first-order selection law, and the curvature-dependent boundary location; for $N=2$ and $κ=1$ they suggest a transition near $τ_*\approx3.2$.

math.AP

StepAudio 3 Realtime Technical Report

Realtime spoken interaction demands deep reasoning, prompt responses, and fluid turn-taking. We present StepAudio 3 Realtime, an audio-language foundation model organized around a continuous listen-converse-think-act loop. Deep Perception captures rich acoustic cues to interpret user intent, while Seamless Duplex models synchronized audio streams to handle pauses, backchannels, and interruptions naturally. Crucially, we resolve the tension between deep deliberation and latency via Think-While-Speaking, executing private reasoning in parallel with spoken delivery. In reasoning mode, StepAudio 3 reaches a 73.0 macro average on StepAudioChat. With Think-While-Speaking, it achieves dialogue and reasoning performance comparable to dedicated reasoning models while speaking in real time. Furthermore, an integrated Voice Agent handles asynchronous tool execution without disrupting the dialogue flow. StepAudio 3 Realtime achieves top-tier performance across key dimensions: an exceptional 90.6 on the MMSU benchmark, 98.9 Overall on the Artificial Analysis Full-Duplex Bench, and a 56.0% macro task-success rate on $τ$-Voice.

cs.SD

Intrinsic {Right} $q$-Radial Vector Derivatives and Localized Fischer Decompositions on Radial Algebras

We define intrinsic right \(q\)-radial vector derivatives on radial algebras of abstract vector variables. For a finite parameter set \(Y\), a scalar Jackson calculus on \(x^2\) and the mixed anticommutators \(\{x,y_i\}\) is combined with the exterior decomposition relative to \(x\). Relabelling covariance and compatibility under inclusions of parameter sets give a direct-limit operator on arbitrary radial algebras. {Under the dimension specialization \(Q=q^M\), its classical limit agrees in every \(M\)-dimensional Clifford realization with the standard right Clifford derivative; for an exterior blade one obtains \([x\wedge y_I]\partial_x=(-1)^{|I|}(M-|I|)y_I\).} Two Fischer-type constructions are considered. The anticommutator with exterior creation is triangular and admits an explicit Green inverse after localization at its diagonal factors. {For right \(q\)-monogenic elements, the derivative is paired with right multiplication \(R_xF=Fx\). The homogeneous operators \(\partial^{Y,\mathrm R}_{x,q}R_x\) have generically nonzero determinants, which yields determinant-localized Fischer decompositions and explicit projectors.} The one- and two-vector determinants are computed explicitly. {After \(Q=q^M\), the two-vector determinant is nonzero for every \(0<q<1\) and every positive integer \(M\). For arbitrary finite support, a support filtration factors the determinant into exact-support terms; in degree zero, support rank \(p\) contributes the factor \([m-p]_q+p\).}

math.CV

StepAudio 3 Gen Technical Report

We introduce StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types within a unified framework. At its core, StepAudio 3 Gen is a discrete autoregressive generator that models audio directly over residual vector quantization (RVQ) tokens, departing from the diffusion Transformer-based continuous generation paradigm prevalent in recent general audio models. Its StepAudio Tokenizer represents general audio at 12.5 Hz in a shared $16 \times 2048$ residual code space, jointly quantizing semantic and waveform-level acoustic features so that each code layer preserves both types of information. For generation, the backbone predicts the first codebook along the time axis using autoregressive modeling, while a lightweight causal Transformer completes the remaining fifteen codebooks along the codebook axis. Our study further identifies three key design principles: (1) interference-aware progressive pretraining for acquiring audio capabilities while preserving the textual abilities of the large language model, (2) RVQ Adaptor for effectively incorporating multi-codebook acoustic representations, and (3) discrete autoregressive modeling over a shared representation across general audio domains. With progressive pretraining, multi-task instruction training, and supervised fine-tuning, StepAudio 3 Gen achieves state-of-the-art performance on both TTS and voice design, while retaining strong generation capabilities across speech, vocals, sound effects, and music. Audio samples are available at https://stepaudiollm.github.io/step-audio-3-gen/.

cs.SD

When Ad Networks Misbehave: Understanding Risks of Semi-Drive-By Splash Ads

We investigate the mobile splash ads ecosystem, i.e., full-screen advertisements shown at app launch, where monetization relies on interaction signals that are difficult to verify end-to-end. This setting is especially sensitive because incidental touches and sensor-driven callbacks are common yet easy to misattribute as engagement. Prior work has largely framed mobile ad fraud as a publisher-side problem, while some studies attribute fraudulent operations to embedded ad libraries. Yet an important risk remains underexplored: ad SDKs control how interaction signals are interpreted, measured, and reported, creating an opportunity to reinterpret ambiguous user or device signals as valid advertising interactions. We uncover a previously less-known form of fraud at the ad-network layer in which splash ads are triggered not by intentional user actions but by incidental or indirect interactions, which we term semi-drive-by splash ads. By translating non-ad interactions into billable engagement events, ad networks can inflate performance metrics, overcharge advertisers, and erode user trust. To expose this behavior in the wild, we design AdHive, an automated honeypot-like analysis framework that induces evasive splash-ad delivery and landing behaviors under realistic device conditions. AdHive reproduces human-like activity through LLM-generated usage traces and sensor dynamics, enabling execution paths that remain hidden in conventional analysis environments. Our large-scale measurement across thousands of popular Android applications shows that semi-drive-by splash ads are widespread and are often triggered by subtle signals such as minor sensor variations. We further confirm real-world impact by working with one of China's largest advertisers, identifying multiple ad networks engaging in this fraud and leading to enforced repayments of about 4 million Yuan (approximately US$600,000).

cs.CR

MedQA-MM: Shortcuts Behind Medical Visual Reasoning

A benchmark score credits final answers, but not the route by which an item can be answered. In medical multimodal multiple-choice questions (MCQs), this distinction matters because a correct answer can be supported by the intended image finding or by benchmark-preserved cues in the wording of answers, non-visual clinical text, visible image text, artificial annotations, or device/context artifacts. We call the resulting score-level overinterpretation reasoning inflation. Here, a route is an observable input path that can support answer selection, not a claim about the model's hidden cognition. Across six medical multimodal MCQ datasets, we separate candidate cues from behavioral evidence through prompt- and image-side audits, modality ablations, and matched repairs that preserve the medical target and answer key. In a 13-configuration open-model panel, full-input accuracy is 62.63%, while text-only and options-only settings achieve 53.96% and 29.71%, respectively. Removing length-gap, absolute/conspicuous, and spatial/prepositional cues lowers accuracy by 6.58, 3.50, and 4.77 percentage points. We also construct MedQA-MM, a 1,000-item shortcut-mitigated subset, where text-only and options-only accuracy fall to 5.21% and 12.33%. This does not imply that models never use images; it shows that medical image-reasoning claims require route-level evidence.

cs.CV

Direct Realization of Near-Ideal Carbyne in Ultrathin Boron Nitride Nanotubes

Carbyne, the sp-hybridized one-dimensional allotrope of carbon, is predicted to be the stiffest known material, with electronic and optical properties set by a single structural parameter, the bond length alternation. However, its intrinsic properties have never been measured: chains synthesized through molecular chemistry carry endgroup and finite-length perturbations that persist even in the longest molecules available, while chains grown inside carbon nanotubes strongly couple to the host, which renormalizes their vibrational frequency by up to 110 cm$^{-1}$ in a diameter-dependent manner. Here, we show that encapsulating and thermally converting hydrogen-capped polyynes inside ultrathin boron nitride nanotubes, structural analogues of carbon nanotubes but electrically insulating, yields carbyne chains in a near-ideal regime, where endgroup, finite length, and host-guest perturbations are reduced to secondary effects. Statistical Raman spectroscopy across 245 locations returns a vibrational frequency distribution an order of magnitude narrower than in carbon nanotubes, an anharmonicity consistent with the universal law for carbyne-like materials, and a bond length alternation matching correlated calculations for the free chain. No photoluminescence is detected, despite the transparent host, as expected for the dipole-forbidden emission of an unperturbed carbyne chain. Boron nitride nanotubes give experimental access to carbyne in its near-ideal form.

cond-mat.mes-hall

GeoContext: One Context Ladder, Two Failure Modes in Vision-Language Geolocation: Flat Reliance on User-Provided Location Context and False Confirmation of Location Claims

Visual geolocation benchmarks typically ask a model where an image was captured without accounting for the location context that users often provide. We introduce GeoContext, a resource supporting two complementary tasks: GeoHint, open-ended localization given a true but coarse location hint, and GeoVerify, binary verification of whether an image was taken within 150 m of a claimed place. GeoContext constructs a context ladder by stratifying nearby reference points according to distance and referenceability, allowing the image to remain fixed while the supplied context varies. The benchmark covers 109 sites in 30 cities and evaluates five vision-language models using 21,933 GeoHint responses and 6,270 GeoVerify responses. Our evaluation reveals three main patterns. First, hint repetition varies by only 1.5 percentage points across referenceability tiers and by less than 3 points across distance bands, while the resulting localization error increases steadily with hint distance. Second, behavior depends strongly on no-context performance: at sites with low no-context accuracy, the median ratio between localization error and hint distance is approximately 1.00, whereas at higher-accuracy sites it ranges from 0.24 to 0.69. After correcting for bias introduced by the site grouping procedure, only one of the five models retains a negative accuracy estimate when given a nearby hint. Third, in GeoVerify, no model reaches d' = 1 for decoys immediately beyond the 150 m tolerance. Model rankings also change when sensitivity is separated from response bias, and 83.8% of false acceptances are reported with confidence of at least 0.8. We release the benchmark, construction pipeline, audit decisions, and scoring code.

cs.CV

Comparables XAI: Faithful Example-based AI Explanations with Counterfactual Trace Adjustments

Explaining with examples is an intuitive way to justify AI decisions. However, it is challenging to understand how a decision value should change relative to the examples with many features differing by large amounts. We draw from real estate valuation that uses Comparables-examples with known values for comparison. Estimates are made more accurate by hypothetically adjusting the attributes of each Comparable and correspondingly changing the value based on factors. We propose Comparables XAI for relatable example-based explanations of AI with Trace adjustments that trace counterfactual changes from each Comparable to the Subject, one attribute at a time, monotonically along the AI feature space. In modelling and user studies, Trace-adjusted Comparables achieved the highest XAI faithfulness and precision, user accuracy, and narrowest uncertainty bounds compared to linear regression, linearly adjusted Comparables, or unadjusted Comparables. This work contributes a new analytical basis for using example-based explanations to improve user understanding of AI decisions.

cs.HC

Spectral properties of the magnetic Robin Laplacian on curvilinear half-plane

We analyse magnetic Robin Laplacian in curvilinear half-plane with a smooth boundary. It is well known that the spectrum of the Robin Laplacian is unstable with respect to boundary deformations. This means that if the boundary is a straight line then the spectrum of the Robin Laplacian is purely essential. From the other hand, the perturbation of the boundary produces eigenvalues below the essential spectrum. In this paper, the Robin-Laplace operator with a compactly supported magnetic field is considered. We prove that the spectrum of the magnetic Robin Laplacian is stable under small and local deformations of the boundary.

math.SP