Searcharxiv⌕ Search

arXiv subjects

Han Wang

Publications and source records attributed to Han Wang.

At least 109 records · Page 6Linked to original sources

Cracking Gravitational Wave Multiple Ringdown Modes in Space

Ringdown signals from perturbed black holes (BHs) offer a clean window into BH spacetime, strong-field gravity, and fundamental physics. Presently the quasi-normal modes of stellar-mass BH ringdowns have been successfully extracted in the ground-based gravitational wave (GW) observations. Looking ahead, the future space-borne observatories will listen to the ringdowns from massive BH binary coalescences more loudly and resolve multiple modes to unprecedented precision, which calls for efficient approaches to mitigate the sharply increasing computational burden. We develop a practical ringdown analysis pipeline for space-borne detectors by implementing FIREFLY, a novel acceleration algorithm validated in ground-based detectors, and for the first time demonstrate its compatibility and effectiveness with the time-delay interferometry (TDI) observables. With high fidelity, we achieve a $\sim 200$-fold speedup for a simulated ringdown signal including six modes, providing a viable and scalable route for multi-mode ringdown analysis in the space context. This new approach has sound statistical interpretation and is extensible to other GW sources in band.

gr-qc↗

CreativeGame:Toward Mechanic-Aware Creative Game Generation

Large language models can generate plausible game code, but turning this capability into \emph{iterative creative improvement} remains difficult. In practice, single-shot generation often produces brittle runtime behavior, weak accumulation of experience across versions, and creativity scores that are too subjective to serve as reliable optimization signals. A further limitation is that mechanics are frequently treated only as post-hoc descriptions, rather than as explicit objects that can be planned, tracked, preserved, and evaluated during generation. This report presents \textbf{CreativeGame}, a multi-agent system for iterative HTML5 game generation that addresses these issues through four coupled ideas: a proxy reward centered on programmatic signals rather than pure LLM judgment; lineage-scoped memory for cross-version experience accumulation; runtime validation integrated into both repair and reward; and a mechanic-guided planning loop in which retrieved mechanic knowledge is converted into an explicit mechanic plan before code generation begins. The goal is not merely to produce a playable artifact in one step, but to support interpretable version-to-version evolution. The current system contains 71 stored lineages, 88 saved nodes, and a 774-entry global mechanic archive, implemented in 6{,}181 lines of Python together with inspection and visualization tooling. The system is therefore substantial enough to support architectural analysis, reward inspection, and real lineage-level case studies rather than only prompt-level demos. A real 4-generation lineage shows that mechanic-level innovation can emerge in later versions and can be inspected directly through version-to-version records. The central contribution is therefore not only game generation, but a concrete pipeline for observing progressive evolution through explicit mechanic change.

cs.AI↗

MERRIN: A Benchmark for Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments

Motivated by the underspecified, multi-hop nature of search queries and the multimodal, heterogeneous, and often conflicting nature of real-world web results, we introduce MERRIN (Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments), a human-annotated benchmark for evaluating search-augmented agents. MERRIN measures AI agents' ability to identify relevant modalities, retrieve multimodal evidence, and perform multi-hop reasoning over noisy web sources. It differs from prior work in three important aspects: (1) using natural language queries without explicit modality cues, (2) incorporating underexplored modalities such as video and audio, and (3) requiring the retrieval of complex, often noisy or conflicting multimodal evidence during web search. We evaluate diverse search agents powered by ten models, including strong closed-source models (e.g., GPT-5.4-mini, Gemini 3/3.1 Flash/Pro) and open-weight models (Qwen3-4B/30B/235B), across three search settings (no search, native search, and agentic search). Our results show that MERRIN is highly challenging: the average accuracy across all agents is 22.3%, with the best-performing agent reaching only 40.1%. We further observe that while stronger agents like Gemini Deep Research achieve higher performance, gains are modest due to over-exploration; they take more steps and use more tools, but are often distracted by conflicting or partially relevant web content, leading to incorrect answers. Compared to humans, these agents consume more resources yet achieve lower accuracy, largely due to inefficient source selection and an overreliance on text modalities. These findings highlight the need for search agents capable of robust search and reasoning across diverse modalities in noisy web environments, making MERRIN a valuable testbed for evaluating such capabilities.

cs.CL↗

Scheduling the Unschedulable: Taming Black-Box LLM Inference at Scale

When output token counts can be predicted at submission time (Gan et al., 2026), client-side scheduling against a black-box LLM API becomes semi-clairvoyant: decisions condition on coarse token priors even though the provider's internals remain hidden. We decompose this boundary problem into three separable concerns: allocation (inter-class share via adaptive DRR), ordering (intra-class sequencing with feasible-set scoring), and overload control (explicit admit/defer/reject on a cost ladder). An information ladder experiment shows that coarse magnitude priors -- not class labels alone -- are the practical threshold for useful client control; removing magnitude inflates short-request P95 by up to $5.8\times$ and degrades deadline satisfaction. Under balanced / high congestion the full stack achieves 100% completion, 100% deadline satisfaction, and useful goodput of $4.2 \pm 1.6$ SLO-meeting requests/s with short P95 within tens of milliseconds of quota-tiered isolation. A predictor-noise sweep confirms graceful degradation under up to 60% multiplicative error. Heavy-dominated regimes separate policies on completion, tail, and interpretable shedding. We further compare short-priority allocation (biased toward interactive traffic) with Fair Queuing (round-robin across classes): Fair Queuing achieves +32% short-request P90 improvement over FIFO with only +17% long-request overhead, versus Short-Priority's +27% / +116% trade-off -- demonstrating that the allocation layer accommodates different fairness objectives without changing the remaining stack. We contribute the three-layer client-side decomposition, controlled evaluation of joint metrics across regimes, allocation-policy alternatives, and overload-policy evidence linking cost-ladder shedding to the stated service objective.

cs.DC↗

Model-Free Power System Stability Enhancement with Dissipativity-Based Neural Control

The integration of converter-interfaced generation introduces new transient stability challenges to modern power systems. Classical Lyapunov- and scalable passivity-based approaches typically rely on restrictive assumptions, and finding storage functions for large grids is generally considered intractable. Furthermore, most methods require an accurate grid dynamics model. To address these challenges, we propose a model-free, nonlinear, and dissipativity-based controller which, when applied to grid-connected virtual synchronous generators (VSGs), enhances power system transient stability. Using input-state data, we train neural networks to learn dissipativity-characterizing matrices that yield stabilizing controllers. Furthermore, we incorporate cost function shaping to improve the performance with respect to the user-specified objectives. Numerical results on a modified, all-VSG Kundur two-area power system validate the effectiveness of the proposed approach.

eess.SY↗

POEMetric: The Last Stanza of Humanity

Large Language Models (LLMs) can compose poetry, but how far are they from human poets? In this paper, we introduce POEMetric, the first comprehensive framework for poetry evaluation, examining 1) basic instruction-following abilities in generating poems according to a certain form and theme, 2) advanced abilities of showing creativity, lexical diversity, and idiosyncrasy, evoking emotional resonance, and using imagery and literary devices, and 3) general appraisal of the overall poem quality and estimation of authorship. We curated a human poem dataset - 203 English poems of 7 fixed forms annotated with meter, rhyme patterns and themes - and experimented with 30 LLMs for poetry generation based on the same forms and themes of the human data, totaling 6,090 LLM poems. Based on POEMetric, we assessed the performance of both human poets and LLMs through rule-based evaluation and LLM-as-a-judge, whose results were validated by human experts. Results show that, though the top model achieved high form accuracy (4.26 out of 5.00, with Gemini-2.5-Pro as a judge; same below) and theme alignment (4.99), all models failed to reach the same level of advanced abilities as human poets, who achieved unparalleled creativity (4.02), idiosyncrasy (3.95), emotional resonance (4.06), and skillful use of imagery (4.49) and literary devices (4.67). Humans also defeated the best-performing LLM in overall poem quality (4.22 vs. 3.20). As such, poetry generation remains a formidable challenge for LLMs. Data and codes are released at https://github.com/Bingru-Li/POEMetric.

cs.CL↗

A Global Spacetime Optimization Approach to the Real-Space Time-Dependent Schrödinger Equation

The time-dependent Schrödinger equation (TDSE) in real space is fundamental to understanding the dynamics of many-electron quantum systems, with applications ranging from quantum chemistry to condensed matter physics and materials science. However, solving the TDSE for complex fermionic systems remains a significant challenge, particularly due to the need to capture the time-evolving many-body correlations, while the antisymmetric nature of fermionic wavefunctions complicates the function space in which these solutions must be represented. We propose a general-purpose neural network framework for solving the real-space TDSE, Fermionic Antisymmetric Spatio-Temporal Network, which treats time as an explicit input alongside spatial coordinates, enabling a unified spatiotemporal representation of complex, antisymmetric wavefunctions for fermionic systems. This approach formulates the TDSE as a global optimization problem, avoiding step-by-step propagation and supporting highly parallelizable training. The method is demonstrated on five benchmark problems, achieving excellent agreement with reference solutions across all cases. These results demonstrate the method's accuracy and flexibility within the bound-state manifold across various dimensions and interaction regimes. While the current localized Ansatz inherently restricts the description of extensive ionization and continuum states, the method demonstrates the capability to stably simulate coherent multi-electron dynamics over extended time windows. Our framework offers a highly expressive alternative to traditional basis-dependent or mean-field methods, opening new possibilities for ab initio simulations of time-dependent quantum systems, with applications in quantum dynamics, molecular control, and ultrafast spectroscopy.

quant-ph↗

The Stellar IMF and Dark Matter Halo of ESO0286: Constraints from Strong Lensing and Dynamics

The internal mass structure of elliptical galaxies offers critical insights into galaxy formation, yet disentangling stellar mass from dark matter and determining the stellar initial mass function (IMF) remains challenging. We present a detailed analysis of ESO0286-G022 ($z=0.0312$), a rare nearby strong-lens system with a fast-rotating elliptical galaxy, combining high-resolution Hubble Space Telescope (HST) imaging with VLT/MUSE integral-field stellar kinematics. We construct axisymmetric and triaxial Schwarzschild orbit-superposition models to reconstruct its intrinsic shape and mass distribution. Despite being a fast rotator, ESO0286 exhibits clear kinematic signatures of intrinsic triaxiality, characterized by rotation along both the major and minor axes, making it only the second such confirmed case. By incorporating the mass enclosed within the Einstein radius from strong lensing as a complementary constraint, we tightly anchor the total mass at large radii. This significantly reduces the uncertainty on the outer mass profile and orbital structure, demonstrating that only models with strong radial anisotropy beyond the IFU field of view are compatible with the data. In the inner regions, we robustly constrain an upper limit for the stellar mass around $r \sim 0.7$ kpc, ruling out an IMF more bottom-heavy than Kroupa, though a gentle gradient toward a slightly heavier central IMF is permitted. This aligns with recent dynamical studies of local massive early-type galaxies but contrasts with heavier IMFs reported for lenses at $z>0.1$. Our work demonstrates the power of combining lensing and dynamical modeling to resolve the detailed inner structure of massive galaxies.

astro-ph.GA↗

Fock space prethermalization and time-crystalline order on a quantum processor

Periodically driven quantum many-body systems exhibit a wide variety of exotic nonequilibrium phenomena and provide a promising pathway for quantum applications. A fundamental challenge for stabilizing and harnessing these highly entangled states of matter is system heating by energy absorption from the drive. Here, we propose and demonstrate a disorder-free mechanism, dubbed Fock space prethermalization (FSP), to suppress heating. This mechanism divides the Fock-space network into linearly many sparse sub-networks, thereby prolonging the thermalization timescale even for initial states at high energy densities. Using 72 superconducting qubits, we observe an FSP-based time-crystalline order that persists over 120 cycles for generic initial Fock states. The underlying kinetic constraint of approximately conserved domain wall (DW) numbers is identified by measuring site-resolved correlators. Further, we perform finite-size scaling analysis for DW and Fock-space dynamics by varying system sizes, which reveals size-independent regimes for FSP-thermalization crossover and links the dynamical behaviors to the eigenstructure of the Floquet unitary. Our work establishes FSP as a robust mechanism for breaking ergodicity, and paves the way for exploring novel nonequilibrium quantum matter and its applications.

quant-ph↗

Bilinear structures of the fourth-order lattice Gel'fand-Dikii equations

In this paper we derive bilinear forms and present their solutions in Casoratians for several fourth-order lattice Gel'fand-Dikii (lattice GD-4) equations. These equations were recently formulated from the direct linearization approach and exhibit the multidimensionally consistent property in multi-component form. Based on the obtained soliton solutions, we are able to extend these equations by introducing a parameter $δ$. These $δ$-extended lattice GD-4 type equations are still consistent around the cube, and their bilinear forms together with Casoratian solutions are provided.

nlin.SI↗

Lap2: Revisiting Laplace DP-SGD for High Dimensions via Majorization Theory

Differentially Private Stochastic Gradient Descent (DP-SGD) is a cornerstone technique for ensuring privacy in deep learning, widely used in both training from scratch and fine-tuning large-scale language models. While DP-SGD predominantly relies on the Gaussian mechanism, the Laplace mechanism remains underutilized due to its reliance on L1 norm clipping. This constraint severely limits its practicality in high-dimensional models because the L1 norm of an n-dimensional gradient can be up to sqrt(n) times larger than its L2 norm. As a result, the required noise scale grows significantly with model size, leading to poor utility or untrainable models. In this work, we introduce Lap2, a new solution that enables L2 clipping for Laplace DP-SGD while preserving strong privacy guarantees. We overcome the dimensionality-driven clipping barrier by computing coordinate-wise moment bounds and applying majorization theory to construct a tight, data-independent upper bound over the full model. By exploiting the Schur-convexity of the moment accountant function, we aggregate these bounds using a carefully designed majorization set that respects the L2 clipping constraint. This yields a multivariate privacy accountant that scales gracefully with model dimension and enables the use of thousands of moments. Empirical evaluations demonstrate that our approach significantly improves the performance of Laplace DP-SGD, achieving results comparable to or better than Gaussian DP-SGD under strong privacy constraints. For instance, fine-tuning RoBERTa-base (125M parameters) on SST-2 achieves 87.88% accuracy at epsilon=0.54, outperforming Gaussian (87.16%) and standard Laplace (48.97%) under the same budget.

cs.CR↗

SwinYNet: A Transformer-based Multi-Task Model for Accurate and Efficient FRB Search

In this study, we present a transformer-based multi-task model for Fast Radio Burst (FRB) detection, signal segmentation, and parameter estimation directly from time-frequency data, without requiring computationally expensive de-dispersion preprocessing. To overcome the scarcity of labeled observational data, we develop an FRB simulator and a rule-based automatic annotation pipeline, enabling training exclusively on simulated data. Evaluations on the FAST-FREX dataset show that our model achieves an F1 score of 97.8%, recall of 95.7%, and precision of 100%, outperforming both conventional tools (e.g., PRESTO, Heimdall) and recent AI-based baselines (e.g., RaSPDAM, DRAFTS) in both accuracy and inference speed. The model supports pixel-level signal segmentation and yields reliable estimates for dispersion measure (DM) and time of arrival (ToA). Large-scale blind searches on CRAFTS data further demonstrate robustness, with an average false positive rate of 0.28% and minimal human verification required. This search has already led to the identification of two pulsar candidates, both confirmed as known pulsars. Processing benchmarks indicate that the model enables real-time searches on a single consumer-grade GPU, making petabyte-scale blind searches feasible. The code is publicly available on GitHub, and the model can be easily integrated with existing tools to automate and streamline radio data analysis beyond FRB or pulsar searches.

astro-ph.GA↗

Greedy-based Value Representation for Optimal Coordination in Multi-agent Reinforcement Learning

Due to the representation limitation of the joint Q value function, multi-agent reinforcement learning methods with linear value decomposition (LVD) or monotonic value decomposition (MVD) suffer from relative overgeneralization. As a result, they can not ensure optimal consistency (i.e., the correspondence between individual greedy actions and the maximal true Q value). In this paper, we derive the expression of the joint Q value function of LVD and MVD. According to the expression, we draw a transition diagram, where each self-transition node (STN) is a possible convergence. To ensure optimal consistency, the optimal node is required to be the unique STN. Therefore, we propose the greedy-based value representation (GVR), which turns the optimal node into an STN via inferior target shaping and further eliminates the non-optimal STNs via superior experience replay. In addition, GVR achieves an adaptive trade-off between optimality and stability. Our method outperforms state-of-the-art baselines in experiments on various benchmarks. Theoretical proofs and empirical results on matrix games demonstrate that GVR ensures optimal consistency under sufficient exploration.

cs.MA↗

On The Fragility of Benchmark Contamination Detection in Reasoning Models

Leaderboards for LRMs have turned evaluation into a competition, incentivizing developers to optimize directly on benchmark suites. A shortcut to achieving higher rankings is to incorporate evaluation benchmarks into the training data, thereby yielding inflated performance, known as benchmark contamination. Surprisingly, our studies find that evading contamination detections for LRMs is alarmingly easy. We focus on the two scenarios where contamination may occur in practice: (I) when the base model evolves into LRM via SFT and RL, we find that contamination during SFT can be originally identified by contamination detection methods. Yet, even a brief GRPO training can markedly conceal contamination signals that most detection methods rely on. Further empirical experiments and theoretical analysis indicate that PPO style importance sampling and clipping objectives are the root cause of this detection concealment, indicating that a broad class of RL methods may inherently exhibit similar concealment capability; (II) when SFT contamination with CoT is applied to advanced LRMs as the final stage, most contamination detection methods perform near random guesses. Without exposure to non-members, contaminated LRMs would still have more confidence when responding to those unseen samples that share similar distributions to the training set, and thus, evade existing memorization-based detection methods. Together, our findings reveal the unique vulnerability of LRMs evaluations: Model developers could easily contaminate LRMs to achieve inflated leaderboards performance while leaving minimal traces of contamination, thereby strongly undermining the fairness of evaluation and threatening the integrity of public leaderboards. This underscores the urgent need for advanced contamination detection methods and trustworthy evaluation protocols tailored to LRMs.

cs.CR↗

How do Visual Attributes Influence Web Agents? A Comprehensive Evaluation of User Interface Design Factors

Web agents have demonstrated strong performance on a wide range of web-based tasks. However, existing research on the effect of environmental variation has mostly focused on robustness to adversarial attacks, with less attention to agents' preferences in benign scenarios. Although early studies have examined how textual attributes influence agent behavior, a systematic understanding of how visual attributes shape agent decision-making remains limited. To address this, we introduce VAF, a controlled evaluation pipeline for quantifying how webpage Visual Attribute Factors influence web-agent decision-making. Specifically, VAF consists of three stages: (i) variant generation, which ensures the variants share identical semantics as the original item while only differ in visual attributes; (ii) browsing interaction, where agents navigate the page via scrolling and clicking the interested item, mirroring how human users browse online; (iii) validating through both click action and reasoning from agents, which we use the Target Click Rate and Target Mention Rate to jointly evaluate the effect of visual attributes. By quantitatively measuring the decision-making difference between the original and variant, we identify which visual attributes influence agents' behavior most. Extensive experiments, across 8 variant families (48 variants total), 5 real-world websites (including shopping, travel, and news browsing), and 4 representative web agents, show that background color contrast, item size, position, and card clarity have a strong influence on agents' actions, whereas font styling, text color, and item image clarity exhibit minor effects.

cs.AI↗

TDCOSMO XXIV. First spatially resolved kinematics of the lens galaxy obtained using JWST-NIRSpec to improve time-delay cosmography

Spatially resolved stellar kinematics has become a key ingredient in time-delay cosmography to break the mass-sheet degeneracy in the mass profile and in turn provide a precise constraint on the Hubble constant and other cosmological parameters. In this paper, we present the first measurements of 2D resolved stellar kinematics for the lens galaxy in the quadruply lensed quasar system RXJ1131$-$1231 using integral field spectroscopy from JWST's Near-Infrared Spectrograph (NIRSpec), marking the first such measurement conducted with JWST. In extracting robust kinematic measurements from this first-of-its-kind dataset, we have made methodological improvements both in the data reduction and kinematic extraction. In our kinematic extraction procedure, we performed joint modeling of the lens galaxy, the quasar, and its host galaxy's contributions in the spectra to deblend the lens galaxy component and robustly constrain its stellar kinematics. Our improved methodological frameworks are released as software pipelines for future use: squirrel, for extracting stellar kinematics, and RegalJumper, for JWST-NIRSpec data reduction. We incorporated additional artifact cleaning beyond the standard JWST pipeline. We compared our measured stellar kinematics from the JWST NIRSpec with previously obtained ground-based measurements from the Keck Cosmic Web Imager integral field unit and find that the two datasets are statistically consistent at a $\sim$1.1$σ$ confidence level. Our measured kinematics will be used in a future study to improve the precision of the Hubble constant measurement.

astro-ph.GA↗

OmniOCR: Generalist OCR for Ethnic Minority Languages

Optical character recognition (OCR) has advanced rapidly with deep learning and multimodal models, yet most methods focus on well-resourced scripts such as Latin and Chinese. Ethnic minority languages remain underexplored due to complex writing systems, scarce annotations, and diverse historical and modern forms, making generalization in low-resource or zero-shot settings challenging. To address these challenges, we present OmniOCR, a universal framework for ethnic minority scripts. OmniOCR introduces Dynamic Low-Rank Adaptation (Dynamic LoRA) to allocate model capacity across layers and scripts, enabling effective adaptation while preserving knowledge.A sparsity regularization prunes redundant updates, ensuring compact and efficient adaptation without extra inference cost. Evaluations on TibetanMNIST, Shui, ancient Yi, and Dongba show that OmniOCR outperforms zero-shot foundation models and standard post training, achieving state-of-the-art accuracy with superior parameter efficiency, and compared with the state-of-the-art baseline models, it improves accuracy by 39%-66% on these four datasets. Code: https://github.com/AIGeeksGroup/OmniOCR.

cs.CV↗

Teleportation transition of surface codes on a superconducting quantum processor

The topological surface code is a leading candidate for harnessing long-range entanglement to protect logical quantum information against errors, and teleportation of logical states is desirable for robust quantum information processing. Nevertheless, scaling up the surface code in quantum teleportation poses a formidable challenge to experiment. Here on a superconducting quantum processor with 125 qubits, we demonstrate the robust teleportation of topological rotated surface code prepared by a linear-depth unitary circuit, with code distances up to 7. We obtain the teleportation phase diagram by tuning the local entangling gates uniformly across a finite threshold. Furthermore, we show that the entangling threshold can be boosted by coherent qubit rotations that inject magic resources beyond the Clifford regime, restoring the duality symmetry of the topological phase, which serves as a guiding principle to minimize the entanglement resource. Our results shed light on simulating and leveraging topological quantum matter on quantum devices, and pave the way to the ultimate goal of distributed fault tolerant quantum computation.

quant-ph↗