SearcharxivSearch

arXiv subjects

Qiang Wu

Publications and source records attributed to Qiang Wu.

At least 19 recordsLinked to original sources

Joint parameters estimation in cubic tensor model

We study joint parameter estimation from a single observation in high-dimensional Gibbs measures with cubic tensor interactions, motivated by dense ERGMs, arithmetic-progression models, and inhomogeneous random hypergraphs. Focusing on the maximum pseudolikelihood estimator, we give checkable conditions for joint consistency and asymptotic ill-conditioning. For the edge-triangle ERGM, pseudolikelihood is ill-conditioned in the ferromagnetic regime with nonnegative field, but consistent in a sufficiently strong antiferromagnetic regime. For the edge-three-star ERGM, it is ill-conditioned for all inverse temperatures and external fields. We also study consistency for arithmetic-progression, and inhomogeneous hypergraph models. Our proofs develop nonlinear large-deviation and mean-field approximation tools for cubic tensor Gibbs measures, which have scope for broad applications.

math.ST

The Convergence and Error Analysis of Coordinate Descent Methods with Compression for Full Configuration Interaction

We study the effect of compression in Coordinate Descent Full Configuration Interaction (CDFCI) within an unconstrained optimization formulation of the full configuration interaction ground-state problem. Under suitable local assumptions, we prove that the compressed iteration converges linearly to the solution of an associated restricted problem. We also characterize the convergence point of the compressed algorithm. Under an additional exponential decay assumption on the target eigenvector, we show that the resulting eigenvalue error is of order $\tau^2$, where $\tau$ denotes the compression threshold. Numerical results support the analysis.

math.NA

An adaptive and evolvable deep reinforcement learning framework for weather prediction

No single AI weather model excels at all variables, pressure levels, and lead times. Rather than building yet another architecture, we reframe the forecasting problem as one of coordination. Here we present Feitian Adaptive Ensemble Weather (FTAE-Weather), a lightweight framework that learns, through deep reinforcement learning, when and where to trust each member of an open pool of pretrained forecasters. A tactical Weight-Agent reads the current atmospheric state and assigns variable- and horizon-specific fusion weights, while a strategic Evolve-Agent periodically prunes underperforming models and absorbs newly released ones. Asynchronous prediction caching keeps training cost independent of the slowest constituent model. Adding fewer than 0.01 percent extra parameters, FTAE-Weather reduces RMSE by from 17.2 percent to 78.3 percent over the best individual model in 10 atmospheric variables and outperforms conventional ensemble baselines across lead times from 72 to 360 hours. The framework thus converts a growing, fragmented inventory of specialist models into a single prediction system that strengthens as the field of AI weather forecasting releases new architectures-turning model diversity from a coordination challenge into a compounding scientific advantage.

physics.ao-ph

Reward as An Agent for Embodied World Models

While RL has become a promising tool for refining world models, existing methods largely rely on conservative rollouts near the training distribution, limiting exploration, behavioral diversity, and richer dynamic discovery. In this work, we challenge this conservative paradigm. We argue that the core limitation is not exploration itself, but the lack of reliable verification strategies to support broader exploration. Without reliable verification, expanded exploration becomes highly susceptible to reward hacking, where policies exploit imperfect rewards without achieving genuine improvement. To evaluate this motivation, we instantiate our method in embodied world models, where physical plausibility, and task completion provide a rigorous testbed for scalable RL under complex dynamics. On the verification side, we introduce Reward as an Agent, an agentic reward framework that actively evaluates generated behaviors to provide robust reward signals and mitigate reward hacking under distribution shifts. On the exploration side, we introduce Dynamic-Aware Rollout Diversification through DynDiff-GRPO, which explicitly expands action-space exploration to diversify trajectories, broaden state-action coverage, and encourage richer embodied behaviors beyond conservative rollout regimes. By unifying Reward as an Agent with DynDiff-GRPO, we enable RL on a more reliable reward foundation with substantially diversified sampling, effectively mitigating reward hacking while yielding significant accuracy gains across multiple open-source world models, thereby demonstrating that broader exploration can scale successfully when grounded in robust verification.

cs.AI

Does DESI prefer Damped Oscillating Dark Energy over Cosmological constant?

We investigate a dark-energy equation of state governed by a damped harmonic oscillator equation, admitting underdamped, critically damped, and overdamped solutions. Confronting the model with Planck CMB distance priors, DESI BAO, BBN, cosmic chronometers, and three Type~Ia supernova compilations, we find that the data select an underdamped solution yielding $H_0 = 70.9 \pm 1.1$ km/s/Mpc with DES-Dovekie and $H_0 = 72.0^{+1.4}_{-2.1}$ km/s/Mpc with Union3, without any local $H_0$ prior. These higher values of $H_0$ arise along the $\Omega_{\rm m}$--$H_0$ degeneracy direction while the sound horizon remains nearly unchanged at $r_{\rm d} \simeq 145$~Mpc, indicating that the enhancement of the late-time expansion rate is a geometrical effect that does not address the early-time calibration of $r_{\rm d}$. In contrast, the Pantheon+ compilation selects a near-critically damped solution with a prior-limited positive $w_0$ and $H_0 = 66.23 \pm 0.85$ km/s/Mpc, highlighting the sensitivity of the model to the low-redshift distance information encoded in the different supernova compilations. The Bayesian evidence relative to $\Lambda$CDM is inconclusive for the DES-Dovekie and Union3 combinations, whereas Pantheon+ shows a strong preference for the damped-oscillator model, driven by the departure from $w=-1$ at $z\lesssim0.1$.

astro-ph.CO

OverFlowLight: Real-Time Gridlock Prevention and Traffic Signal Optimization for Urban Intersections

Queue overflow, a severe consequence of urban traffic congestion, occurs when vehicle queues exceed intersection capacity, obstructing upstream traffic and triggering cascading gridlocks. Prevailing traffic signal control (TSC) algorithms, primarily optimized for throughput, often fail to address overflow during peak hours, exacerbating congestion and creating safety hazards. We propose OverFlowLight, a real-time framework designed to preemptively resolve overflow and enhance overall TSC performance. It first introduces a mechanism to accurately detect overflow in real-time by leveraging multi-modal sensing from cameras and radars. Upon detection, it dynamically generates and inserts dedicated overflow phases into the signal cycle to clear the blocking queues. This is orchestrated by a hybrid control design that combines rapid rule-based overflow intervention with controller back ends such as reinforcement learning (RL) for longer-horizon efficiency. We conducted extensive real-world deployments of OverFlowLight across 43 intersections in three major cities. The framework demonstrates seamless integration with existing RL-based TSC agents, highlighting its modularity and practical applicability. Empirical results show that OverFlowLight reduces overflow incidents by 60.4% and increases network throughput by 18.2% compared to deployed baselines. Furthermore, it substantially diminishes the need for manual intervention common with expert-tuned signal plans. This work presents the first practical, scalable, and data-driven framework for actively preventing traffic gridlock, offering a crucial component for building resilient and efficient urban transportation systems. Our demonstration videos, codes and datasets are available at the anonymous URL, https://anonymous.4open.science/r/OverFlowLight-FBF9.

cs.LG

Constraints on Schwarzschild Black Hole in a Generalized Dehnen-Type $(1,4,\gamma)$ Dark Matter Halo via the S2 Star Orbit around Sgr A$^\star$

The distribution of dark matter (DM) halo around supermassive black holes (BHs) may leave observable imprints on stellar dynamics near galactic centers. Motivated by this, we investigate the orbital motion of the S2 star in the spacetime of a recently derived generalized Schwarzschild BH solution embedded in a Dehnen-type $(1,4,\gamma)$ DM halo, considering it as a possible model for Sgr A$^{\star}$ at the center of the Milky Way. Unlike previous studies restricted to specific values of the halo parameter $\gamma$, the present solution describes the fully generalized case with arbitrary $\gamma$. We derive the corresponding equations of motion and obtain the associated perihelion shift over one orbital period. Using observational data of the S2 star, we constrain the parameters of the Schwarzschild--Dehnen BH-DM system through a Markov Chain Monte Carlo (MCMC) analysis. Our results yield the best-fit values $\gamma = 1.18^{+1.03}_{-0.81}$ $(1.23^{+1.01}_{-0.85})$, $\rho_s = 0.37^{+0.42}_{-0.29}$ $(0.31^{+0.44}_{-0.26})$, and $r_s = 0.05^{+0.05}_{-0.03}$ $(0.14^{+0.18}_{-0.10})$ for observational data of Do et al.~\cite{Do19} and Gillessen et al.~\cite{Gillessen17ApJ}, respectively. We further obtain the corresponding 95\% confidence upper bounds: $\gamma < 2.66$ $(2.67)$, $\rho_s < 0.93$ $(0.92)$, and $r_s < 0.16$ $(0.52)$. These results demonstrate that precise stellar orbit measurements can provide meaningful constraints on the DM halo distributions surrounding supermassive BHs and may offer insights into the DM environment of Sgr A$^{\star}$ at the center of the Milky Way.

gr-qc

AccelCIM: Systematic Dataflow Exploration for SRAM Compute-in-Memory Accelerator

SRAM-based compute-in-memory (CIM) offers high computational density and energy efficiency for deep neural network (DNN) accelerators, but its limited capacity causes on/off-chip data movement overhead for large DNN models. Existing CIM accelerator studies typically assume that DNN models fit entirely on-chip, leaving efficient dataflow design largely untapped. This paper introduces AccelCIM, a systematic dataflow exploration framework for SRAM CIM accelerator, which addresses two key limitations of prior work. (1) It formulates a systematic dataflow design space spanning CIM macro configurations and macro-array organizations. (2) It introduces rigorous design evaluation using cycle-accurate architectural simulation and post-layout PPA analysis. We conduct an extensive design space exploration and apply AccelCIM to representative LLM applications, providing practical insights for the principled design of CIM accelerators.

cs.AR

A Full-Stack Performance Evaluation Infrastructure for 3D-DRAM-based LLM Accelerators

Large language models (LLMs) exhibit memory-intensive behavior during decoding, making it a key bottleneck in LLM inference. To accelerate decoding execution, hybrid-bonding-based 3D-DRAM has been adopted in LLM accelerators. While this emerging technology provides strong performance gains over existing hardware, current 3D-DRAM accelerators (3D-Accelerators) rely on closed-source evaluation tools, limiting access to publicly available performance analysis methods. Moreover, existing designs are highly customized for specific scenarios, lacking a general and reusable full-stack modeling for 3D-Accelerators across diverse usecases. To bridge this fundamental gap, we present ATLAS, the first silicon-proven Architectural Three-dimesional-DRAM-based LLM Accelerator Simulation framework. Built on commercially deployed multi-layer 3D-DRAM technology, ATLAS introduces unified abstractions for both 3D-Accelerator system architecture and programming primitives to support arbitrary LLM inference scenarios. Validation against real silicon shows that ATLAS achieves $\le$8.57% simulation error and 97.26-99.96\% correlation with measured performance. Through design space exploration with ATLAS, we demonstrate its ability to guide architecture design and distill key takeaways for both 3D-DRAM memory system and 3D-Accelerator microarchitecture across scenarios. ATLAS will be open-sourced upon publication, enabling further research on 3D-Accelerators.

cs.AR

Hardware-Software Co-design for 3D-DRAM-based LLM Serving Accelerator

Large language models (LLMs) have been widely deployed for online generative services, where numerous LLM instances jointly handle workloads with fluctuating request arrival rates and variable request lengths. To efficiently execute coexisting compute-intensive and memory-intensive operators, near-memory processing (NMP) based computing paradigm has been extensively proposed. However, existing NMP designs adopt coarse-grained KV cache management and inflexible attention execution flow. Such limitations hinder these proposals from efficiently handling \textit{highly dynamic} LLM serving workloads, limiting their ability to accelerate LLM serving. To tackle these problems, we propose Helios, a Hybrid-bonding-based \uline{L}LM \uline{S}erving accelerator. Helios aims to bridge the fundamental gap between the dynamic nature of KV cache management in LLM serving and the distributed, non-uniform memory abstraction among NMP processing engines (PEs). To this end, we design both the intra-PE execution flow and the inter-PE communication primitives for distributed tiled attention execution. We further propose \textit{spatially-aware} KV cache allocation mechanism to balance the attention workload distribution while minimizing the inter-PE data transfer overhead. Compared with existing GPU/NMP designs, Helios achieves 3.25 times (geomean) speedup and 3.36 times (geomean) better energy efficiency, along with up to 72%/76% P50/P99 time-between-tokens degradation.

cs.AR

Gravitational wave signatures from periodic orbits around a Schwarzschild-Bertotti-Robinson black hole

In this paper, we investigate periodic bound orbits and gravitational wave (GW) emission in the Schwarzschild-Bertotti-Robinson (Schwarzschild-BR) spacetime-an exact electrovacuum solution describing a static black hole (BH) immersed in a uniform magnetic field. We explore how the background magnetic field qualitatively alters the BH's gravitational dynamics, affecting timelike geodesics such as the marginally bound orbit (MBO) and the innermost stable circular orbit (ISCO). We then analyze periodic bound orbits using the frequency ratio ${\omega_{\varphi}}/{\omega_{r}}$, which characterizes the orbits by their azimuthal and radial motions. Based on the numerical kludge method we further compute the gravitational waveforms emitted from periodic orbits around a supermassive Schwarzschild-BR BH. We show that the background magnetic field significantly changes orbital frequencies, resonance conditions, zoom-whirl structures, and the resulting waveforms. Finally, we examine the frequency spectra in the mHz range and the detectability of these GW signals by computing the characteristic strain via a discrete Fourier transform on the time-domain waveforms, comparing the results with the sensitivity curves of space-based GW detectors such as LISA, Taiji, and TianQin. Our results show that intrinsically magnetic fields modify spacetime and leave observable imprints on extreme mass-ratio inspiral GWs, which may be tested by future observations.

gr-qc

Signatures of Quantum-Corrected Black Holes in Gravitational Waves from Periodic Orbits

We investigate gravitational wave emission from periodic timelike orbits of a test particle around a loop quantum gravity-inspired Schwarzschild black hole. The spacetime is characterised by a holonomy-correction parameter that modifies the radial metric component while preserving asymptotic flatness and the classical location of the horizon. The bound geodesics are systematically classified using the zoom--whirl representation labelled by three integers $(z,w,v)$. Gravitational waveforms are computed within a numerical framework that combines exact geodesic motion with the quadrupole approximation, which is suitable for extreme mass ratio inspirals. We demonstrate that the quantum corrections lead to distinct phase shifts, amplitude variations, and modifications to the harmonic structure of the waveforms, with increasingly complex features for orbits with larger zoom numbers. The corresponding frequency spectra and characteristic strain peak, which fall within the millihertz band, are within the sensitivity ranges of space-based detectors such as LISA, Taiji, and TianQin. For specific orbital configurations and values of the quantum-correction parameter, the characteristic strain exceeds the projected detector noise, indicating potential observability. Our results demonstrate that gravitational waves from periodic orbits provide a sensitive probe of quantum-corrected black hole spacetimes in the strong-field regime.

gr-qc

Spatial-spectral mapping for long-duration broadband terahertz pulse generation in on-chip waveguide arrays

Conventional approaches to terahertz (THz) pulse generation are restricted by the Fourier-transform limit, which hinders the creation of sources that combine long duration with broad bandwidth--a capability crucial for many spectroscopic and sensing applications. In this work, we overcome this challenge in the terahertz domain using an on-chip gradient waveguide array. The key is to spectrally disperse the pulse into spatially separated channels within a lithium niobate chip, effectively decoupling the design of temporal and spectral properties. We validate the source by distinguishing amino acid mixtures, demonstrating its tailored biosensing potential. This work establishes a novel mechanism for integrated THz generation, offering considerable promise for broadband spectroscopy and on-chip photonics.

physics.optics

CFLight: Enhancing Safety with Traffic Signal Control through Counterfactual Learning

Traffic accidents result in millions of injuries and fatalities globally, with a significant number occurring at intersections each year. Traffic Signal Control (TSC) is an effective strategy for enhancing safety at these urban junctures. Despite the growing popularity of Reinforcement Learning (RL) methods in optimizing TSC, these methods often prioritize driving efficiency over safety, thus failing to address the critical balance between these two aspects. Additionally, these methods usually need more interpretability. CounterFactual (CF) learning is a promising approach for various causal analysis fields. In this study, we introduce a novel framework to improve RL for safety aspects in TSC. This framework introduces a novel method based on CF learning to address the question: ``What if, when an unsafe event occurs, we backtrack to perform alternative actions, and will this unsafe event still occur in the subsequent period?'' To answer this question, we propose a new structure causal model to predict the result after executing different actions, and we propose a new CF module that integrates with additional ``X'' modules to promote safe RL practices. Our new algorithm, CFLight, which is derived from this framework, effectively tackles challenging safety events and significantly improves safety at intersections through a near-zero collision control strategy. Through extensive numerical experiments on both real-world and synthetic datasets, we demonstrate that CFLight reduces collisions and improves overall traffic performance compared to conventional RL methods and the recent safe RL model. Moreover, our method represents a generalized and safe framework for RL methods, opening possibilities for applications in other domains. The data and code are available in the github https://github.com/AdvancedAI-ComplexSystem/SmartCity/tree/main/CFLight.

cs.LG

VideoCoF: Unified Video Editing with Temporal Reasoner

Existing video editing methods face a critical trade-off: expert models offer precision but rely on task-specific priors like masks, hindering unification; conversely, unified temporal in-context learning models are mask-free but lack explicit spatial cues, leading to weak instruction-to-region mapping and imprecise localization. To resolve this conflict, we propose VideoCoF, a novel Chain-of-Frames approach inspired by Chain-of-Thought reasoning. VideoCoF enforces a ``see, reason, then edit" procedure by compelling the video diffusion model to first predict reasoning tokens (edit-region latents) before generating the target video tokens. This explicit reasoning step removes the need for user-provided masks while achieving precise instruction-to-region alignment and fine-grained video editing. Furthermore, we introduce a RoPE alignment strategy that leverages these reasoning tokens to ensure motion alignment and enable length extrapolation beyond the training duration. We demonstrate that with a minimal data cost of only 50k video pairs, VideoCoF achieves state-of-the-art performance on VideoCoF-Bench, validating the efficiency and effectiveness of our approach. Our code, weight, data are available at https://github.com/knightyxp/VideoCoF.

cs.CV

Spinning Primordial Black Holes and Scalar Induced Gravitational Waves from Single Field Inflation

We investigate the formation of primordial black holes (PBHs), their spin and abundance, in a single-field inflationary model based on a mutated hilltop potential inserted with a small step-like feature. This step induces a brief phase of ultra-slow-roll inflation, producing the large enhancement of the scalar power spectrum required for an appreciable amount of PBH abundance. Instead of the commonly used analytical power spectra, we compute the primordial power spectrum accurately by numerically solving the Mukhanov-Sasaki equation. Using the obtained power spectrum, we apply peak theory with $\nabla^2 \zeta$ treated as a Gaussian random field and parametrize the curvature profile by its amplitude $\mu$ and characteristic width $K$. Confining the study to Type-I PBH, the threshold value is calculated using two robust methods: the average of the compaction function and the q-function method. Using the result, the dimensionless spin parameter of the resulting PBHs is calculated at linear order and found to be $\sqrt{\langle a_\star^2 \rangle} \sim 10^{-3}$; however, it can be higher for smaller masses. We present detailed predictions for two representative parameter sets, calculate the present-day PBH mass function $f_{\rm PBH}(M)$ and the associated scalar-induced gravitational waves (SIGW). The first produces PBHs of mass $M \simeq 10^{-13}\,M_\odot$ that can account for $100\%$ of dark matter, while the second yields $M \simeq 10^{-2}\,M_\odot$ PBHs contributing approximately $2.4\%$ of the dark matter density. The predicted signals of SIGWs lie within the sensitivity bands of future experiments such as LISA, DECIGO, BBO, and SKA. In particular, the second parameter set produces a SIGWs compatible with the recent NANOGrav evidence for a low-frequency gravitational-wave signal.

astro-ph.CO

Probing Loop Quantum Gravity black holes through gravitational lensing

We investigate strong gravitational lensing by a charged loop quantum gravity (LQG) black hole obtained through the polymerisation scheme of Borges \textit{et al.} \cite{Borges:2023fog}. These effective geometries replace the Reissner--Nordstr\"om singularity with a symmetric transition surface and admit an extremal, cold remnant determined by the minimal area gap in LQG. In turn, we derive the null geodesic equations, investigate the photon effective potential, and obtain expressions for the photon-sphere radius and critical impact parameter. We compute the weak-field deflection angle and Einstein ring size, highlighting the deviations induced by the polymerisation parameter and the Barbero--Immirzi parameter. In the strong-field regime, we compute the strong deflection coefficients $(\bar{a},\bar{b})$ and evaluate the lensing observables $\theta_\infty$, $s$, and $r_{\rm mag}$. Unlike the Reissner--Nordstr\"om case, the LQG corrections enhance the deflection angle and increase the angular separation of relativistic images, with deviations growing as the geometry approaches the LQG remnant limit. We further compute the corresponding observables for Sgr~A* and M87*, finding that the quantum-gravity modifications lie within the potential sensitivity of next-generation VLBI facilities. For M87*, the angular separation $s\in(0.05712,0.19123)\,\mu\text{as}$, while it is $s\in(0.07595,0.25426)\,\mu\text{as}$ for Sgr A*. The relative flux ratio is found to lie in the range, $r_{\rm mag}\in(4.49272,5.96397)$. Our analysis demonstrates that LQG-induced corrections leave characteristic strong and weak-lensing imprints, offering a promising observational pathway to probe quantum gravity using near-future high-resolution observations.

gr-qc

OTARo: Once Tuning for All Precisions toward Robust On-Device LLMs

Large Language Models (LLMs) fine-tuning techniques not only improve the adaptability to diverse downstream tasks, but also mitigate adverse effects of model quantization. Despite this, conventional quantization suffers from its structural limitation that hinders flexibility during the fine-tuning and deployment stages. Practical on-device tasks demand different quantization precisions (i.e. different bit-widths), e.g., understanding tasks tend to exhibit higher tolerance to reduced precision compared to generation tasks. Conventional quantization, typically relying on scaling factors that are incompatible across bit-widths, fails to support the on-device switching of precisions when confronted with complex real-world scenarios. To overcome the dilemma, we propose OTARo, a novel method that enables on-device LLMs to flexibly switch quantization precisions while maintaining performance robustness through once fine-tuning. OTARo introduces Shared Exponent Floating Point (SEFP), a distinct quantization mechanism, to produce different bit-widths through simple mantissa truncations of a single model. Moreover, to achieve bit-width robustness in downstream applications, OTARo performs a learning process toward losses induced by different bit-widths. The method involves two critical strategies: (1) Exploitation-Exploration Bit-Width Path Search (BPS), which iteratively updates the search path via a designed scoring mechanism; (2) Low-Precision Asynchronous Accumulation (LAA), which performs asynchronous gradient accumulations and delayed updates under low bit-widths. Experiments on popular LLMs, e.g., LLaMA3.2-1B, LLaMA3-8B, demonstrate that OTARo achieves consistently strong and robust performance for all precisions.

cs.LG