SearcharxivSearch

arXiv subjects

Jie An

Publications and source records attributed to Jie An.

At least 19 recordsLinked to original sources

How Powerful are LLMs in Generating Formal Program Specifications?

Formal verification provides strong guarantees of software correctness, but its adoption is limited by the high cost of writing precise formal specifications. While recent large language models (LLMs) have shown strong capabilities in theorem proving and verified code generation, their true ability to generate program specifications remains unclear. Existing evaluations require either verifying implementation conformance or proving semantic equivalence between specifications, both of which are formidably difficult and may conflate proof difficulty with specification quality. To address this problem, we introduce Coins, a Rocq based evaluation framework that assesses specification quality by instantiating specifications under evaluation on trusted test cases and generating concrete proof obligations. This design aligns with the asymmetric nature of formal reasoning, where successful proofs provide reliable evidence while proof failures are inherently ambiguous. Using Coins, we conduct a large scale study on HumanEval with a curated set of human written Rocq specifications. Our results show that specification generation remains a formidable challenge, and that verification complexity can obscure genuine differences in specification quality. Overall, we find that accurate specification evaluation, rather than model scaling alone, is central to understanding the power of LLMs for specification synthesis, and that test case based formal reasoning offers a more faithful and discriminative measure of progress.

cs.SE

Mural: Transferring LLM knowledge to image generation via Mixture-of-Transformers

Leveraging capabilities of large language models (LLMs) in text-to-image (T2I) synthesis is an important research direction. In this work we investigate whether the knowledge of a frozen LLM can be effectively utilized in T2I generation when trained exclusively on standard text-image pairs. We integrate a frozen, reasoning-capable LLM with a diffusion-based image generator via shared attention within the Mixture-of-Transformers (MoT) architecture. Our experiments span two critical questions: (1) what degree of the LLM's intrinsic knowledge remains accessible during T2I training, and (2) what novel capabilities emerge in the resulting system. Across established benchmarks, our models achieve strong performance among unified understanding-generation systems: 0.85 on GenEval, 86.75 on DPG-Bench, and 0.66 on WISE with inference-time reasoning, using only text-image data. Remarkably, we uncover emergent behaviors absent from training data, including cross-lingual image generation, color-guided composition, emoji / ASCII scene construction, and generation directed by world knowledge. These results demonstrate that pretrained LLM knowledge can guide image synthesis under standard text-to-image training paradigms, without interleaved multimodal signals or explicit reasoning supervision. Our findings open new avenues for harnessing frozen model capabilities in resource-constrained multimodal learning.

cs.CV

EP251023a: A fast X-ray transient featuring a magnetar-powered optical internal plateau followed by a steep decay

EP251023a is an extragalactic fast X-ray transient (eFXT) detected solely by EP without a gamma-ray counterpart. The prompt emission consists of a main emission with a duration $T_{90}=292\pm19$ s, followed by a long-lasting tail emission that persists until the observation ends at $T_0+1571$ s. With the upper limit of Konus--Wind, we derived a conservative upper limit on the isotropic gamma-ray energy $E_{\gamma,\rm{iso}}$ of $5.7 \times 10^{52}$ erg for the main emission phase. A redshift of $z = 2.232\pm0.001$ is identified from strong absorption features in the Keck spectrum, which also indicate a relatively low host-galaxy HI column density. Based on the broadband spectral energy distribution, the late-time light curves show an achromatic plateau, followed by an extremely steep decay with a slope of 3.99 after a break at about 49 ks, which is consistent with a rapidly spinning millisecond magnetar engine. Under the isotropic wind scenario, we obtain the initial period $P_0<2.27$~ms and the magnetic field strength $B_p<8.33\times10^{14}$~G for the magnetar; whereas considering a jet collimation with a typical opening angle of 0.1 rad relaxes these constraints to $P_0<32.15$~ms and $B_p<1.18\times10^{16}$~G. Together with GRB\,070707, EP251023a may represent a rare class of optical magnetar-powered internal plateaus with little external-shock contamination, unlike previous examples detected primarily in X-rays. Future discoveries of similar events will help clarify the relationship between magnetar-powered internal emission observed in the optical band and that detected only in X-rays.

astro-ph.HE

GRB 250424A: A Case Study of Energy Injection with Multiwavelength Observations

We present a comprehensive multiwavelength analysis of the long-duration gamma-ray burst (GRB) 250424A. Our dataset spans from the prompt gamma-ray emission to late-time optical monitoring, including spectra obtained with the Keck 10\,m telescope. We find that the afterglow light curves display a prominent, simultaneous shallow decay phase in both X-ray and optical bands, followed by an achromatic transition to a standard decay regime. The broadband spectral energy distributions are well-modeled by a single power-law function, indicating a common synchrotron origin for the emission across frequencies. We interpret the afterglow evolution within the framework of a relativistic forward shock refreshed by continuous energy injection. This scenario successfully reproduces the observed temporal and spectral behavior, yielding an isotropic equivalent kinetic energy of $E_{\rm K,iso} \approx 5.5 \times 10^{52}$ erg and an injection index of $q\approx 0.34$ in a constant-density circumburst environment. The shallow decay phase is consistent with sustained energy injection lasting $\sim$ 9 ks. Despite the relatively low redshift, late-time optical observations reveal no distinct supernova component; however, our derived upper limits do not strictly rule out the presence of a typical GRB-associated supernova.

astro-ph.HE

Failed jet breakout in the metal-poor broad-lined type Ic supernova 2026gzf

A long-standing question in the death of massive stars is the role of relativistic jets. While many gamma-ray bursts and some fast X-ray transients seem to be associated with broad-lined type Ic supernovae, the opposite is not true. The lack of observable jet emission in those Ic-BL SNe can be explained by invoking off-axis jets, choked jets that inject all their energy into the stellar envelope, baryon-loaded jets for which the prompt high-energy emission is strongly suppressed, or non-jetted SNe. The lack of exact explosion time in the majority of SNe presents an obstacle to distinguish between these scenarios. Here we report the properties of SN 2026gzf associated with the X-ray thermal Einstein Probe shock-breakout EP260321a at z=0.0343. The absence of compelling shocked cocoon and radio emission up to 54 days, combined with initial expansion velocities of ~30,000 km/s and a circumstellar shell of ~0.07 M$_\odot$, favour a scenario for SN 2026gzf in which a jet was choked in the circumstellar shell. Our high-spatial resolution images of the SN environment show that the progenitor was located between two highly star-forming regions with a metallicity lower than any previously known Ic-BL SN. As the first case of a Ic-BL SN associated with high-energy prompt emission without the signature of a jet, SN 2026gzf provides a unique perspective to understand the successful launch of relativistic jets during the deaths of massive stars.

astro-ph.HE

X-rays breaking out of pre-explosion ejecta mark a supernova's first light

Massive stars die as core-collapse supernovae, whose optical light emerges days after the implosion. Theory predicts that the initial collapse-driven shock, upon breaking through the star and dense circumstellar medium, emits a brief thermal flash of soft X-rays and ultraviolet. Yet these elusive first signals have remained largely undetected, owing to limited wide-field soft X-ray monitoring. Here we report the discovery of a soft X-ray flash, EP260321a, followed days later by a broad-lined supernova from an envelope-stripped progenitor. Its X-ray spectrum, best modeled with blackbody, establishes it as the long-sought archetypal shock breakout. The burst's duration and energetics place the breakout at a radius of 300 solar radii, tracing a dense surrounding shell and revealing abrupt mass ejection within the final month before collapse.

astro-ph.HE

GRB 260310A / SN 2026fgk: A Multi-Wavelength Study of a Nearby Underluminous Long GRB and SN with a Complex Afterglow

We present a comprehensive multi-wavelength study of GRB 260310A / SN 2026fgk, a nearby ($z=0.153$), long-duration gamma-ray burst (GRB) with an exceptionally underluminous prompt $\gamma$-ray emission and a Comptonized spectrum. The burst occurred at the edge of a blue host galaxy at a projected distance of 15 kpc, which is one of the largest offsets reported for a long GRB. The bright optical afterglow, with dense coverage from COLIBR\'I, likely peaked at a few to several hours post-burst, followed by a shallow decay not expected from canonical afterglow models. Both the optical and X-ray light curves show a brief chromatic plateau from $4-7$ days. We show that the subsequent rebrightening observed at $\sim20$ days is best explained by the combined contribution of the associated Type Ic-BL supernova, identified in GTC spectra, and a late-time refreshed shock. The broadband optical to X-ray spectral energy distribution is well described by synchrotron emission from the forward shock, while the radio observations demand an additional emission component. We model the afterglow using (a) an on-axis uniform jet from a dirty fireball with late-time energy injection and (b) a misaligned jet with power-law angular structure, both having material emitting along our line-of-sight (LOS) moving with an initial Lorentz factor of $\Gamma_0\sim20-35$. We conclude that at more typical GRB distances ($z\gtrsim0.5$) the prompt $\gamma$-ray emission from this source would likely have escaped detection, whereas its optical afterglow would have remained observable, making the event appear as an orphan afterglow or a gamma-ray quiet fast X-ray transient.

astro-ph.HE

ClarifySTL: An Interactive LLM Agent Framework for STL Transformation through Requirements Clarification

Signal Temporal Logic (STL) is a formal language for specifying real-time behaviors of cyber-physical systems (CPS). Automatically transforming natural language requirements into STL specifications has received growing attention. Recent efforts leveraging large language models (LLMs) have demonstrated impressive performance, but some natural language requirements in practice contain vague or ambiguous information, which remains challenging for LLMs to handle. To address these challenges, we propose ClarifySTL, an interactive LLM-agent framework that enhances STL transformation through requirements clarification. ClarifySTL first detects vague expressions that indicate underspecified information in a requirement. If any vagueness is detected, it generates targeted clarification queries to guide users in supplementing the requirement until all necessary details are provided. Subsequently, if ClarifySTL detects ambiguities, it formulates focused ambiguity clarification queries and updates the requirements based on user feedback until all ambiguities are resolved. Finally, the requirements with vagueness and ambiguity clarified are transformed into STL specifications using LLMs. This interactive framework ensures that the resulting STL formulas faithfully capture user intent while reducing the burden on the user. We evaluate ClarifySTL on the representative benchmarks DeepSTL and STL-DivEn, as well as our newly introduced AmbiEval benchmark, which is specifically designed to assess the performance of the agents in handling vagueness and ambiguity, including both detection and query generation. The experimental results show that ClarifySTL is effective.

cs.SE

A fast X-ray transient with chromatic flares: signatures of violent collisions induced by late-time central engine reactivation

Extragalactic Fast X-ray Transients (EFXTs) represent an emerging class of high-energy phenomena characterized by X-ray outbursts lasting from tens to hundreds of seconds. However, for more than half of the EFXTs, their physical origins remain elusive. In this Letter, we report the discovery of EP250302a, a luminous EFXT detected by the Einstein Probe (EP) at a redshift of $z = 1.131$. The multi-wavelength light curves of EP250302a reveal remarkable temporal features that distinguish it from the previously known EP-detected EFXT population, most notably a needle-like X-ray flare accompanied by smooth optical rebrightening during the afterglow phase. We suggest that the distinct X-ray and optical behaviors constitute the first observed instance of late-time violent collision of two relativistic shells in an EFXT. Drawing on insights from GRB studies, such a collision process strongly indicates the reactivation of a central engine, making EP250302a-like transients a unique laboratory for probing the late-time activity and jet physics of EFXT central engines.

astro-ph.HE

STLts-Div: Diversified Trace Synthesis from STL Specifications Using MILP (Extended Version)

Modern cyber-physical systems are complex, and requirements are often written in Signal Temporal Logic (STL). Writing the right STL is difficult in practice; engineers benefit from concrete executions that illustrate what a specification actually admits. Trace synthesis addresses this need, but a single witness rarely suffices to understand intent or explore edge cases - diverse satisfying behaviors are far more informative. We introduce diversified trace synthesis: the automatic generation of sets of behaviorally diverse traces that satisfy a given STL formula. Building on a MILP encoding of STL and system model, we formalize three complementary diversification objectives - Boolean distance, random Boolean distance, and value distance - all captured by an objective function and solved iteratively. We implement these ideas in STLts-Div, a lightweight Python tool that integrates with Gurobi.

eess.SY

Exact Moment Estimation of Stochastic Differential Dynamics

Moment estimation for stochastic differential equations (SDEs) is fundamental to the formal reasoning and verification of stochastic dynamical systems, yet remains challenging and is rarely available in closed form. In this paper, we study time-homogeneous SDEs with polynomial drift and diffusion, and investigate when their moments can be computed exactly. We formalize the notion of moment-solvable SDEs and propose a generic symbolic procedure that, for a given monomial, attempts to construct a finite linear ordinary differential equation (ODE) system governing its moment, thereby enabling exact computation. We introduce a syntactic class of pro-solvable SDEs, characterized by a block-triangular structure, and prove that all polynomial moments of any pro-solvable SDE admit such finite ODE representations. This class strictly generalizes linear SDEs and includes many nonlinear models. Experimental results demonstrate the effectiveness of our approach.

eess.SY

Counterexample Classification for Signal Temporal Logic Specifications

Signal Temporal Logic (STL) has been widely adopted as a specification language for specifying desirable behaviors of hybrid systems. One of the most common uses of STL is falsification, which attempts to generate counterexample signals that demonstrate how the system violates a given STL specification. A number of falsification methods and tools are available for efficient generation of counterexamples, which can be examined by the engineer to identify potential defects in the system. However, some of these counterexamples may be considered similar to each other in that they describe system behavior that stems from the same underlying causes or defects. Since examining counterexamples can be a labor-intensive task, a tool that presents a distinct set of counterexamples and avoids showing repetitive ones could reduce the amount of effort that the engineer spends in debugging. In this paper, we propose a counterexample classification method for STL specifications. Our approach is based on a novel criterion for classifying given counterexamples into a finite set of classes, each of which corresponds to a set of signals that share a common behavioral pattern. In particular, each class is represented by a formula in parametric signal temporal logic (PSTL), which provides a concise description of the signals in the class; then, the problem of checking whether a given signal belongs to a particular class can be formulated as finding parameter values for the corresponding PSTL such that the signal satisfies the formula. We propose an algorithm for automatically identifying classes from a given set of counterexamples and an efficient pruning method that leverages the concept of an inclusion relation between different classes. We demonstrate the efficiency of our algorithm and its utility on three hybrid systems, including automatic transmission, abstract fuel control, and robot navigation.

cs.SE

Minutes-long soft X-ray prompt emission from a compact object merger

Compact object mergers are multi-messenger sources and known progenitors of some gamma-ray bursts, bright flashes of high-energy radiation powered by a central engine, either an accreting black hole or a neutron star. Our understanding of these events has so far been shaped primarily by observations in the gamma-ray band, leaving their prompt phase poorly constrained at lower energies. A long-lasting ($\approx$100 s) engine-driven X-ray emission was discussed to explain rapidly fading X-ray afterglows following several ($\approx$30%) bursts of short ($\lesssim$2 s) duration. However, this prompt X-ray component was not directly observed and past candidates were not confirmed. Here we report the discovery of EP250704a containing a minutes-long ($\sim$560 s) flash of soft (0.5--4 keV) X-rays immediately following the short ($\sim$0.4 s) GRB 250704B. The variability and spectral shape of this emission are inconsistent with the canonical picture of a hard, accretion-powered spike followed by a standard external-shock afterglow. Instead, the long-soft bump points to a distinct phase of prompt emission in X-rays, which would not have been detected without the soft X-ray coverage of Einstein Probe. The detection of a prompt soft X-ray counterpart in an otherwise ordinary short GRB shows that long-lasting X-ray emission is likely a common feature of merger-driven bursts and a promising electromagnetic counterpart to gravitational wave sources.

astro-ph.HE

Quantifier Elimination Meets Treewidth

In this paper, we address the complexity barrier inherent in Fourier-Motzkin elimination (FME) and cylindrical algebraic decomposition (CAD) when eliminating a block of (existential) quantifiers. To mitigate this, we propose exploiting structural sparsity in the variable dependency graph of quantified formulas. Utilizing tools from parameterized algorithms, we investigate the role of treewidth, a parameter that measures the graph's tree-likeness, in the process of quantifier elimination. A novel dynamic programming framework, structured over a tree decomposition of the dependency graph, is developed for applying FME and CAD, and is also extensible to general quantifier elimination procedures. Crucially, we prove that when the treewidth is a constant, the framework achieves a significant exponential complexity improvement for both FME and CAD, reducing the worst-case complexity bound from doubly exponential to single exponential. Preliminary experiments on sparse linear real arithmetic (LRA) and nonlinear real arithmetic (NRA) benchmarks confirm that our algorithm outperforms the existing popular heuristic-based approaches on instances exhibiting low treewidth.

cs.LO

EP250827b/SN 2025wkm: An X-ray Flash-Supernova Powered by a Central Engine and Circumstellar Interaction

We present the discovery of EP250827b/SN 2025wkm, an X-ray Flash (XRF) discovered by the Einstein Probe (EP), accompanied by a broad-line Type Ic supernova (SN Ic-BL) at $z = 0.1194$. EP250827b possesses a prompt X-ray luminosity of $\sim 10^{45} \, \rm{erg \, s^{-1}}$, lasts over 1000 seconds, and has a peak energy $E_{\rm{p}} < 1.5$ keV at 90\% confidence. SN 2025wkm possesses a double-peaked optical light curve (LC), though its bolometric luminosity plateaus after its initial peak for $\sim 20$ days, consistent with a central engine injecting additional energy into the explosion. Its spectrum transitions from a blue to red continuum with clear blueshifted broad absorption features consistent with a SN Ic-BL classification. We do not detect any transient radio emission and rule out the existence of an on-axis, energetic jet $\gtrsim 10^{50}~$erg assuming a typical LGRB circumburst constant density ($n \approx 10^{-3}$--$10^{-1}~{\rm cm}^{-3}$) and microphysical parameters ($\epsilon_{\rm e} = 0.1$ and $\epsilon_{\rm B} = 0.01$). In the model we invoke, the collapse gives rise to a long-lived magnetar, potentially surrounded by an accretion disk. Magnetically--driven winds from the magnetar and the disk mix together and break out with a velocity $\sim 0.35c$ and interact with an extended circumstellar medium with radius $\sim 10^{13}$ cm, generating X-ray breakout emission through non-thermal free-free processes. The disk outflows and magnetar winds power blackbody photospheric emission as they cool adiabatically and thermalize, producing the first SN peak. The spin-down luminosity of the magnetar and radioactive decay of $^{56}$Ni powers the late-time emission. We end by discussing the landscape of XRF-SNe within the context of EP's recent discoveries.

astro-ph.HE

Runtime Safety and Reach-avoid Prediction of Stochastic Systems via Observation-aware Barrier Functions

Stochastic dynamical systems have emerged as fundamental models across numerous application domains, providing powerful mathematical representations for capturing uncertain system behavior. In this paper, we address the problem of runtime safety and reach-avoid probability prediction for discrete-time stochastic systems with online observations, i.e., estimating the probability that the system satisfies a given safety or reach-avoid specification. Unlike traditional approaches that rely solely on offline models, we propose a framework that incorporates real-time observations to dynamically refine probability estimates for safety and reach-avoid events. By introducing observation-aware barrier functions, our method adaptively updates probability bounds as new observations are collected, combining efficient offline computation with online backward iteration. This approach enables rigorous and responsive prediction of safety and reach-avoid probabilities under uncertainty. In addition to the theoretical guarantees, experimental results on benchmark systems demonstrate the practical effectiveness of the proposed method.

eess.SY

RESTL: Reinforcement Learning Guided by Multi-Aspect Rewards for Signal Temporal Logic Transformation

Signal Temporal Logic (STL) is a powerful formal language for specifying real-time specifications of Cyber-Physical Systems (CPS). Transforming specifications written in natural language into STL formulas automatically has attracted increasing attention. Existing rule-based methods depend heavily on rigid pattern matching and domain-specific knowledge, limiting their generalizability and scalability. Recently, Supervised Fine-Tuning (SFT) of large language models (LLMs) has been successfully applied to transform natural language into STL. However, the lack of fine-grained supervision on atomic proposition correctness, semantic fidelity, and formula readability often leads SFT-based methods to produce formulas misaligned with the intended meaning. To address these issues, we propose RESTL, a reinforcement learning (RL)-based framework for the transformation from natural language to STL. RESTL introduces multiple independently trained reward models that provide fine-grained, multi-faceted feedback from four perspectives, i.e., atomic proposition consistency, semantic alignment, formula succinctness, and symbol matching. These reward models are trained with a curriculum learning strategy to improve their feedback accuracy, and their outputs are aggregated into a unified signal that guides the optimization of the STL generator via Proximal Policy Optimization (PPO). Experimental results demonstrate that RESTL significantly outperforms state-of-the-art methods in both automatic metrics and human evaluations.

cs.FL

Control Synthesis of Cyber-Physical Systems for Real-Time Specifications through Causation-Guided Reinforcement Learning

In real-time and safety-critical cyber-physical systems (CPSs), control synthesis must guarantee that generated policies meet stringent timing and correctness requirements under uncertain and dynamic conditions. Signal temporal logic (STL) has emerged as a powerful formalism of expressing real-time constraints, with its semantics enabling quantitative assessment of system behavior. Meanwhile, reinforcement learning (RL) has become an important method for solving control synthesis problems in unknown environments. Recent studies incorporate STL-based reward functions into RL to automatically synthesize control policies. However, the automatically inferred rewards obtained by these methods represent the global assessment of a whole or partial path but do not accumulate the rewards of local changes accurately, so the sparse global rewards may lead to non-convergence and unstable training performances. In this paper, we propose an online reward generation method guided by the online causation monitoring of STL. Our approach continuously monitors system behavior against an STL specification at each control step, computing the quantitative distance toward satisfaction or violation and thereby producing rewards that reflect instantaneous state dynamics. Additionally, we provide a smooth approximation of the causation semantics to overcome the discontinuity of the causation semantics and make it differentiable for using deep-RL methods. We have implemented a prototype tool and evaluated it in the Gym environment on a variety of continuously controlled benchmarks. Experimental results show that our proposed STL-guided RL method with online causation semantics outperforms existing relevant STL-guided RL methods, providing a more robust and efficient reward generation framework for deep-RL.

cs.AI