SearcharxivSearch

arXiv subjects

Ting Li

Publications and source records attributed to Ting Li.

At least 19 recordsLinked to original sources

Ecdysis: Efficient and Effective Training of Runtime Harnesses for LLM Agents

Self-evolving runtime harnesses can substantially improve the capabilities of large language model (LLM) agents and provide a promising paradigm for optimizing agent execution. Existing harness evolution methods typically rely on iterative search, repeatedly evaluating and revising candidate harnesses based on execution feedback from task instances. While this paradigm enables continuous harness optimization, it incurs substantial time overhead due to repeated agent executions and code modifications, and may overfit to observed tasks and specific failure patterns, resulting in degraded generalization to unseen tasks. We identify the lack of principled failure diagnosis as a key bottleneck in harness evolution: an observed failure can reflect either model-specific deficiencies or systematic harness deficiencies, and directly optimizing against individual failures can lead to unnecessary model-specific accommodation. We therefore propose Ecdysis, an efficient and effective framework that distinguishes model-specific accommodation from harness-level repair and biases adaptation toward systematic harness deficiencies by identifying recurring cross-task failure patterns. Ecdysis adopts a batch-level cross-instance failure aggregation paradigm to jointly analyze failure evidence from multiple task instances and further introduces Failure-Driven Collaborative Refinement to diagnose failure causes and iteratively refine harness modification specifications. By combining cross-instance failure analysis with multi-role diagnosis, Ecdysis enables more effective harness evolution with lower training time. Experiments show that Ecdysis achieves up to a 1.84x speedup in harness training compared with existing harness evolution methods, while improving the reasoning accuracy of the resulting harnesses by 18.56%.

cs.SE

Empirical-Bayes Elastic-Net Computation for Exponential Random Graph Models

Exponential random graph models (ERGMs) describe dependence among network ties, but inference becomes difficult when the likelihood is intractable and candidate network statistics are strongly correlated. We introduce BERGM Elastic Net, an adaptive empirical-Bayes approach that combines lasso shrinkage with ridge stabilization in a Bayesian ERGM. A latent-variable formulation supports approximate exchange sampling, while empirical-Bayes updates adapt the amount of regularization to the observed network. We connect the proposed prior to elastic-net penalized likelihood and clarify the interpretation of thresholded reporting and coefficient grouping. The method is developed for over-specified network models containing many related structural and covariate effects.

stat.ME

Statistical Study of Solar Prominence Plumes Based on NVST H$\alpha$ Observations

Plumes are one of the most representative dynamic features observed in prominences and play a key role in mass and magnetic transport within them. However, their physical nature and triggering processes remain actively debated. Based on limb H$\alpha$ observations from the New Vacuum Solar Telescope (NVST) during 2013--2025, we statistically investigated 34 plumes with clear and complete evolutions by developing an automated image-processing pipeline. It is revealed that plume lifetimes mainly range from 300 s to 700 s, with vertical displacements between 3--7 Mm. The mean widths and velocities are concentrated in the range of 0.5--1.5 Mm and 10--20 km s$^{-1}$, respectively. Besides wide distribution ranges, plume parameters exhibit irregular evolution fluctuations, indicating that the formation and evolution of various plumes may exhibit different physical patterns. Correlation analysis among the parameters further reveals that: (1) Positive correlations were found among lifetime, vertical displacement, and mean width, indicating an intrinsic coupling between the temporal and spatial scales of plumes. (2) Trajectory curvature is negatively correlated with lifetime, vertical displacement, and velocity. Accelerating and width-contracting plumes typically have lower curvature, suggesting that curvature may reflect environmental influences and the stability of plumes. (3) Plumes with higher initial velocities were more likely to be accompanied by precursor brightening, suggesting that these plumes may be triggered by magnetic reconnection. Furthermore, we infer that some plumes in non-bubble regions may be inherently driven by mini-filament eruptions. These results establish a statistical framework for prominence plumes and reveal diversity in their dynamical evolution and triggering mechanisms.

astro-ph.SR

Offline-Online Curriculum RL for Multimodal Reasoning

Multimodal large language models exhibit capabilities on reasoning tasks, yet often produce flawed intermediate steps while yielding correct final answers. This behavior undermines interpretability and reliability, suggesting reliance on spurious shortcuts rather than faithful reasoning. Although efforts have explored step-level supervision, distinguishing decisive steps from redundant ones remains challenging. We propose $O^2$-CritiCuRL, a novel curriculum reinforcement learning framework that introduces critical-step awareness through an iterative offline-online paradigm. In the offline stage, $O^2$-CritiCuRL conducts multi-rollout analysis over step-annotated trajectories to estimate step-level importance, allowing the framework to distill critical reasoning steps and filter out redundant ones. In the online stage, we employ a progressive step-level reinforcement learning strategy, where truncated chains guide the model to infer missing steps and refine its reasoning, thereby sharpening its focus on critical steps and overcoming the limitations of static supervision. Extensive experiments on multimodal reasoning benchmarks show that our method achieves state-of-the-art performance while delivering superior training and inference efficiency. Code is available at https://github.com/kk0013/CritiCuRL.

cs.AI

Efficient Online LLM Watermark Detection via Rao-Blackwellized E-Processes

As large language models (LLMs) are increasingly deployed, reliable and efficient mechanisms for distinguishing AI-generated text from human-written content have become essential. Statistical watermarking has emerged as a promising solution, yet most existing methods are typically fixed-horizon procedures, precluding valid early stopping in streaming generation. In this paper, we develop an efficient online watermark detection framework with anytime-valid inference based on Rao-Blackwellized e-processes, enabling recursive token-level evidence updates without storing the full history. In particular, we instantiate the framework for the Gumbel-max watermark and reduce the original token-level dependence testing problem to a pivot-induced sequential testing problem with an explicit null distribution. Theoretically, we prove anytime-valid Type I error control under arbitrary optional stopping and establish positive asymptotic log-growth under watermarking, implying consistency of the proposed stopping rules. Simulations and experiments on real LLM-generated text demonstrate efficient online detection with rigorous anytime-valid guarantees.

stat.ML

Unprecedent fast winking of solar flares triggered by bursty magnetic reconnection

Flare ribbons form as a result of energy deposition associated with particles accelerated in low layers of the solar atmosphere. The fine-scale structures of flare ribbons, also called ribbon kernels, offer a potentially powerful diagnostic of the flare reconnection process, however to date the dynamic evolution of ribbon kernels has not been fully characterized in statistical studies. Here, we checked the state-of-the-art observations (cadence $\leq$ 2.5 seconds) of solar flares in the ultraviolet from space by Interface Region Imaging Spectrograph (IRIS) over the past 12 years. Our results showed the first statistical study of multiple spatially-resolved flare kernel quasi-periodic pulsation events for 31 flares, with the period of 6-24 seconds. The ribbon kernels have a spatial scale of 480$-$1200 km and some kernels exhibit unprecedent fast ``winking" process, i.e., quasi-periodic pulsation-like flashing of individual kernels. The shortest heating time reaches about 2$-$3 s, implying that the energy is deposited only in a small localized region within flare ribbons, persisting for only a few seconds. Meanwhile, some ribbon kernels were observed to slip along the ribbon at speeds of 20-1800 km s$^{-1}$. These observations strongly imply a joint picture for the dynamics and the bursty nature of ribbon kernels as being due to coupled effects of plasmoid formation and three-dimensional (3D) magnetic reconnection in the overlaying coronal current sheet. We suggest that the observed flare behaviors provide strong observational evidences of 3D bursty reconnection.

astro-ph.SR

Coarse-to-Fine: A Hybrid Self-Supervised Method for Non-rigid 3D Shape Matching

Non-rigid 3D shape matching is a fundamental task in computer vision and graphics. In this paper, we propose a hybrid self-supervised method based on a coarse-to-fine strategy, which ensures consistency between the coarse mapping and the refined correspondence produced by our refinement module. The architecture features a dual-branch design, consisting of two symmetric functional map learning streams: one based on the Laplacian basis and the other utilizing the elastic basis. Extensive experiments show that our approach not only maintains computational efficiency, but also achieves state-of-the-art performance across a variety of challenging scenarios, including non-isometric deformations and topological noise. Finally, we rigorously demonstrate that contrastive energies promote feature discrimination. Furthermore, integrating these energies with existing methods yields consistent improvements, validating the overall efficacy of our approach. Our code is available at https://github.com/LuoFeifan77/Coarse-to-Fine-Hybrid-Self-Supervised-Matching.

cs.CV

Revisiting Approaches to Stellar White-Light Flare Energy Based on Spatiotemporally Resolved Solar Observations

Accurately estimating the bolometric energy of solar and stellar white-light flares (WLFs) is crucial for understanding their physical nature and impact on surrounding planets. However, the lack of spatial resolution in stellar observations forced pioneering stellar WLF studies to adopt simplified energy estimation methods, typically assuming either a constant flare temperature or a fixed radiating area. To assess the physical plausibility of these assumptions, we utilize high-spatiotemporal-resolution solar observations to analyze the true evolution of source region's radiating area and temperature of 70 solar WLFs. It is revealed that both area and temperature of most solar WLFs undergo significant temporal evolution, and the flare area strongly correlates with the flare's peak optical continuum flux. Therefore, we propose a new energy estimation method that permits both flare area and temperature to evolve. Compared with existing methods, our dynamic approach yields systematically lower flare energies, which then prompts us to revisit classical macroscopic scaling laws related to the flare energy. It is further revealed that different energy estimation approaches can systematically alter these scaling relations, calling for a re-examination of these established statistical results and their targeted testing or revision in future work.

astro-ph.SR

Global Average Treatment Effects for Individualized Randomization Experiments with Aggregate Data

Individualized randomized experiments are central to online platforms for optimizing personalized decisions in complex environments. In two-sided markets, however, standard treatment effect estimation is often invalid due to strong temporal and cross-unit interference, a challenge compounded when only aggregated data are available because of privacy or system constraints. To address these issues, we identify the Global Average Treatment Effect (GATE) using only group-level data from treatment and control groups. We first establish identification conditions based on aggregated observations, and then propose the Individualized Randomized Experiment Varying Coefficient Decision Process (IRE-VCDP) model, which accounts for interference through supply-demand dynamics. Building on this framework, we develop a complete procedure for estimation and statistical inference of the GATE, along with theoretical guarantees for the proposed test. Extensive simulations and real-world experiments using data from a leading ridesharing platform demonstrate the effectiveness of our approach.

stat.ME

Distill: Uncovering the True Intent behind Human-Robot Communication

As robots become increasingly integrated into everyday environments, intuitive communication paradigms such as natural language and end-user programming have become indispensable for specifying autonomous robot behavior. However, these mechanisms are ineffective at fully capturing user intent: natural language is imprecise and ambiguous, whereas end-user programming can be overly specific. As a result, understanding what users truly mean when they interact with robots remains a central challenge for human-AI communication systems. To address this issue, we propose the Distill approach for human-robot communication interfaces. Given a task specification provided by the user, Distill (1) removes unnecessary steps; (2) generalizes the meaning behind individual steps; and (3) relaxes ordering constraints between steps. We implemented Distill on a web interface and, through a crowdsourcing study, demonstrated its ability to elicit and refine user intent from initial task specifications.

cs.RO

Robust Sequential Experimental Design for A/B Testing

Experimental design has emerged as a powerful approach for improving the sample efficiency of A/B testing, yet existing designs rely critically on correctly specified models. We study robust sequential experimental design under model misspecification and develop a unified framework that covers both contextual bandit and dynamic settings. Theoretically, we prove that our design bounds the worst-case mean squared error of the estimated treatment effect. Empirically, we demonstrate the effectiveness of the proposed approach using synthetic and real-world datasets from a leading technology company.

stat.ML

Learning Perturbations to Extrapolate Your LLM

Recent advancements in large language models demonstrate that injecting perturbations can substantially enhance extrapolation performance. However, current approaches often rely on discrete perturbations with fixed designs, which limits their flexibility. In this work, we propose a framework where token prefixes are perturbed by a learnable transformation of a continuous latent vector within an embedding space. To overcome the challenge of an intractable marginal likelihood, we derive unbiased estimating equations for model parameters and optimize them via stochastic gradient descent. We establish the statistical properties of the resulting estimator in over-parameterized regimes. Empirical evaluations on both synthetic and real-world datasets demonstrate that our proposal yields significant gains in out-of-domain settings over a range of state-of-the-art baseline methods.

stat.ML

Can We Distinguish the Source Region Location of Filament/Prominence Eruptions from the Sun-as-a-star H$\alpha$ Spectrum?

Solar filament/prominence eruptions can significantly perturb geospace when originating from favorable source locations and directions. While stellar analogs have been recently reported, the disk locations and magnetic environments of their source regions remain spatially unresolved on other stars. To bridge this gap, we investigate the typical Sun-as-a-star H$\alpha$ temporal spectral characteristics of solar filament/prominence eruptions with different source region locations (on-disk vs. limb, active region vs. quiet-Sun region). It is revealed that limb eruptions are characterized by blueshifted/redshifted emission caused by the bright off-limb erupting structures, whereas on-disk eruptions may show blueshifted absorptions due to the dark erupting filaments. Among the limb eruptions, front-side limb eruptions usually display line center emission before the blueshifted/redshifted emission, while far-side limb eruptions show the opposite sequence. Moreover, the magnetic environment at source also shapes the spectral characteristics. On-disk filament eruptions from active region exhibit much more intense flare-ribbon-dominated line center emission features compared with those from quiet-Sun region. Limb active region eruptions often show single-wing emissions, whereas large-scale quiet-Sun region (quiescent) prominence eruptions frequently display expansion-induced emission in both wings followed by line center absorption due to the disappearance of bright prominence. These distinct Sun-as-a-star H$\alpha$ spectral characteristics, dependent on eruption location, provide a diagnostic basis for inferring source regions of stellar filament/prominence eruptions from spatially unresolved H$\alpha$ spectra.

astro-ph.SR

Magnetic Evolution of Highly-Sheared Region in Active Region 13842 Producing Large X9.0 Flare

Shearing motion and magnetic flux cancellation around the polarity inversion line (PIL) play significant roles in the build-up of free magnetic energy and magnetic flux rope (MFR) in source region of major solar flares. Here we investigate the magnetic evolution of a highly-sheared PIL in active region (AR) 13842, hosting the largest X9.0 flare of Solar Cycle 25. Since 2024 September 29, a positive-polarity pore persistently drifted northward along the western side of the AR's main negative-polarity sunspot. The main sunspot remained stationary until negative-polarity patches successively emerged to its east and approached. Rear-ended by these same-polarity patches, the sunspot then began moving westward toward the opposite-polarity pore around October 1, forming a collisional PIL. Meanwhile, on the PIL's other side, the pore was also rear-ended by same-polarity patches sequentially emerging behind it, accelerating the shearing motion around the PIL, where frequent flux cancellations were also observed. Synchronous rapid accumulation of free magnetic energy and formation of MFR were then observed in the PIL, where multiple major flares successively occurred within two days. Before these large flares, the area and total free energy of the high-free-energy-density PIL region gradually decreased in the photosphere, which could be caused by the initial ascent of MFR before eruption and serve as a precursor of solar eruptions. These results suggest that persistent flux emergences with cross separation directions facilitates rapid formation of collisional shearing PIL and frequent flux cancellations, leading to repeated MFR formations and multiple large flares in a relatively short time.

astro-ph.SR

Solar Energetic Particle Events and Associated Type II Radio Bursts from Different Source Regions

Large solar energetic particle (SEP) events are thought to originate from the shocks driven by fast coronal mass ejections (CMEs) and thus generally accompanied by type II radio bursts. However, a significant proportion of type II radio bursts is not accompanied by SEP events. To study the relationship between SEPs and type II radio bursts and the associated physical mechanisms, we statistically analyze 43 SEP halo-CMEs and 131 non-SEP halo-CMEs observed from 2010 to 2024, and check the related properties of type II radio bursts and solar source region. We find nearly all SEP events and approximately two-thirds of non-SEP events are accompanied by type II radio bursts. Type II radio bursts associated with SEP events usually have longer duration and lower ending frequencies. The starting frequency exhibits a clear source region dependence, being highest for ''single active region (AR)'', intermediate for ''multiple ARs'', and lowest for ''outside of ARs''. Furthermore, the spectra of both protons and electrons exhibit a similar softening trend in the three types of source regions. Joint analysis of spectra and type II radio bursts reveals that the proton spectra index has a good anti-correlation with the starting frequency of the type II radio bursts. Our statistical results have important implications for the mechanisms behind SEP acceleration

astro-ph.SR

Riemannian Motion Generation: A Unified Framework for Human Motion Representation and Generation via Riemannian Flow Matching

Human motion generation is often learned in Euclidean spaces, although valid motions follow structured non-Euclidean geometry. We present Riemannian Motion Generation (RMG), a unified framework that represents motion on a product manifold and learns dynamics via Riemannian flow matching. RMG factorizes motion into several manifold factors, yielding a scale-free representation with intrinsic normalization, and uses geodesic interpolation, tangent-space supervision, and manifold-preserving ODE integration for training and sampling. On HumanML3D, RMG achieves state-of-the-art FID in the HumanML3D format (0.043) and ranks first on all reported metrics under the MotionStreamer format. On MotionMillion, it also surpasses strong baselines (FID 5.6, R@1 0.86). Ablations show that the compact $\mathscr{T}+\mathscr{R}$ (translation + rotations) representation is the most stable and effective, highlighting geometry-aware modeling as a practical and scalable route to high-fidelity motion generation.

cs.CV

Designing Time Series Experiments in A/B Testing with Transformer Reinforcement Learning

A/B testing has become a gold standard for modern technological companies to conduct policy evaluation. Yet, its application to time series experiments, where policies are sequentially assigned over time, remains challenging. Existing designs suffer from two limitations: (i) they do not fully leverage the entire history for treatment allocation; (ii) they rely on strong assumptions to approximate the objective function (e.g., the mean squared error of the estimated treatment effect) for optimizing the design. We first establish an impossibility theorem showing that failure to condition on the full history leads to suboptimal designs, due to the dynamic dependencies in time series experiments. To address both limitations simultaneously, we next propose a transformer reinforcement learning (RL) approach which leverages transformers to condition allocation on the entire history and employs RL to directly optimize the MSE without relying on restrictive assumptions. Empirical evaluations on synthetic data, a publicly available dispatch simulator, and a real-world ridesharing dataset demonstrate that our proposal consistently outperforms existing designs.

cs.LG

Operational Solar Flare Forecasting System Using an Explainable Large Language Model

This study focuses on forecasting major (>=M-class) solar flares that can severely impact the near-Earth environment. We construct two types of datasets using the Space Weather HMI Active Region Patches (SHARP), and develop a flare prediction network based on large language model (LLMFlareNet). We apply SHapley Additive exPlanations (SHAP) to explain the model predictions. We develop an operational forecasting system based on the LLMFlareNet model. We adopt a daily mode for performance comparison across various operational forecasting systems under identical active region (AR) number and prediction date, using daily operational observational data. The main results are as follows. (1) Through ablation experiments and comparison with baseline models, LLMFlareNet achieves the best TSS scores of 0.720 +/- 0.040 on the ten cross-validation (CV) dataset with mixed ARs. (2) By both global and local SHAP analyses, we identify that R_VALUE is the most influential physical feature for the prediction of LLMFlareNet, aligning with flare magnetic reconnection theory. (3) In daily mode, LLMFlareNet achieves TSS scores of 0.680/0.571 (0.689/0.661, respectively) on the dataset with single/mixed ARs, markedly outperforming NASA/CCMC (SolarFlareNet, respectively). This work introduces the first application of a large language model as a universal computation engine with explainability method in this domain, and presents the first comparison between operational flare forecasting systems in daily mode. The proposed LLMFlareNet-based system demonstrates substantial improvements over existing systems.

astro-ph.SR