SearcharxivSearch

arXiv subjects

Ziqi Wu

Publications and source records attributed to Ziqi Wu.

11 recordsLinked to original sources

DART-SD: Diamond-topology Aware Retrieval and Tuning for Self-Distillation of Multi-Turn Tool-Calling Agents

Equipping Large Language Models (LLMs) with multi-turn tool-calling capabilities is essential for building autonomous agents. However, progress is fundamentally limited by the reliance on full-length trajectory imitation. For tasks involving multiple order-independent sub-goals, the optimal solution space forms a vast combinatorial diamond lattice. Forcing this rich topology into monolithic trajectories causes a severe topological collapse, indiscriminately penalizing valid alternative explorations and severely degrading policy diversity. To address this, we propose DART-SD (Diamond-topology Aware Retrieval and Tuning for Self-Distillation), a novel framework that shifts the paradigm from global forcing to topology-guided localized correction. DART-SD first models the execution process as a converging Interaction-State Transition Graph (ISTG), faithfully capturing the inherent diamond topology of successful and failed exploratory paths. During autonomous rollouts, the framework identifies the Critical Topological Breakpoint (CTB) and retrieves success-supported recovery references. Finally, we introduce a progressive self-distillation paradigm through CTB-guided localized supervision, ensuring that the training loss is calculated exclusively on the generated recovery steps while strictly protecting the valid reasoning prefix from destructive gradient updates. Experiments on complex multi-turn tool-calling benchmarks demonstrate that DART-SD significantly outperforms traditional full-trajectory baselines.

cs.CL

From Storage to Access: Verifiable Activation of Parametric Knowledge in LLMs via Explicit Priming and Implicit Reasoning

Although Large Language Models (LLMs) encode rich factual knowledge in their parameters, reliably recalling and verifying such knowledge remains a key bottleneck in factual question answering. Existing end-to-end methods entangle knowledge elicitation with reasoning, making it difficult to determine whether correct answers arise from parametric knowledge or the input context. To address this challenge, we propose VAKE (Verifiable Activation of Parametric KnowledgE), a two-stage reinforcement-learning framework that externalizes latent parametric knowledge through explicit Priming and transfers the acquired elicitation capability to implicit Reasoning. Given a query and an insufficient retrieved subgraph, the Priming policy explicitly inserts bridging triples as verifiable evidence, with supervision provided by rewards derived from answers generated by a separate frozen model over the augmented subgraph. Building on the policy learned during Priming, the Reasoning stage trains the model to answer from the original input, testing whether the capability acquired through explicit knowledge elicitation transfers to implicit reasoning. Experiments across seven benchmarks and models from 3B to 14B show that VAKE consistently outperforms standard baselines, including when transferring directly from HotpotQA to OOD datasets. LLM-based evaluation further shows that over 80% of the inserted triples provide factual bridging knowledge not derivable from the retrieved context, while more than half elicit knowledge inaccessible through direct prompting. These results suggest that VAKE activates latent parametric knowledge rather than copying the input context or memorizing dataset-specific associations.

cs.CL

NanoMorph-3D: An End-to-End Physics-Driven Unrolling Framework for Nanomaterial Reconstruction

Precise 3D characterization of nanomaterials is essential for unlocking structure-property relationships. However, standard electron tomography is fundamentally limited by the missing wedge problem. Consequently, conventional algorithms suffer from severe geometric distortions, a challenge further complicated by pervasive noise interference. Current learning-based methods either rely on physics-blind post-processing or employ end-to-end architectures constrained by local receptive fields, failing to capture complex 3D topologies. We propose NanoMorph-3D, a unified end-to-end framework grounded in a comprehensive Nanomorphological Taxonomy. Powered by a large-scale synthetic dataset explicitly modeling non-linear electron attenuation, we design a Physics-Driven Unrolled Network mapping proximal gradient descent into a learnable architecture. To capture complex internal topologies, we formulate a hierarchical attention mechanism with Physics-Normalization for long-range 3D dependencies and scale invariance. Crucially, our Dual-Domain strategy leverages Sinusoidal Attention to explicitly model physical projection trajectories, enforcing strict sinogram consistency to mitigate missing wedge artifacts. Finally, an unsupervised dual-stream mechanism bridges the simulation-to-reality gap. Experiments demonstrate NanoMorph-3D reconstructs diverse topologies with superior fidelity and speed.

cs.CV

A Non-Spherical Model for the Solar Coronal Magnetic Field

The coronal magnetic field plays a fundamental role in governing coronal activities, driving space-weather events, and shaping the heliosphere. Due to a lack of direct observations, extrapolation models such as the Potential Field Source Surface (PFSS) model become the primary method to obtain the three-dimensional magnetic field distribution in the corona. However, the PFSS model cannot solve the long-standing open-flux problem, in which the extrapolated open magnetic flux is significantly lower than that inferred from in-situ measurements. To address this issue, we develop a Non-Spherical Potential Field (NSPF) model. The model introduces a Non-Spherical Source Surface (NSSS) defined as an isosurface of the total magnetic field. The NSSS naturally forms concave structures beneath external current sheets, enabling the model to generate substantially more open magnetic flux while yielding a physically plausible distribution of open field regions. As a result, the NSPF model successfully reproduces complex coronal magnetic topologies, interplanetary magnetic field properties, and solar wind source mappings. Our refined coronal magnetic model provides a useful framework for future research on solar and heliospheric magnetic coupling.

astro-ph.SR

Imaging magnetically driven astrospheres: a forward modelling approach

An astrosphere is a vast, tailed bubble-like volume around a star, formed through the interaction between the stellar magnetic field, the stellar wind, and the interstellar medium (ISM). Detecting and characterizing astrospheres are essential for constraining stellar wind properties, understanding stellar evolution, and assessing the habitability of surrounding exoplanetary systems. Charge exchanges between ionized stellar wind particles and cold ISM hydrogen atoms populate the astrosphere with neutral hydrogen, which can leave observable signatures in the Lyman-$α$ (Ly$α$) line absorption profile. Previous studies have inferred stellar mass-loss rates by measuring Ly$α$ absorption in stellar spectra caused by astrospheric neutral hydrogen. However, our knowledge of the global morphology of astrospheres remains limited and largely dependent on sometimes contradictory simulations. Here we investigate the feasibility of detecting Ly$α$ emission generated by resonant scattering from \NH{} surrounding the star, enabling the construction of a two-dimensional map of the astrosphere. With a three-dimensional magnetohydrodynamic astrosphere model, we perform forward modelling of the Ly$α$ emission and assess the observation feasibility according to the observational limits of the {\it Hubble Space Telescope} (HST). We further discuss the influence of varied line-of-sight orientations and averaged ISM velocity along the line-of-sight. The spatially resolved circumstellar Ly$α$ emission could provide important constraints on the astrospheric configuration and stellar wind properties, such as the bow shock standing distance, the stellar wind symmetry, and the shape of the astro-tail. Our results highlight Ly$α$ astrosphere detections as a promising science case for {\it HST} and future missions such as the \textit{Habitable Worlds Observatory}.}

astro-ph.SR

Statistical Characteristic-Guided Denoising for Rapid High-Resolution Transmission Electron Microscopy Imaging

High-Resolution Transmission Electron Microscopy (HRTEM) enables atomic-scale observation of nucleation dynamics, which boosts the studies of advanced solid materials. Nonetheless, due to the millisecond-scale rapid change of nucleation, it requires short-exposure rapid imaging, leading to severe noise that obscures atomic positions. In this work, we propose a statistical characteristic-guided denoising network, which utilizes statistical characteristics to guide the denoising process in both spatial and frequency domains. In the spatial domain, we present spatial deviation-guided weighting to select appropriate convolution operations for each spatial position based on deviation characteristic. In the frequency domain, we present frequency band-guided weighting to enhance signals and suppress noise based on band characteristics. We also develop an HRTEM-specific noise calibration method and generate a dataset with disordered structures and realistic HRTEM image noises. It can ensure the denoising performance of models on real images for nucleation observation. Experiments on synthetic and real data show our method outperforms the state-of-the-art methods in HRTEM image denoising, with effectiveness in the localization downstream task. Code will be available at https://github.com/HeasonLee/SCGN.

cs.CV

Current Helicity in Response to Coronal Mass Ejections

Coronal mass ejections (CMEs), powerful solar eruptions with massive plasma ejected into the interplanetary space, are caused by the release of the magnetic free enengy stored in coronal electric currents. Photospheric current helicity, defined as the integral of the product of vertical electric current density and vertical magnetic field ($H_c=\int j_zB_z\ dS$), serves as a key parameter in understanding the eruptions. Using a 3D magnetohydrodynamic model, we identify a current helicity reversal pattern associated with the eruption: a pre-eruption decrease and a post-eruption increase. This helicity reversal is attributed to the redistribution of electric currents: before the eruption, currents concentrate toward the polarity inversion line (PIL); after the eruption they move away from the PIL, consistent with the flare ribbon separation, which is caused by the upward progression reconnection site. To validate this pattern, we conducted an observational analysis of 50 $\geq$M5.0 eruptive flares. The results reveal that 58\% of cases exhibited a pre-eruption decrease and 92\% showed the post-eruption increase in current helicity. Detailed analysis of two cases with this reversal suggests that they share the same current redistribution pattern, consistent with the mechanism identified in the simulations. Moreover, the pre-eruption decrease could be observed clearly even in the long-term evolution of the two cases. Current helicity can serve as an indicator of when electric currents are built up for the subsequent eruption, and it has the potential to predict CMEs to some extent.

astro-ph.SR

Noise Calibration and Spatial-Frequency Interactive Network for STEM Image Enhancement

Scanning Transmission Electron Microscopy (STEM) enables the observation of atomic arrangements at sub-angstrom resolution, allowing for atomically resolved analysis of the physical and chemical properties of materials. However, due to the effects of noise, electron beam damage, sample thickness, etc, obtaining satisfactory atomic-level images is often challenging. Enhancing STEM images can reveal clearer structural details of materials. Nonetheless, existing STEM image enhancement methods usually overlook unique features in the frequency domain, and existing datasets lack realism and generality. To resolve these issues, in this paper, we develop noise calibration, data synthesis, and enhancement methods for STEM images. We first present a STEM noise calibration method, which is used to synthesize more realistic STEM images. The parameters of background noise, scan noise, and pointwise noise are obtained by statistical analysis and fitting of real STEM images containing atoms. Then we use these parameters to develop a more general dataset that considers both regular and random atomic arrangements and includes both HAADF and BF mode images. Finally, we design a spatial-frequency interactive network for STEM image enhancement, which can explore the information in the frequency domain formed by the periodicity of atomic arrangement. Experimental results show that our data is closer to real STEM images and achieves better enhancement performances together with our network. Code will be available at https://github.com/HeasonLee/SFIN}{https://github.com/HeasonLee/SFIN.

cs.CV

The Solar Origin of an Intense Geomagnetic Storm on 2023 December 1st: Successive Slipping and Eruption of Multiple Magnetic Flux Ropes

The solar eruption that occurred on 2023 November 28 (SOL2023-11-28) triggered an intense geomagnetic storm on Earth on 2023 December 1. The associated Earth's auroras manifested at the most southern latitudes in the northern hemisphere observed in the past two decades. In order to explore the profound geoeffectiveness of this event, we conducted a comprehensive analysis of its solar origin to offer potential factors contributing to its impact. Magnetic flux ropes (MFRs) are twisted magnetic structures recognized as significant contributors to coronal mass ejections (CMEs), thereby impacting space weather greatly. In this event, we identified multiple MFRs in the solar active region and observed distinct slipping processes of the three MFRs: MFR1, MFR2, and MFR3. All three MFRs exhibit slipping motions at a speed of 40--137 km s$^{-1}$, extending beyond their original locations. Notably, the slipping of MFR2 extends to $\sim$30 Mm and initiate the eruption of MFR3. Ultimately, MFR1's eruption results in an M3.4-class flare and a CME, while MFR2 and MFR3 collectively produce an M9.8-class flare and another halo CME. This study shows the slipping process in a multi-MFR system, showing how one MFR's slipping can trigger the eruption of another MFR. We propose that the CME--CME interactions caused by multiple MFR eruptions may contribute to the significant geoeffectiveness.

astro-ph.SR

Walle: An End-to-End, General-Purpose, and Large-Scale Production System for Device-Cloud Collaborative Machine Learning

To break the bottlenecks of mainstream cloud-based machine learning (ML) paradigm, we adopt device-cloud collaborative ML and build the first end-to-end and general-purpose system, called Walle, as the foundation. Walle consists of a deployment platform, distributing ML tasks to billion-scale devices in time; a data pipeline, efficiently preparing task input; and a compute container, providing a cross-platform and high-performance execution environment, while facilitating daily task iteration. Specifically, the compute container is based on Mobile Neural Network (MNN), a tensor compute engine along with the data processing and model execution libraries, which are exposed through a refined Python thread-level virtual machine (VM) to support diverse ML tasks and concurrent task execution. The core of MNN is the novel mechanisms of operator decomposition and semi-auto search, sharply reducing the workload in manually optimizing hundreds of operators for tens of hardware backends and further quickly identifying the best backend with runtime optimization for a computation graph. The data pipeline introduces an on-device stream processing framework to enable processing user behavior data at source. The deployment platform releases ML tasks with an efficient push-then-pull method and supports multi-granularity deployment policies. We evaluate Walle in practical e-commerce application scenarios to demonstrate its effectiveness, efficiency, and scalability. Extensive micro-benchmarks also highlight the superior performance of MNN and the Python thread-level VM. Walle has been in large-scale production use in Alibaba, while MNN has been open source with a broad impact in the community.

cs.LG

MNN: A Universal and Efficient Inference Engine

Deploying deep learning models on mobile devices draws more and more attention recently. However, designing an efficient inference engine on devices is under the great challenges of model compatibility, device diversity, and resource limitation. To deal with these challenges, we propose Mobile Neural Network (MNN), a universal and efficient inference engine tailored to mobile applications. In this paper, the contributions of MNN include: (1) presenting a mechanism called pre-inference that manages to conduct runtime optimization; (2)deliveringthorough kernel optimization on operators to achieve optimal computation performance; (3) introducing backend abstraction module which enables hybrid scheduling and keeps the engine lightweight. Extensive benchmark experiments demonstrate that MNN performs favorably against other popular lightweight deep learning frameworks. MNN is available to public at: https://github.com/alibaba/MNN.

cs.CV