SearcharxivSearch

arXiv subjects

Hongwei Chen

Publications and source records attributed to Hongwei Chen.

At least 19 recordsLinked to original sources

CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval

Reusable skill libraries allow large language model (LLM) agents to reuse procedural knowledge across tasks, but they also turn memory access into a challenging retrieval problem. Full-library prompting preserves coverage at high context cost, vector retrieval returns compact neighborhoods but treats skills as independent text, and graph-based retrieval can recover workflow context only when the edges that carry relevance are reliable. We propose CaSKG, a counterfactual-causal skill graph framework that calibrates procedural relations before retrieval. CaSKG first builds a high-recall directed candidate graph from semantic, lexical, input/output, and structural evidence, with repair evidence and an optional LLM judge further refining candidate scores. It then applies direction-conditioned textual counterfactual probes that remove, substitute, and reorder skill pairs, aggregates the evidence with Bayesian smoothing, and publishes a state-filtered weighted graph for task-conditioned expansion. The graph is constructed offline and used without changing the downstream agent policy or task interface. Across six LLM backbones on ALFWorld ID-140 and ScienceWorld U211, CaSKG achieves the highest task score in all twelve combinations of model and benchmark. Relative to Graph-of-Skills (GoS), it improves the six-model macro-average ScienceWorld score from 72.62 to 80.50 and ALFWorld success from 80.01\% to 86.79\%, while reducing mean environment steps on both benchmarks. Qualitative and ablation analyses further show that calibrated edges help retrieval preserve prerequisites, state-changing actions, verification routines, and final completion steps. These results position edge-confidence calibration as an effective route to compact and executable skill retrieval at scale\footnote{Code is available at: https://github.com/ZhiyuanLi218/Caskg }.

cs.AI

All-Optical Wide-Field Magnetometry with Van Der Waals Quantum Sensor

Negatively charged boron vacancy ($V_B^-$) centers in hexagonal boron nitride ($h$-BN) have attracted wide-range interests owing to their van der Waals lattice and their potentials for $in$-$situ$ quantum sensing. Here we propose and experimentally demonstrate an all-optical strategy for wide-field magnetometry based on $V_B^-$ centers. This strategy exploits the magnetically sensitive ground-state level anti-crossing (GSLAC) of $V_B^-$ centers, which induces a strong electron spin transition between $m_S = 0$ and $m_S = -1$ states, enabling microwave-free magnetic field measurement. By monitoring the shift of GSLAC feature, the external magnetic field can be precisely determined. Using this technique, we demonstrate all-optical wide-field imaging of near-field DC magnetic field distribution from current-carrying circuits over an area of around 42 $\times$ 21 $\mu$m$^2$. An estimated photon shot-noise-limited sensitivity of 67.1 $\mu$T/$\sqrt{\text{Hz}}$ is achieved for a single pixel, which is an approximately threefold improvement over the ODMR method, along with a spatial resolution of about 1 $\mu$m per pixel. Our approach expands the applicability of $V_B^-$ centers in quantum sensing, paving the way for robust and convenient magnetometry under extreme conditions.

quant-ph

Photon-Atom Granularity Noise Thermometry

We propose granularity noise thermometry (GNT), a fluctuation-based optical thermometry scheme that exploits the intrinsic fluctuations of susceptibility arising from atomic discreteness. The power spectral density of transmitted light exhibits an excess noise above the shot-noise limit that scales linearly with the photon-to-atom ratio $\mathcal{R}$. Consequently, varying the incident power (hence $\mathcal{R}$) yields the slope $\mathcal{K}$ of this linear scaling, which directly encodes the temperature. Closed-form expressions for the polarizability moments are derived via the plasma dispersion function, which yield distinct temperature scalings: $\mathcal{K}\propto P_{\mathrm{v}}(T)/T^2$ for thermal vapors and $\mathcal{K}\propto T^{2}$ for cold atoms. While practical implementation requires careful control of technical noise and system parameters, the present framework provides a noise-based pathway for optical thermometry using atomic ensembles.

physics.atom-ph

Meter-long broadband chirped Bragg gratings for on-chip dispersion control and pulse shaping

Precise on-chip dispersion control is essential for advanced integrated photonic technologies, enabling applications ranging from high-speed communications and sensing to signal processing and biomedical imaging. However, existing on-chip dispersion control methods still suffer from substantial loss and a limited dispersion-bandwidth product (DBP) far from application needs. As a result, on-chip systems continue to rely exclusively on off-chip dispersion control solutions provided by optical fiber or bulky free-space optics. To overcome these limitations, we design and fabricate meter-long chirped spiral Bragg gratings (CSBGs) on the ultra-low-loss silicon nitride (SiN) photonic platform for advanced dispersion control. Our device achieves a 10-nanosecond group delay with customizable bandwidths exceeding 10 nanometers within a compact footprint of only 30 $\text {mm} ^2$, surpassing the physical limits of fiber-based grating devices. More importantly, CSBGs can simultaneously possess the characteristics of high stability, low latency, and a large DBP, thanks to the ultra-low-loss SiN platform with a loss of only 0.3 dB/m. Leveraging the precise and stable dispersion profile, we demonstrate high-fidelity pulse shaping and compression of electro-optic frequency combs (EOCs) with a 1-GHz repetition rate centered across the entire reflection bandwidth. The compressed pulse has an on-chip peak (average) power of 21.6 watts (580 milliwatts). Furthermore, we showcase for the first time the application of on-chip pulse-compressed EOC in wavelength-swept coherent anti-Stokes Raman scattering (CARS) microscopy. Our work provides integrated photonics with a long-sought, scalable, and robust solution for high-performance on-chip dispersion control, empowering a new generation of on-chip functionalities.

physics.optics

MatryoshkaThinking: Recursive Test-Time Scaling Enables Efficient Reasoning

Test-time scaling has emerged as a promising paradigm in language modeling, wherein additional computational resources are allocated during inference to enhance model performance. Recent approaches, such as DeepConf, have demonstrated the efficacy of this strategy, however, they often incur substantial computational overhead to achieve competitive results. In this work, we propose MatryoshkaThinking, a novel method that significantly reduces computational cost while maintaining state-of-the-art performance. Specifically, MatryoshkaThinking attains a score of 99.79 on AIME2025 using only 4% of the computation required by DeepConf. The core of our approach lies in the recursive exploitation of the model's intrinsic capabilities in reasoning, verification, and summarization, which collectively enhance the retention of correct solutions and reduce the disparity between Pass@k and Pass@1. Comprehensive evaluations across multiple open-source models and challenging multi-modal reasoning benchmarks validate the effectiveness and generality of our method. These findings offer new insights into the design of efficient and scalable test-time inference strategies for advanced language models.

cs.CL

Host dependence of PL5 ensemble in 4H-SiC

Color center PL5 in 4H silicon carbide (4H-SiC) has drawn significant attention due to its room-temperature quantum coherence properties and promising potential of quantum sensing applications. The preparation of PL5 ensemble is a critical prerequisite for practical applications. In this work, we investigated the formation of PL5 ensembles in types of 4H-SiC wafers, focusing on their suitability as hosts for PL5 ensemble. Results demonstrate that PL5 signals are exclusively observed in high-purity semi-insulating (HPSI) substrates, whereas divacancies PL1-PL4 can be detected in both HPSI and epitaxial samples. The type of in-plane stress in HPSI and epitaxial hosts is compressive in the same order of magnitude. Defects like stacking faults and dislocations are not observed simultaneously in the PL5 ensemble. Notably, the PL5 ensemble exhibits a relatively uniform distribution in the HPSI host, highlighting its readiness for integration into quantum sensing platforms. Furthermore, signal of PL5 can always be detected in the HPSI samples with different doses of electron irradiation, which suggests that HPSI wafers are more suitable hosts for the production of PL5 ensemble. This work provides critical insights into the material-specific requirements for PL5 ensemble formation and advances the development of 4H-SiC-based quantum technologies.

cond-mat.mtrl-sci

Evidence for field induced quantum spin liquid behavior in a spin-1/2 honeycomb magnet

One of the most important issues in modern condensed matter physics is the realization of fractionalized excitations, such as the Majorana excitations in the Kitaev quantum spin liquid. To this aim, the 3d-based Kitaev material Na2Co2TeO6 is a promising candidate whose magnetic phase diagram of B // a* contains a field-induced intermediate magnetically disordered phase within 7.5 T < |B| < 10 T. The experimental observations, including the restoration of the crystalline point group symmetry in the angle-dependent torque and the coexisting magnon excitations and spinon-continuum in the inelastic neutron scattering spectrum, provide strong evidence that this disordered phase is a field induced quantum spin liquid with partially polarized spins. Our variational Monte Carlo simulation with the effective K-J1-{\Gamma}-{\Gamma}'-J3 model reproduces the experimental data and further supports this conclusion.

cond-mat.str-el

Fiber Transmission Model with Parameterized Inputs based on GPT-PINN Neural Network

In this manuscript, a novelty principle driven fiber transmission model for short-distance transmission with parameterized inputs is put forward. By taking into the account of the previously proposed principle driven fiber model, the reduced basis expansion method and transforming the parameterized inputs into parameterized coefficients of the Nonlinear Schrodinger Equations, universal solutions with respect to inputs corresponding to different bit rates can all be obtained without the need of re-training the whole model. This model, once adopted, can have prominent advantages in both computation efficiency and physical background. Besides, this model can still be effectively trained without the needs of transmitted signals collected in advance. Tasks of on-off keying signals with bit rates ranging from 2Gbps to 50Gbps are adopted to demonstrate the fidelity of the model.

cs.AI

Principle Driven Parameterized Fiber Model based on GPT-PINN Neural Network

In cater the need of Beyond 5G communications, large numbers of data driven artificial intelligence based fiber models has been put forward as to utilize artificial intelligence's regression ability to predict pulse evolution in fiber transmission at a much faster speed compared with the traditional split step Fourier method. In order to increase the physical interpretabiliy, principle driven fiber models have been proposed which inserts the Nonlinear Schodinger Equation into their loss functions. However, regardless of either principle driven or data driven models, they need to be re-trained the whole model under different transmission conditions. Unfortunately, this situation can be unavoidable when conducting the fiber communication optimization work. If the scale of different transmission conditions is large, then the whole model needs to be retrained large numbers of time with relatively large scale of parameters which may consume higher time costs. Computing efficiency will be dragged down as well. In order to address this problem, we propose the principle driven parameterized fiber model in this manuscript. This model breaks down the predicted NLSE solution with respect to one set of transmission condition into the linear combination of several eigen solutions which were outputted by each pre-trained principle driven fiber model via the reduced basis method. Therefore, the model can greatly alleviate the heavy burden of re-training since only the linear combination coefficients need to be found when changing the transmission condition. Not only strong physical interpretability can the model posses, but also higher computing efficiency can be obtained. Under the demonstration, the model's computational complexity is 0.0113% of split step Fourier method and 1% of the previously proposed principle driven fiber model.

cs.AI

Fiber neural networks for the intelligent optical fiber communications

Optical neural networks have long cast attention nowadays. Like other optical structured neural networks, fiber neural networks which utilize the mechanism of light transmission to compute can take great advantages in both computing efficiency and power cost. Though the potential ability of optical fiber was demonstrated via the establishing of fiber neural networks, it will be of great significance of combining both fiber transmission and computing functions so as to cater the needs of future beyond 5G intelligent communication signal processing. Thus, in this letter, the fiber neural networks and their related optical signal processing methods will be both developed. In this way, information derived from the transmitted signals can be directly processed in the optical domain rather than being converted to the electronic domain. As a result, both prominent gains in processing efficiency and power cost can be further obtained. The fidelity of the whole structure and related methods is demonstrated by the task of modulation format recognition which plays important role in fiber optical communications without losing the generality.

eess.SP

Data-free Multi-label Image Recognition via LLM-powered Prompt Tuning

This paper proposes a novel framework for multi-label image recognition without any training data, called data-free framework, which uses knowledge of pre-trained Large Language Model (LLM) to learn prompts to adapt pretrained Vision-Language Model (VLM) like CLIP to multilabel classification. Through asking LLM by well-designed questions, we acquire comprehensive knowledge about characteristics and contexts of objects, which provides valuable text descriptions for learning prompts. Then we propose a hierarchical prompt learning method by taking the multi-label dependency into consideration, wherein a subset of category-specific prompt tokens are shared when the corresponding objects exhibit similar attributes or are more likely to co-occur. Benefiting from the remarkable alignment between visual and linguistic semantics of CLIP, the hierarchical prompts learned from text descriptions are applied to perform classification of images during inference. Our framework presents a new way to explore the synergies between multiple pre-trained models for novel category recognition. Extensive experiments on three public datasets (MS-COCO, VOC2007, and NUS-WIDE) demonstrate that our method achieves better results than the state-of-the-art methods, especially outperforming the zero-shot multi-label recognition methods by 4.7% in mAP on MS-COCO.

cs.CV

REPOFUSE: Repository-Level Code Completion with Fused Dual Context

The success of language models in code assistance has spurred the proposal of repository-level code completion as a means to enhance prediction accuracy, utilizing the context from the entire codebase. However, this amplified context can inadvertently increase inference latency, potentially undermining the developer experience and deterring tool adoption - a challenge we termed the Context-Latency Conundrum. This paper introduces REPOFUSE, a pioneering solution designed to enhance repository-level code completion without the latency trade-off. REPOFUSE uniquely fuses two types of context: the analogy context, rooted in code analogies, and the rationale context, which encompasses in-depth semantic relationships. We propose a novel rank truncated generation (RTG) technique that efficiently condenses these contexts into prompts with restricted size. This enables REPOFUSE to deliver precise code completions while maintaining inference efficiency. Through testing with the CrossCodeEval suite, REPOFUSE has demonstrated a significant leap over existing models, achieving a 40.90% to 59.75% increase in exact match (EM) accuracy for code completions and a 26.8% enhancement in inference speed. Beyond experimental validation, REPOFUSE has been integrated into the workflow of a large enterprise, where it actively supports various coding tasks.

cs.SE

Code-Based English Models Surprising Performance on Chinese QA Pair Extraction Task

In previous studies, code-based models have consistently outperformed text-based models in reasoning-intensive scenarios. When generating our knowledge base for Retrieval-Augmented Generation (RAG), we observed that code-based models also perform exceptionally well in Chinese QA Pair Extraction task. Further, our experiments and the metrics we designed discovered that code-based models containing a certain amount of Chinese data achieve even better performance. Additionally, the capabilities of code-based English models in specified Chinese tasks offer a distinct perspective for discussion on the philosophical "Chinese Room" thought experiment.

cs.CL

3D Heisenberg universality in the Van der Waals antiferromagnet NiPS$_3$

Van der Waals (vdW) magnetic materials are comprised of layers of atomically thin sheets, making them ideal platforms for studying magnetism at the two-dimensional (2D) limit. These materials are at the center of a host of novel types of experiments, however, there are notably few pathways to directly probe their magnetic structure. We report the magnetic order within a single crystal of NiPS$_3$ and show it can be accessed with resonant elastic X-ray diffraction along the edge of the vdW planes in a carefully grown crystal by detecting structurally forbidden resonant magnetic X-ray scattering. We find the magnetic order parameter has a critical exponent of $\beta\sim0.36$, indicating that the magnetism of these vdW crystals is more adequately characterized by the three-dimensional (3D) Heisenberg universality class. We verify these findings with first-principle density functional theory, Monte-Carlo simulations, and density matrix renormalization group calculations.

cond-mat.str-el

CodeFuse-13B: A Pretrained Multi-lingual Code Large Language Model

Code Large Language Models (Code LLMs) have gained significant attention in the industry due to their wide applications in the full lifecycle of software engineering. However, the effectiveness of existing models in understanding non-English inputs for multi-lingual code-related tasks is still far from well studied. This paper introduces CodeFuse-13B, an open-sourced pre-trained code LLM. It is specifically designed for code-related tasks with both English and Chinese prompts and supports over 40 programming languages. CodeFuse achieves its effectiveness by utilizing a high quality pre-training dataset that is carefully filtered by program analyzers and optimized during the training process. Extensive experiments are conducted using real-world usage scenarios, the industry-standard benchmark HumanEval-x, and the specially designed CodeFuseEval for Chinese prompts. To assess the effectiveness of CodeFuse, we actively collected valuable human feedback from the AntGroup's software development process where CodeFuse has been successfully deployed. The results demonstrate that CodeFuse-13B achieves a HumanEval pass@1 score of 37.10%, positioning it as one of the top multi-lingual code LLMs with similar parameter sizes. In practical scenarios, such as code generation, code translation, code comments, and testcase generation, CodeFuse performs better than other models when confronted with Chinese prompts.

cs.SE

Kernel Fusion in Atomistic Spin Dynamics Simulations on Nvidia GPUs using Tensor Core

In atomistic spin dynamics simulations, the time cost of constructing the space- and time-displaced pair correlation function in real space increases quadratically as the number of spins $N$, leading to significant computational effort. The GEMM subroutine can be adopted to accelerate the calculation of the dynamical spin-spin correlation function, but the computational cost of simulating large spin systems ($>40000$ spins) on CPUs remains expensive. In this work, we perform the simulation on the graphics processing unit (GPU), a hardware solution widely used as an accelerator for scientific computing and deep learning. We show that GPUs can accelerate the simulation up to 25-fold compared to multi-core CPUs when using the GEMM subroutine on both. To hide memory latency, we fuse the element-wise operation into the GEMM kernel using $\mathtt{CUTLASS}$ that can improve the performance by 26% $\sim$ 33% compared to implementation based on $\mathtt{cuBLAS}$. Furthermore, we perform the on-the-fly calculation in the epilogue of the GEMM subroutine to avoid saving intermediate results on global memory, which makes the large-scale atomistic spin dynamics simulation feasible and affordable.

physics.comp-ph

Static and dynamical properties of the spin-5/2 nearly ideal triangular lattice antiferromagnet Ba3MnSb2O9

We study the ground state and spin excitations in Ba3MnSb2O9, an easy-plane S = 5/2 triangular lattice antiferromagnet. By combining single-crystal neutron scattering, electric spin resonance (ESR), and spin wave calculations, we determine the frustrated quasi-two-dimensional spin Hamiltonian parameters describing the material. While the material has a slight monoclinic structural distortion, which could allow for isosceles-triangular exchanges and biaxial anisotropy by symmetry, we observe no deviation from the behavior expected for spin waves in the in-plane 120o state. Even the easy-plane anisotropy is so small that it can only be detected by ESR in our study. In conjunction with the quasi-two-dimensionality, our study establishes that Ba3MnSb2O9 is a nearly ideal triangular lattice antiferromagnet with the quasi-classical spin S = 5/2, which suggests that it has the potential for an experimental study of Z- or Z2-vortex excitations.

cond-mat.str-el

On Ultrafast X-ray Methods for Magnetism

With the introduction of x-ray free electron laser sources around the world, new scientific approaches for visualizing matter at fundamental length and time-scales have become possible. As it relates to magnetism and "magnetic-type" systems, advanced methods are being developed for studying ultrafast magnetic responses on the time-scales at which they occur. We describe three capabilities which have the potential to seed new directions in this area and present original results from each: pump-probe x-ray scattering with low energy excitation, x-ray photon fluctuation spectroscopy, and ultrafast diffuse x-ray scattering. By combining these experimental techniques with advanced modeling together with machine learning, we describe how the combination of these domains allows for a new understanding in the field of magnetism. Finally, we give an outlook for future areas of investigation and the newly developed instruments which will take us there.

cond-mat.mtrl-sci