SearcharxivSearch

arXiv subjects

Kang Xu

Publications and source records attributed to Kang Xu.

At least 19 recordsLinked to original sources

VERDICT: Agreement Beats Pixel-Space Verification in Real-Document OCSR

Optical Chemical Structure Recognition (OCSR) converts 2D molecular depictions in the published literature into SMILES, and is increasingly important for constructing large-scale chemical training datasets. Automation at that scale requires identifying unreliable predictions in the absence of ground truth. Three families of label-free signals were compared on $263$ ACS journal depictions with verified ground truth: model confidence, re-rendering similarity, and agreement among recognizers. Pixel-space re-rendering performed little better than chance (AUROC $0.547$, $95\%$ CI $[0.465,0.629]$), and an oracle-tuned threshold on it reduced correct labels per image from $0.745$ to $0.205$. Agreement among four architecturally distinct recognizers instead reached an AUROC of $0.916$ ($[0.880,0.952]$). The two-of-four rule accepted $81.7\%$ of images at $88.8\%$ precision, the three-of-four rule $52.1\%$ at $98.5\%$. The same pattern held on CLEF-IP, UOB, and USPTO. This distinction is obscured on synthetic benchmarks, where re-rendered predictions naturally resemble their inputs. A substance filter removed $2{,}193$ false agreements on wildcards and R-group fragments, after which the three-of-four rule rejected all $68$ generic depictions. VERDICT was then applied to PMC Open Access, producing $6{,}146$ structure labels for $4{,}833$ molecules; chemist adjudication of $400$ released labels in two independent samples yielded precisions of $0.995$ for the three-of-four tier and $0.958$ for the two-of-four tier. VERDICT therefore enables validated labels for multimodal molecular databases linking structure images, machine-readable representations, and source-publication information. In SES AI's Molecular Universe platform, VERDICT further serves as an image-based interface for searching and retrieving molecular records.

cs.CV

Real Data Closes Synthetic-to-Real Gap in Optical Chemical Structure Recognition

Millions of chemical structures appear in patents and papers only as drawings, and using that information at scale requires reading the drawings. OCSR appears nearly solved on synthetic images yet remains difficult on real documents: the starting recognizer, Qwen2.5-VL-7B, exceeds 91% accuracy on synthetic renders but falls below 16% on three real-world benchmarks (ACS, CLEF-IP, USPTO). To identify the main source of improvement, 21 recognizers were fine-tuned on mixtures of synthetically rendered structures and labeled real depictions from patents, journal figures, and hand-drawn collections, varying the vision language model (VLM) base, the fraction of real training data, and the vision-tower adaptation strategy. Labeled real training images make the largest difference. For Qwen2.5-VL, ACS exact match rises from 0.15 with no real data to 0.37 at 9.5% and 0.46 at 50.2%; a controlled experiment across three base models reproduces the trend. A vision-tower LoRA, in contrast, does nothing for Qwen (+0.00, paired p=1.00), substantially helps InternVL3-8B (+22.8 to +34.6 pt), and modestly helps GLM-4.1V-9B (+1.0 to +9.6 pt), so its value depends on the base model. The best configuration reaches 0.96 exact match on clean renders and 0.49, 0.65, 0.84, and 0.76 on ACS, CLEF-IP, UOB, and USPTO, respectively. Gaps between base models are largest without real data (0.21), shrink to 0.06 at 70% real data, and reorder the ranking; base model and real-data mixture must therefore be selected together. Small-scale experiments on handwritten image-to-LaTeX recognition and chart-to-table conversion show that base-model rankings also vary beyond chemistry. More generally, model and adaptation choices for visual structure recognition should be evaluated on the target task.

cs.LG

Predictive Simulation of Interphases on Li Metal Surface

Interphases remain the least understood components in advanced batteries. Although their properties dictate whether a new battery chemistry could perform as designed, there has never been a reliable way to predict what an interphase could arise from a new electrolyte system due to the absence of atomistic level knowledge about interphasial formation process. In this work, we attempt to develop a simulation method that can universally predict interphasial chemistries formed on Li metal surface, so that the electrolyte engineering would no longer need lengthy Edisonian approaches. By combining a transferable universal polarizable force field and a universal machine learning force field, we simulate interphasial chemistry across chemically diverse electrolyte formulations, and successfully replicate the experimental observation that fluorinated solvents promote the formation of LiF-rich interphases, whereas interphases of more organic origin arise from conventional carbonate-based electrolytes. By directly capturing these spontaneous interfacial reactions behind these interphasial chemistries, our simulations establish molecular-level relationships between electrolyte chemistry, salt concentration, decomposition pathways, and SEI properties, and opens a route toward universal and high-throughput predictive simulation of interphases that is the foundation for AI-driven electrolyte discoveries.

cond-mat.mtrl-sci

Observation of long-lived spin order in nanoconfined water

Liquids confined to nanometer-scale geometries exhibit behavior that departs markedly from their bulk counterparts, yet studying their dynamics under controlled conditions remains experimentally challenging. Here, we use nitrogen-vacancy (NV) center nuclear magnetic resonance (NMR) spectroscopy to probe water confined in 5.6 nm channels as a function of temperature. The system remains liquid throughout the investigated temperature range and exhibits strongly suppressed diffusivity, enabling direct detection of its 1H NMR spectrum. Occasionally, the proton resonance transforms into a doublet with a splitting of several tens of kilohertz, which we tentatively attribute to hyperfine interactions mediated by long-lived paramagnetic charge complexes, in turn seeded by solvated electrons optically injected during laser illumination. The intermittent appearance of this feature suggests a metastable state comprising a correlated population of charge-hydration complexes extending throughout the confined liquid.

cond-mat.mes-hall

A High-Precision Frequency Locking Method Based on All-Phase FFT Demonstrated on a Crystal Oscillator with Rubidium Clock Reference

This article proposes a novel frequency-locking method based on frequency-domain unbiased phase estimation (FDUPE) for high-precision frequency control. By performing weighted recombination of the acquired data followed by Fourier-transform processing, the phase at the center of the data segment can be estimated without bias, making the method suitable for frequency-locking applications. The principle of the proposed method is analyzed, and an electronic prototype is developed to experimentally validate its feasibility. In the prototype, analog-to-digital converters (ADCs) are used for signal digitization, and a field-programmable gate array (FPGA) is used to implement the FDUPE algorithm. A digital proportional-integral-derivative (PID) controller is also implemented on the FPGA to provide feedback for accurate frequency locking. In the experiment, a (10~\mathrm{MHz}) voltage-controlled oscillator (VCO) with a free-running Allan deviation of (1 \times 10^{-9}) at (1~\mathrm{s}) is used as the device under test (DUT), while a rubidium atomic clock with an Allan deviation of (2 \times 10^{-11}) at (1~\mathrm{s}) serves as the high-stability reference source. Experimental results show that the proposed system achieves excellent locking performance, reducing the standard deviation of frequency fluctuations from (12.75~\mathrm{mHz}) root-mean-square (rms) in the free-running state to (0.88~\mu\mathrm{Hz}) rms after locking. Correspondingly, the Allan deviation at (10~\mathrm{s}) is reduced from (9.6 \times 10^{-10}) to (1.45 \times 10^{-14}), representing a five-order-of-magnitude improvement in frequency stability.

physics.ins-det

The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence

We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The flagship M2 contains 229.9B total parameters with only 9.8B activated per token. Designed end-to-end for agentic deployment, the M2 series rests on three components: (i) agent-driven data pipelines producing large-scale, verifiable trajectories across agentic coding and agentic cowork, each grounded in an executable workspace and an artifact-aligned reward; (ii) Forge, a scalable agent-native RL system that adapts to long-horizon agent trajectories, paired with windowed-FIFO scheduling, prefix-tree merging, inference optimization, and a clean training-inference-agent decoupling that supports both white-box and black-box agents; (iii) the latest M2.7 checkpoint takes an early step toward self-evolution -- autonomously debugging training runs and modifying its own scaffold. Across M2 through M2.7, this combination translates a mini-activation footprint into frontier-tier performance on agentic coding, deep search, office-task, and reasoning benchmarks.

cs.AI

Wave packet landscape in open quantum systems

We formulate a landscape theory for the long-time wave packet spreading of free and harmonically trapped particles with quantum fluctuations and its related dissipation. We show that the diffusion, localization, and collapse of wave packets arise from symmetry structures of an underlying landscape in covariance space. The geometry of this landscape determines the asymptotic fate of the wave packet. In the quantum landscape description, the trapping potential and bath fluctuation break the landscape symmetry in distinct ways: the former lifts the valley-like landscape of a fluctuation-free free particle into a bowl-like landscape, leading to collapse, whereas the latter tilts the valley and turns localization into diffusion. The resulting landscape symmetry breaking accounts for the noncommuting long-time limits and abrupt changes in the asymptotic wave-packet width. This establishes landscape symmetry breaking as a unified geometric origin of wave-packet diffusion, localization, and collapse in quantum Brownian motion.

quant-ph

Erbium Probes of Magnetic Order in a Layered van der Waals Material

There is growing interest in characterizing magnetic order and dynamics in two-dimensional magnets, yet most efforts to date rely on external probes that interrogate the sample from tens of nanometers away and inevitably average over that length scale. Here we use internal, lattice-embedded Er3+ defects in CrSBr as atomic-scale probes, accessing their telecom-band photoluminescence with spectroscopy and temperature-dependent confocal imaging to read out magnetism from within the material. At room temperature we observe narrow, long-lived photoluminescence (PL) lines in the telecom band, characteristic of erbium emitters. Upon cooling to 3 K and reheating, the Er3+ PL intensity and excited-state lifetime display pronounced thermal hysteresis with a minimum near 132 K, at the reported antiferromagnetic (AFM) transition of CrSBr. Remarkably, we observe magnetic signatures persisting over a broader temperature range than expected from bulk benchmarks, suggesting nanoscale magnetic order that locally survives beyond the nominal phase boundary. Further, a moderate in-plane field of 0.3 T shifts the PL minimum by +8 K, which we tentatively associate to field-biased ferromagnetic correlations.

cond-mat.mes-hall

MicLog: Towards Accurate and Efficient LLM-based Log Parsing via Progressive Meta In-Context Learning

Log parsing converts semi-structured logs into structured templates, forming a critical foundation for downstream analysis. Traditional syntax and semantic-based parsers often struggle with semantic variations in evolving logs and data scarcity stemming from their limited domain coverage. Recent large language model (LLM)-based parsers leverage in-context learning (ICL) to extract semantics from examples, demonstrating superior accuracy. However, LLM-based parsers face two main challenges: 1) underutilization of ICL capabilities, particularly in dynamic example selection and cross-domain generalization, leading to inconsistent performance; 2) time-consuming and costly LLM querying. To address these challenges, we present MicLog, the first progressive meta in-context learning (ProgMeta-ICL) log parsing framework that combines meta-learning with ICL on small open-source LLMs (i.e., Qwen-2.5-3B). Specifically, MicLog: i) enhances LLMs' ICL capability through a zero-shot to k-shot ProgMeta-ICL paradigm, employing weighted DBSCAN candidate sampling and enhanced BM25 demonstration selection; ii) accelerates parsing via a multi-level pre-query cache that dynamically matches and refines recently parsed templates. Evaluated on Loghub-2.0, MicLog achieves 10.3% higher parsing accuracy than the state-of-the-art parser while reducing parsing time by 42.4%.

cs.SE

Probing interfacial water via color-center-enabled spin magnetometry

Understanding the behavior of confined water at liquid-solid interfaces is central to numerous physical, chemical, and biological processes, yet remains experimentally challenging. Here, we utilize shallow nitrogen-vacancy (NV) centers in diamond to investigate the nanoscale dynamics of interfacial water confined between the diamond surface and an overlying fluorinated oil droplet. Using NV-based nuclear magnetic resonance protocols selectively sensitive to 1H and 19F, we independently track water and oil near the interface under ambient conditions. Comparing opposite sides of a doubly-implanted diamond membrane - one exposed to oil, the other not - we uncover a slow, multi-day process in which the interfacial water layer is gradually depleted. This desorption appears to be driven by sustained interactions with the fluorinated oil and is supported by molecular dynamics simulations and surface-sensitive X-ray spectroscopies. Our findings provide molecular-level insight into long-timescale hydration dynamics and underscore the power of NV-NMR for probing liquid-solid heterointerfaces with chemical specificity.

physics.chem-ph

Unsupervised Skill Discovery through Skill Regions Differentiation

Unsupervised Reinforcement Learning (RL) aims to discover diverse behaviors that can accelerate the learning of downstream tasks. Previous methods typically focus on entropy-based exploration or empowerment-driven skill learning. However, entropy-based exploration struggles in large-scale state spaces (e.g., images), and empowerment-based methods with Mutual Information (MI) estimations have limitations in state exploration. To address these challenges, we propose a novel skill discovery objective that maximizes the deviation of the state density of one skill from the explored regions of other skills, encouraging inter-skill state diversity similar to the initial MI objective. For state-density estimation, we construct a novel conditional autoencoder with soft modularization for different skill policies in high-dimensional space. Meanwhile, to incentivize intra-skill exploration, we formulate an intrinsic reward based on the learned autoencoder that resembles count-based exploration in a compact latent space. Through extensive experiments in challenging state and image-based tasks, we find our method learns meaningful skills and achieves superior performance in various downstream tasks.

cs.LG

3D Gaussian Inverse Rendering with Approximated Global Illumination

3D Gaussian Splatting shows great potential in reconstructing photo-realistic 3D scenes. However, these methods typically bake illumination into their representations, limiting their use for physically-based rendering and scene editing. Although recent inverse rendering approaches aim to decompose scenes into material and lighting components, they often rely on simplifying assumptions that fail when editing. We present a novel approach that enables efficient global illumination for 3D Gaussians Splatting through screen-space ray tracing. Our key insight is that a substantial amount of indirect light can be traced back to surfaces visible within the current view frustum. Leveraging this observation, we augment the direct shading computed by 3D Gaussians with Monte-Carlo screen-space ray-tracing to capture one-bounce indirect illumination. In this way, our method enables realistic global illumination without sacrificing the computational efficiency and editability benefits of 3D Gaussians. Through experiments, we show that the screen-space approximation we utilize allows for indirect illumination and supports real-time rendering and editing. Code, data, and models will be made available at our project page: https://wuzirui.github.io/gs-ssr.

cs.GR

OmniScience: A Domain-Specialized LLM for Scientific Reasoning and Discovery

Large Language Models (LLMs) have demonstrated remarkable potential in advancing scientific knowledge and addressing complex challenges. In this work, we introduce OmniScience, a specialized large reasoning model for general science, developed through three key components: (1) domain adaptive pretraining on a carefully curated corpus of scientific literature, (2) instruction tuning on a specialized dataset to guide the model in following domain-specific tasks, and (3) reasoning-based knowledge distillation through fine-tuning to significantly enhance its ability to generate contextually relevant and logically sound responses. We demonstrate the versatility of OmniScience by developing a battery agent that efficiently ranks molecules as potential electrolyte solvents or additives. Comprehensive evaluations reveal that OmniScience is competitive with state-of-the-art large reasoning models on the GPQA Diamond and domain-specific battery benchmarks, while outperforming all public reasoning and non-reasoning models with similar parameter counts. We further demonstrate via ablation experiments that domain adaptive pretraining and reasoning-based knowledge distillation are critical to attain our performance levels, across benchmarks.

cs.AI

TANGO: A Robust Qubit Mapping Algorithm via Two-Stage Search and Bidirectional Look

Current quantum devices typically lack full qubit connectivity, making it difficult to directly execute logical circuits on quantum devices. This limitation necessitates quantum circuit mapping algorithms to insert SWAP gates, dynamically remapping logical qubits to physical qubits and transforming logical circuits into physical circuits that comply with device connectivity constraints. However, the insertion of SWAP gates increases both the gate count and circuit depth, ultimately reducing the fidelity of quantum algorithms. To achieve a balanced optimization of these two objectives, we propose the TANGO algorithm. By incorporating a layer-weight allocation strategy, the algorithm first formulates an evaluation function that balances the impact of qubit mapping on both mapped and unmapped nodes, thereby enhancing the quality of the initial mapping. Next, we design an innovative two-stage routing algorithm that prioritizes the number of executable gates as the primary evaluation metric while also considering quantum gate distance, circuit depth, and a novel bidirectional-look SWAP strategy, which optimizes SWAP gate selection in conjunction with preceding gates, improving the effectiveness of the mapping algorithm. Finally, by integrating advanced quantum gate optimization techniques, the algorithm's overall performance is further enhanced. Experimental results demonstrate that, compared to state-of-the-art methods, the proposed algorithm achieves multi-objective co-optimization of gate count and circuit depth across various benchmarks and quantum devices, exhibiting significant performance advantages.

quant-ph

An Efficient Iterative Algorithm for Qubit Mapping via Layer-Weight Assignment and Search Space Reduction

Current quantum devices support interactions only between physically adjacent qubits, preventing quantum circuits from being directly executed on these devices. Therefore, SWAP gates are required to remap logical qubits to physical qubits, which in turn increases both quantum resource consumption and error rates. To minimize the insertion of additional SWAP gates, we propose HAIL, an efficient iterative qubit mapping algorithm. Leveraging the inherent parallelism in quantum circuits, a new layer-weight assignment method is integrated with subgraph isomorphism to derive an optimal initial qubit mapping. Moreover, we present a two-stage SWAP sequence search algorithm that effectively identifies the most efficient SWAP sequence by distilling feasible SWAP sequences at different stages. The whole qubit mapping algorithm is then refined through a few iterative bidirectional traversals, further reducing the number of SWAP gates required. Experimental results on the IBM Q20 architecture and various benchmarks show that HAIL-3 reduces the number of additional gates inserted in the $\mathcal{B}_{23}$ by 20.62\% compared to state-of-the-art algorithms. Moreover, we propose a partially extended SWAP sequence strategy combined with HAIL to reduce its time complexity, with experiments on the sparsely connected Google Sycamore architecture demonstrating reductions in both algorithm runtime and additional SWAP gates.

quant-ph

Online Preference Alignment for Language Models via Count-based Exploration

Reinforcement Learning from Human Feedback (RLHF) has shown great potential in fine-tuning Large Language Models (LLMs) to align with human preferences. Existing methods perform preference alignment from a fixed dataset, which can be limited in data coverage, and the resulting reward model is hard to generalize in out-of-distribution responses. Thus, online RLHF is more desirable to empower the LLM to explore outside the support of the initial dataset by iteratively collecting the prompt-response pairs. In this paper, we study the fundamental problem in online RLHF, i.e. \emph{how to explore} for LLM. We give a theoretical motivation in linear reward assumption to show that an optimistic reward with an upper confidence bound (UCB) term leads to a provably efficient RLHF policy. Then, we reformulate our objective to direct preference optimization with an exploration term, where the UCB-term can be converted to a count-based exploration bonus. We further propose a practical algorithm, named \emph{Count-based Online Preference Optimization (COPO)}, which leverages a simple coin-flip counting module to estimate the pseudo-count of a prompt-response pair in previously collected data. COPO encourages LLMs to balance exploration and preference optimization in an iterative manner, which enlarges the exploration space and the entire data coverage of iterative LLM policies. We conduct online RLHF experiments on Zephyr and Llama-3 models. The results on instruction-following and standard academic benchmarks show that COPO significantly increases performance.

cs.LG

Slow water in engineered nano-channels revealed by color-center-enabled sensing

Nanoscale confinement of molecules in a fluid can result in enhanced viscosity, local fluidic order, or collective motion. Confinement also affects ion transport and/or the rate and equilibrium concentration in a chemical reaction, all of which makes it the subject of broad interest. Studying these effects, however, is notoriously difficult, mainly due to the lack of experimental methods with the required sensitivity and spatial or time resolution. Here we leverage shallow nitrogen-vacancy (NV) centers in diamond to probe the dynamics of room-temperature water molecules entrapped within ~6-nm-tall channels formed between the diamond crystal and a suspended hexagonal boron nitride (hBN) flake. NV-enabled nuclear magnetic resonance measurements of confined water protons reveal a much reduced H2O self-diffusivity, orders of magnitude lower than in bulk water. We posit the slow dynamics stem from the accumulation of photogenerated carriers at the interface and trapped fluid, a notion we support with the help of molecular dynamics modeling. Our results provide feedback for theories describing interfacial water, and lay out a route for investigating other fluids under confinement.

cond-mat.mes-hall

ODRL: A Benchmark for Off-Dynamics Reinforcement Learning

We consider off-dynamics reinforcement learning (RL) where one needs to transfer policies across different domains with dynamics mismatch. Despite the focus on developing dynamics-aware algorithms, this field is hindered due to the lack of a standard benchmark. To bridge this gap, we introduce ODRL, the first benchmark tailored for evaluating off-dynamics RL methods. ODRL contains four experimental settings where the source and target domains can be either online or offline, and provides diverse tasks and a broad spectrum of dynamics shifts, making it a reliable platform to comprehensively evaluate the agent's adaptation ability to the target domain. Furthermore, ODRL includes recent off-dynamics RL algorithms in a unified framework and introduces some extra baselines for different settings, all implemented in a single-file manner. To unpack the true adaptation capability of existing methods, we conduct extensive benchmarking experiments, which show that no method has universal advantages across varied dynamics shifts. We hope this benchmark can serve as a cornerstone for future research endeavors. Our code is publicly available at https://github.com/OffDynamicsRL/off-dynamics-rl.

cs.LG