SearcharxivSearch

arXiv subjects

Xin-Yu Zhang

Publications and source records attributed to Xin-Yu Zhang.

14 recordsLinked to original sources

Strongly entangled Quantum Spin Rings driven by Hückel rule

Quantum spin rings represent an intriguing platform for studying unconventional magnetic order and exotic quantum phases, and they are also promising materials for emerging quantum technologies. Conventional spin systems consist of a set of weakly interacting localized spins that are well described by the Heisenberg spin models. Here, we demonstrate that strong interactions between radical centers in macrocycles of different sizes lead to fluctuations in the total number of unpaired electrons and to non-trivial antiferromagnetic order that extends beyond the Heisenberg picture. We demonstrate that the electronic structure of these spin rings is governed by the concept of 4n/4n+2 Hückel (anti)aromaticity for even-membered rings, whereas odd-membered rings possess a highly degenerate frustrated magnetic ground state. The strongly coupled spin rings are experimentally realized through the on-surface synthesis of π-magnetic carbon-based macrocycles, which consist of [2]triangulene units. The close correlation between the electronic structure and the Hückel aromaticity rule is revealed by scanning tunneling spectroscopy and multireference calculations. This work establishes a novel design principle employing the concept of Hückel aromaticity for quantum spin macrocycles.

cond-mat.mes-hall

A generalized rate law for inhomogeneous system and turbulence-chemistry decoupling of reaction rate calculation in combustion

In this work, the rate law for inhomogeneous concentration distributions has been formulated, by applying spatial integration over the products of species concentrations. Reaction rates for typical reactions have been investigated by assuming a linear concentration distribution in the grid. A few examples of one-dimensional concentration distributions, straight line, piecewise, and sine function, for a selected second order reaction have been taken to illustrate the validations of the method developed. Difference between the reaction rates by spatial integration and by mean concentrations have been discussed. It is revealed that the chemical reaction rates for combustion simulation can be calculated by appropriate sub-grid modeling of concentration distributions, without needs of the explicit consideration of turbulent combustion interactions, and the reaction rates for the species transport equation in turbulent combustion simulations can be accurately calculated if the concentration distributions of species within the grid are correctly defined.

physics.chem-ph

Learnware of Language Models: Specialized Small Language Models Can Do Big

The learnware paradigm offers a novel approach to machine learning by enabling users to reuse a set of well-trained models for tasks beyond the models' original purposes. It eliminates the need to build models from scratch, instead relying on specifications (representations of a model's capabilities) to identify and leverage the most suitable models for new tasks. While learnware has proven effective in many scenarios, its application to language models has remained largely unexplored. At the same time, large language models (LLMs) have demonstrated remarkable universal question-answering abilities, yet they face challenges in specialized scenarios due to data scarcity, privacy concerns, and high computational costs, thus more and more specialized small language models (SLMs) are being trained for specific domains. To address these limitations systematically, the learnware paradigm provides a promising solution by enabling maximum utilization of specialized SLMs, and allowing users to identify and reuse them in a collaborative and privacy-preserving manner. This paper presents a preliminary attempt to apply the learnware paradigm to language models. We simulated a learnware system comprising approximately 100 learnwares of specialized SLMs with 8B parameters, fine-tuned across finance, healthcare, and mathematics domains. Each learnware contains an SLM and a specification, which enables users to identify the most relevant models without exposing their own data. Experimental results demonstrate promising performance: by selecting one suitable learnware for each task-specific inference, the system outperforms the base SLMs on all benchmarks. Compared to LLMs, the system outperforms Qwen1.5-110B, Qwen2.5-72B, and Llama3.1-70B-Instruct by at least 14% in finance domain tasks, and surpasses Flan-PaLM-540B (ranked 7th on the Open Medical LLM Leaderboard) in medical domain tasks.

cs.LG

Diverse Inference and Verification for Advanced Reasoning

Reasoning LLMs such as OpenAI o1, o3 and DeepSeek R1 have made significant progress in mathematics and coding, yet find challenging advanced tasks such as International Mathematical Olympiad (IMO) combinatorics problems, Abstraction and Reasoning Corpus (ARC) puzzles, and Humanity's Last Exam (HLE) questions. We use a diverse inference approach that combines multiple models and methods at test time. We find that verifying mathematics and code problems, and rejection sampling on other problems is simple and effective. We automatically verify correctness of solutions to IMO problems by Lean, and ARC puzzles by code, and find that best-of-N effectively answers HLE questions. Our approach increases answer accuracy on IMO combinatorics problems from 33.3% to 77.8%, accuracy on HLE questions from 8% to 37%, and solves 80% of ARC puzzles that 948 humans could not and 26.5% of ARC puzzles that o3 high compute does not. Test-time simulations, reinforcement learning, and meta-learning with inference feedback improve generalization by adapting agent graph representations and varying prompts, code, and datasets. Our approach is reliable, robust, and scalable, and in the spirit of reproducible research, we will make it publicly available upon publication.

cs.AI

AI-Driven Review Systems: Evaluating LLMs in Scalable and Bias-Aware Academic Reviews

Automatic reviewing helps handle a large volume of papers, provides early feedback and quality control, reduces bias, and allows the analysis of trends. We evaluate the alignment of automatic paper reviews with human reviews using an arena of human preferences by pairwise comparisons. Gathering human preference may be time-consuming; therefore, we also use an LLM to automatically evaluate reviews to increase sample efficiency while reducing bias. In addition to evaluating human and LLM preferences among LLM reviews, we fine-tune an LLM to predict human preferences, predicting which reviews humans will prefer in a head-to-head battle between LLMs. We artificially introduce errors into papers and analyze the LLM's responses to identify limitations, use adaptive review questions, meta prompting, role-playing, integrate visual and textual analysis, use venue-specific reviewing materials, and predict human preferences, improving upon the limitations of the traditional review processes. We make the reviews of publicly available arXiv and open-access Nature journal papers available online, along with a free service which helps authors review and revise their research papers and improve their quality. This work develops proof-of-concept LLM reviewing systems that quickly deliver consistent, high-quality reviews and evaluate their quality. We mitigate the risks of misuse, inflated review scores, overconfident ratings, and skewed score distributions by augmenting the LLM with multiple documents, including the review form, reviewer guide, code of ethics and conduct, area chair guidelines, and previous year statistics, by finding which errors and shortcomings of the paper may be detected by automated reviews, and evaluating pairwise reviewer preferences. This work identifies and addresses the limitations of using LLMs as reviewers and evaluators and enhances the quality of the reviewing process.

cs.AI

Atomic-Scale Imaging of Fractional Spinon Quasiparticles in Open-Shell Triangulene Spin-$\frac{1}{2}$ Chains

The emergence of spinon quasiparticles, which carry spin but lack charge, is a hallmark of collective quantum phenomena in low-dimensional quantum spin systems. While the existence of spinons has been demonstrated through scattering spectroscopy in ensemble samples, real-space imaging of these quasiparticles within individual spin chains has remained elusive. In this study, we construct individual Heisenberg antiferromagnetic spin-$\frac{1}{2}$ chains using open-shell [2]triangulene molecules as building blocks. Each [2]triangulene unit, owing to its sublattice imbalance, hosts a net spin-$\frac{1}{2}$ in accordance with Lieb's theorem, and these spins are antiferromagnetically coupled within covalent chains with a coupling strength of $J = 45$ meV. Through scanning tunneling microscopy and spectroscopy, we probe the spin states, excitation gaps, and their spatial excitation weights within covalent spin chains of varying lengths with atomic precision. Our investigation reveals that the excitation gap decreases as the chain length increases, extrapolating to zero for long chains, consistent with Haldane's gapless prediction. Moreover, inelastic tunneling spectroscopy reveals an m-shaped energy dispersion characteristic of confined spinon quasiparticles in a one-dimensional quantum box. These findings establish a promising strategy for exploring the unique properties of excitation quasiparticles and their broad implications for quantum information.

cond-mat.mtrl-sci

Real-time superresolution interferometric measurement enabled by structured nonlinear optics

Optical interferometers are pillars of modern precision metrology, but their resolution is limited by the wavelength of the light source, which cannot be infinitely reduced. Magically, this limitation can be circumvented by using an entangled multiphoton source because interference produced by an N-photon amplitude features a reduced de Broglie wavelength λ/N. However, the extremely low efficiency in multiphoton state generation and coincidence counts actually negates the potential of using multiphoton states in practical measurements. Here, we demonstrate a novel interferometric technique based on structured nonlinear optics, i.e., parametric upconversion of a structured beam, capable of superresolution measurement in real time. The main principle relies in that the orbital angular momentum (OAM) state and associated intramodal phase within the structured beam are both continuously multiplied in cascading upconversion to mimic the superresolved phase evolution of a multiphoton amplitude. Owing to the use of bright sensing beams and OAM mode projection, up to a 12-photon de Broglie wavelength with almost perfect visibility is observed in real time and, importantly, by using only a low-cost detector. Our results open the door to real-time superresolution interferometric metrology and provide a promising way toward multiphoton superiority in practical applications.

physics.optics

On Training Implicit Models

This paper focuses on training implicit models of infinite layers. Specifically, previous works employ implicit differentiation and solve the exact gradient for the backward propagation. However, is it necessary to compute such an exact but expensive gradient for training? In this work, we propose a novel gradient estimate for implicit models, named phantom gradient, that 1) forgoes the costly computation of the exact gradient; and 2) provides an update direction empirically preferable to the implicit model training. We theoretically analyze the condition under which an ascent direction of the loss landscape could be found, and provide two specific instantiations of the phantom gradient based on the damped unrolling and Neumann series. Experiments on large-scale tasks demonstrate that these lightweight phantom gradients significantly accelerate the backward passes in training implicit models by roughly 1.7 times, and even boost the performance over approaches based on the exact gradient on ImageNet.

cs.LG

Structured Sparsification with Joint Optimization of Group Convolution and Channel Shuffle

Recent advances in convolutional neural networks(CNNs) usually come with the expense of excessive computational overhead and memory footprint. Network compression aims to alleviate this issue by training compact models with comparable performance. However, existing compression techniques either entail dedicated expert design or compromise with a moderate performance drop. In this paper, we propose a novel structured sparsification method for efficient network compression. The proposed method automatically induces structured sparsity on the convolutional weights, thereby facilitating the implementation of the compressed model with the highly-optimized group convolution. We further address the problem of inter-group communication with a learnable channel shuffle mechanism. The proposed approach can be easily applied to compress many network architectures with a negligible performance drop. Extensive experimental results and analysis demonstrate that our approach gives a competitive performance against the recent network compression counterparts with a sound accuracy-complexity trade-off.

cs.CV

Semi-Supervised Learning with Meta-Gradient

In this work, we propose a simple yet effective meta-learning algorithm in semi-supervised learning. We notice that most existing consistency-based approaches suffer from overfitting and limited model generalization ability, especially when training with only a small number of labeled data. To alleviate this issue, we propose a learn-to-generalize regularization term by utilizing the label information and optimize the problem in a meta-learning fashion. Specifically, we seek the pseudo labels of the unlabeled data so that the model can generalize well on the labeled data, which is formulated as a nested optimization problem. We address this problem using the meta-gradient that bridges between the pseudo label and the regularization term. In addition, we introduce a simple first-order approximation to avoid computing higher-order derivatives and provide theoretic convergence analysis. Extensive evaluations on the SVHN, CIFAR, and ImageNet datasets demonstrate that the proposed algorithm performs favorably against state-of-the-art methods.

cs.LG

Res2Net: A New Multi-scale Backbone Architecture

Representing features at multiple scales is of great importance for numerous vision tasks. Recent advances in backbone convolutional neural networks (CNNs) continually demonstrate stronger multi-scale representation ability, leading to consistent performance gains on a wide range of applications. However, most existing methods represent the multi-scale features in a layer-wise manner. In this paper, we propose a novel building block for CNNs, namely Res2Net, by constructing hierarchical residual-like connections within one single residual block. The Res2Net represents multi-scale features at a granular level and increases the range of receptive fields for each network layer. The proposed Res2Net block can be plugged into the state-of-the-art backbone CNN models, e.g., ResNet, ResNeXt, and DLA. We evaluate the Res2Net block on all these models and demonstrate consistent performance gains over baseline models on widely-used datasets, e.g., CIFAR-100 and ImageNet. Further ablation studies and experimental results on representative computer vision tasks, i.e., object detection, class activation mapping, and salient object detection, further verify the superiority of the Res2Net over the state-of-the-art baseline methods. The source code and trained models are available on https://mmcheng.net/res2net/.

cs.CV

Learnable Cost Volume Using the Cayley Representation

Cost volume is an essential component of recent deep models for optical flow estimation and is usually constructed by calculating the inner product between two feature vectors. However, the standard inner product in the commonly-used cost volume may limit the representation capacity of flow models because it neglects the correlation among different channel dimensions and weighs each dimension equally. To address this issue, we propose a learnable cost volume (LCV) using an elliptical inner product, which generalizes the standard inner product by a positive definite kernel matrix. To guarantee its positive definiteness, we perform spectral decomposition on the kernel matrix and re-parameterize it via the Cayley representation. The proposed LCV is a lightweight module and can be easily plugged into existing models to replace the vanilla cost volume. Experimental results show that the LCV module not only improves the accuracy of state-of-the-art models on standard benchmarks, but also promotes their robustness against illumination change, noises, and adversarial perturbations of the input signals.

cs.CV

Dependency Aware Filter Pruning

Convolutional neural networks (CNNs) are typically over-parameterized, bringing considerable computational overhead and memory footprint in inference. Pruning a proportion of unimportant filters is an efficient way to mitigate the inference cost. For this purpose, identifying unimportant convolutional filters is the key to effective filter pruning. Previous work prunes filters according to either their weight norms or the corresponding batch-norm scaling factors, while neglecting the sequential dependency between adjacent layers. In this paper, we further develop the norm-based importance estimation by taking the dependency between the adjacent layers into consideration. Besides, we propose a novel mechanism to dynamically control the sparsity-inducing regularization so as to achieve the desired sparsity. In this way, we can identify unimportant filters and search for the optimal network architecture within certain resource budgets in a more principled manner. Comprehensive experimental results demonstrate the proposed method performs favorably against the existing strong baseline on the CIFAR, SVHN, and ImageNet datasets. The training sources will be publicly available after the review process.

cs.CV

AdaSample: Adaptive Sampling of Hard Positives for Descriptor Learning

Triplet loss has been widely employed in a wide range of computer vision tasks, including local descriptor learning. The effectiveness of the triplet loss heavily relies on the triplet selection, in which a common practice is to first sample intra-class patches (positives) from the dataset for batch construction and then mine in-batch negatives to form triplets. For high-informativeness triplet collection, researchers mostly focus on mining hard negatives in the second stage, while paying relatively less attention to constructing informative batches. To alleviate this issue, we propose AdaSample, an adaptive online batch sampler, in this paper. Specifically, hard positives are sampled based on their informativeness. In this way, we formulate a hardness-aware positive mining pipeline within a novel maximum loss minimization training protocol. The efficacy of the proposed method is evaluated on several standard benchmarks, where it demonstrates a significant and consistent performance gain on top of the existing strong baselines.

cs.CV