SearcharxivSearch

arXiv subjects

Qing Yao

Publications and source records attributed to Qing Yao.

11 recordsLinked to original sources

What's in a Name? Morphological Shortcuts by LLMs in Pharmacology

The morphological form of a word can often give cues to its meaning, but purely relying on these mappings can lead to overgeneralization in high-stakes domains. In the medical domain, for instance, LLMs can confidently reason about fictitious drugs from their affixes alone (e.g., wugcillin) and generate plausible-looking clinical content. We present a behavioral and mechanistic study of LLM "affix heuristics" in pharmacology. Using fictitious drug names built from real affixes, we show that affix signals alone elicit class-level pharmacological responses. We introduce a framework for identifying whether a model's drug semantics are driven mainly by the affix, the stem, or the drug name as a whole. Applied across 653 drugs, our framework reveals that models often induce drug meaning primarily through affix cues, yet rarely explicitly indicate this reliance, and sometimes incorrectly conflate properties among affix-sharing drugs. Activation patching across models further localizes this behavior to early-mid layers. These findings show that morphological shortcuts pose a subtle but measurable risk to safety.

cs.CL

France or Spain or Germany or France: A Neural Account of Non-Redundant Redundant Disjunctions

Sentences like "She will go to France or Spain, or perhaps to Germany or France." appear formally redundant, yet become acceptable in contexts such as "Mary will go to a philosophy program in France or Spain, or a mathematics program in Germany or France." While this phenomenon has typically been analyzed using symbolic formal representations, we aim to provide an account grounded in artificial neural mechanisms. We first present new behavioral evidence from humans and large language models demonstrating the robustness of this apparent non-redundancy across contexts. We then show that, in language models, redundancy avoidance arises from two interacting mechanisms: models learn to bind contextually relevant information to repeated lexical items, and Transformer induction heads selectively attend to these context-licensed representations. We argue that this neural explanation sheds light on the mechanisms underlying context-sensitive semantic interpretation, and that it complements existing symbolic analyses.

cs.CL

Holographic codes seen through ZX-calculus

We re-visit the pentagon holographic quantum error correcting code from a ZX-calculus perspective. By expressing the underlying tensors as ZX-diagrams, we study the stabiliser structure of the code via Pauli webs. In addition, we obtain a diagrammatic understanding of its logical operators, encoding isometries, R\'enyi entropy and toy models of black holes/wormholes. Then, motivated by the pentagon holographic code's ZX-diagram, we introduce a family of codes constructed from ZX-diagrams on its dual hyperbolic tessellations and study their logical error rates using belief propagation decoders. Finally, we show how to construct spacetime ZX-diagrams that realise a fault-tolerant quantum channel in which every internal edge is a protected fault location.

quant-ph

Regularized Schr\"odinger Bridge via Distortion-Perception Perturbation for High-Fidelity Speech Enhancement

Speech enhancement (SE) requires high-fidelity reconstruction of clean speech that preserves linguistic and paralinguistic cues while maintaining high perceptual quality. Recently, Schr\"odinger Bridge (SB), a family of diffusion-based generative models, has advanced SE by bridging degraded and clean speech distributions in a principled formulation, enabling higher-quality reconstructions with fewer sampling steps. However, diffusion-based SE methods still face two challenges: (1) the fidelity-realism tradeoff, where they often prioritize perceptual realism encouraged by the learned speech prior, at the expense of fidelity; and (2) the exposure bias issue, where iterative multi-step sampling causes early-step prediction errors to accumulate along the sampling trajectory and degrade enhanced speech quality. In this paper, we analyze standard SB training and show that it induces a systematic prediction drift, which biases the multi-step trajectory and amplifies error accumulation. To address this, we propose Regularized Schr\"odinger Bridge (RSB) for high-fidelity SE, a generative approach that reconciles fidelity and realism while mitigating exposure bias. RSB regularizes training with a Distortion-Perception Perturbation that constructs time-varying targets by interpolating between clean speech and posterior-mean estimates, and trains the network on perturbed intermediate states to correct toward the ground truth progressively. By simulating inference-time prediction errors, this perturbation mitigates the training-inference mismatch and thereby alleviates exposure bias. It also injects posterior-mean estimates as fidelity-preserving guidance, thereby improving reconstruction fidelity.

cs.LG

Both Direct and Indirect Evidence Contribute to Dative Alternation Preferences in Language Models

Language models (LMs) tend to show human-like preferences on a number of syntactic phenomena, but the extent to which these are attributable to direct exposure to the phenomena or more general properties of language is unclear. We explore this with the English dative alternation (DO: "gave Y the X" vs. PO: "gave the X to Y"), using a controlled rearing paradigm wherein we iteratively train small LMs on systematically manipulated input. We focus on two properties that affect the choice of alternant: length and animacy. Both properties are directly present in datives but also reflect more global tendencies for shorter elements to precede longer ones and animates to precede inanimates. First, by manipulating and ablating datives for these biases in the input, we show that direct evidence of length and animacy matters, but easy-first preferences persist even without such evidence. Then, using LMs trained on systematically perturbed datasets to manipulate global length effects (re-linearizing sentences globally while preserving dependency structure), we find that dative preferences can emerge from indirect evidence. We conclude that LMs' emergent syntactic preferences come from a mix of direct and indirect sources.

cs.CL

Random space-time sampling and reconstruction of sparse bandlimited graph diffusion field

In this work, we investigate the sampling and reconstruction of spectrally $s$-sparse bandlimited graph signals governed by heat diffusion processes. We propose a random space-time sampling regime, referred to as {randomized} dynamical sampling, where a small subset of space-time nodes is randomly selected at each time step based on a probability distribution. To analyze the recovery problem, we establish a rigorous mathematical framework by introducing the parameter \textit{the dynamic spectral graph weighted coherence}. This key parameter governs the number of space-time samples needed for stable recovery and extends the idea of variable density sampling to the context of dynamical systems. By optimizing the sampling probability distribution, we show that as few as $\mathcal{O}(s \log(k))$ space-time samples are sufficient for accurate reconstruction in optimal scenarios, where $k$ denotes the bandwidth of the signal. Our framework encompasses both static and dynamic cases, demonstrating a reduction in the number of spatial samples needed at each time step by exploiting temporal correlations. Furthermore, we provide a computationally efficient and robust algorithm for signal reconstruction. Numerical experiments validate our theoretical results and illustrate the practical efficacy of our proposed methods.

math.NA

Effects of syndication network on specialisation and performance of venture capital firms

The Chinese venture capital (VC) market is a young and rapidly expanding financial subsector. Gaining a deeper understanding of the investment behaviours of VC firms is crucial for the development of a more sustainable and healthier market and economy. Contrasting evidence supports that either specialisation or diversification helps to achieve a better investment performance. However, the impact of the syndication network is overlooked. Syndication network has a great influence on the propagation of information and trust. By exploiting an authoritative VC dataset of thirty-five-year investment information in China, we construct a joint-investment network of VC firms and analyse the effects of syndication and diversification on specialisation and investment performance. There is a clear correlation between the syndication network degree and specialisation level of VC firms, which implies that the well-connected VC firms are diversified. More connections generally bring about more information or other resources, and VC firms are more likely to enter a new stage or industry with some new co-investing VC firms when compared to a randomised null model. Moreover, autocorrelation analysis of both specialisation and success rate on the syndication network indicates that clustering of similar VC firms is roughly limited to the secondary neighbourhood. When analysing local clustering patterns, we discover that, contrary to popular beliefs, there is no apparent successful club of investors. In contrast, investors with low success rates are more likely to cluster. Our discoveries enrich the understanding of VC investment behaviours and can assist policymakers in designing better strategies to promote the development of the VC industry.

physics.soc-ph

Emergence of universal scaling in weather extreme events

The frequency and magnitude of weather extreme events have increased significantly during the past few years in response to anthropogenic climate change. However, global statistical characteristics and underlying physical mechanisms are still not fully understood. Here, we adopt a statistical physics and probability theory based method to investigate the nature of extreme weather events, particularly the statistics of the day-to-day air temperature differences. These statistical measurements reveal that the distributions of the magnitudes of the extreme events satisfy a universal \textit{Gumbel} distribution, while the waiting time of those extreme events is governed by a universal \textit{Gamma} function. Further finite-size effects analysis indicates robust scaling behaviours. We additionally unveil that the cumulative distribution of logarithmic waiting times between the record events follows an \textit{Exponential} distribution and that the evolution of this climate system is directional where the underlying dynamics are related to a decelerating release of tension. The universal scaling laws are remarkably stable and unaffected by global warming. Counterintuitively, unlike as expected for record dynamics, we find that the number of quakes of the extreme temperature variability does not decay as one over time but with deviations relevant to large-scale climate extreme events. Our theoretical framework provides a fresh perspective on the linkage of universality, scaling, and climate systems. The findings throw light on the nature of the weather variabilities and could guide us to better forecast extreme events.

physics.ao-ph

Emergence of community structures through biased random walks rewiring

Community structures have been identified in various complex real-world networks, for example, communication, information, internet and shareholder networks. The scaling of community size distribution indicates the heterogeneity in the topological structures of the network. The current network generating or growing models can reproduce some properties, including degree distributions, large clustering coefficients and communities. However, the scaling behaviour of the community size lacks investigation, especially from the perspectives of local interactions. Based on the assumption that heterogeneous nodes behave differently and result in different topological positions of the networks, we propose a model of designed random walks in directed networks to explain the features in the observed networks. The model highlights that two different dynamics can mimic the local interactions, and a hidden layer is essential when reproducing the characteristics of real complex networks. The key features the model can explain include community size distribution, degree distribution, percolation properties, distribution of average path length and dependence of the above properties on the labels of nodes in the data.

physics.soc-ph

Higher-order temporal network effects through triplet evolution

We study the evolution of networks through `triplets' - three-node graphlets. We develop a method to compute a transition matrix to describe the evolution of triplets in temporal networks. To identify the importance of higher-order interactions in the evolution of networks, we compare both artificial and real-world data to a model based on pairwise interactions only. The significant differences between the computed matrix and the calculated matrix from the fitted parameters demonstrate that non-pairwise interactions exist for various real-world systems in space and time, such as our data sets. Furthermore, this also reveals that different patterns of higher-order interaction are involved in different real-world situations. To test our approach, we then use these transition matrices as the basis of a link prediction algorithm. We investigate our algorithm's performance on four temporal networks, comparing our approach against ten other link prediction methods. Our results show that higher-order interactions in both space and time play a crucial role in the evolution of networks as we find our method, along with two other methods based on non-local interactions, give the best overall performance. The results also confirm the concept that the higher-order interaction patterns, i.e., triplet dynamics, can help us understand and predict the evolution of different real-world systems.

physics.soc-ph

How the network properties of shareholders vary with investor type and country

We construct two examples of shareholder networks in which shareholders are connected if they have shares in the same company. We do this for the shareholders in Turkish companies and we compare this against the network formed from the shareholdings in Dutch companies. We analyse the properties of these two networks in terms of the different types of shareholder. We create a suitable randomised version of these networks to enable us to find significant features in our networks. For that we find the roles played by different types of shareholder in these networks, and also show how these roles differ in the two countries we study.

q-fin.GN