SearcharxivSearch

arXiv subjects

Yanbo Zhang

Publications and source records attributed to Yanbo Zhang.

At least 19 recordsLinked to original sources

Erd\H{o}s--Ko--Rado theorems in $\ell_2$-norm for three finite spaces

Let $\mathcal{F}$ be a $k$-uniform hypergraph. The famous Erd\H{o}s--Ko--Rado (1961) theorem determines the maximum size and extremal structure for $\mathcal{F}$ being $t$-intersecting, that is, $|F_1 \cap F_2| \ge t$ for any two edges $F_1, F_2$ of $\mathcal F$. The codegree squared sum $\mathrm{co}_2(\mathcal{F})$ is the square of the $\ell_2$-norm of the codegree vector of all $(k-1)$-sets in $\mathcal{F}$, which was initially introduced for Tur\'an problems of hypergraphs. Recently, Brooks and Linz (2026), as well as Wu and Zhang (2026) investigated the maximum value of $\mathrm{co}_2(\mathcal{F})$ and corresponding extremal structures for $\mathcal{F}$ being $t$-intersecting. Moreover, Brooks and Linz asked if the classical results on intersecting families can be extended to $\mathrm{co}_2(\mathcal{F})$. In this paper, by developing the spectral techniques for incidence matrices, we study the extremal problems of $\mathrm{co}_2(\mathcal{F})$ for $\mathcal{F}$ being intersecting families in finite vector spaces, affine spaces, and attenuated spaces, and establish the Erd\H{o}s--Ko--Rado theorems in $\ell_2$-norm for the three finite spaces.

math.CO

A note on tree-cycle Ramsey numbers

Let $R(T_n,C_m)$ denote the Ramsey number of a tree $T_n$ on $n$ vertices versus a cycle $C_m$ of length $m$. Burr, Erd\H{o}s, Faudree, Rousseau, and Schelp (1982) asked for the least function $f(m)$ such that $R(T_n,C_m)=2n-1$ for every odd $m\ge 3$ whenever $n\ge f(m)$. They proved that $f(m)\le 756m^{10}$. This bound was later improved to $25m$ by Brennan (2016) and to $4m-8$ by Fan and Lin (2025). In this note, we show that $f(m)\le 2m-4$ by using a different method and conjecture that $f(m)=\lceil (2m-1)/3\rceil$.

math.CO

Intelligence from Learnable Novelty

Intelligence appears under different names in different fields: as data compression in statistics and machine learning, as universal computation in dynamical systems, and as adaptive behavior in agents. Each field carries its own objective, and the two most influential drives often fail in mirror image: novelty search, which seeks surprise, is transfixed by a noisy television screen, while the free-energy principle, which avoids surprise, is most content in a dark room. Both failures have a single cause: each objective treats as one quantity the surprise a learner can convert into knowledge and the surprise it never can. Here we show that the learnable part of that information, which we call learnable novelty, yields the seemingly disparate projections of intelligence, and we give a closed-form estimator of it built on a cheap and differentiable reservoir computer. Used as a measure, with no supervision of any kind, the estimator recovers decades of complexity classification, ranking the Turing-complete rule~110 highest among the elementary cellular automata. Used as an objective, its gradient carries a neural cellular automaton from simple dynamics into a regime of solitons, the traveling, colliding structures by which rule~110 computes, as well as organizes the representation of an image encoder around the ten digit classes of MNIST, fully unsupervised: no label ever enters training. Handed to a reinforcement-learning agent as an intrinsic reward, it supplies the exploration that task rewards lack, improving on the task baseline in nine of ten environments and collapsing in none. Complexity generation, abstraction, and exploration, ordinarily pursued with unrelated objectives in separate fields, thus emerge from ascent on one differentiable quantity, and the projections of intelligence gain a common quantitative footing.

cs.LG

Snatcher: Apple Find My Network Exposes Your Lost Devices To Strangers

Apple's Find My network connects nearly one billion devices to locate missing property via Bluetooth Low Energy (BLE). This paper reveals that insecure BLE advertisements and design tradeoffs allow unauthorized discovery and physical theft of lost Apple devices. We develop Snatcher, an attack and analysis framework implemented fully on Android smartphones without specialized hardware. Snatcher identifies vulnerabilities in unencrypted BLE advertisements, unauthenticated acoustic triggers, and slow MAC address randomization. Through three levels - sound-based direction finding, RSSI-IMU sensor-fusion navigation, and spatial-temporal clustering - our Android-based platform physically tracks and locates lost Apple accessories and devices in real-world tests. Our results highlight a crucial conflict between privacy protection, anti-stalking design, and physical security, urging Apple to strengthen Find My defenses.

cs.CR

Controllable Lung Nodule Synthesis via Histogram-Regularized Latent Diffusion Models

While automated diagnosis systems have achieved remarkable success in computed tomography (CT)-based lung cancer screening, their development remains limited by the scarcity of diverse, annotated pulmonary nodule datasets. Diffusion-based generative models offer a promising strategy for data synthesis; however, many existing conditional approaches primarily optimize spatial reconstruction losses, which encourage voxel-wise similarity but may inadequately constrain lesion-level intensity distributions. As a result, these methods may produce over-smoothed texture profiles and underrepresent the distinct attenuation characteristics of different nodule subtypes, including solid, part-solid, and ground-glass nodules. To address this challenge, we propose a controllable latent diffusion model that synthesizes pulmonary nodules within full 3D CT volumes while accurately modeling nodule-specific intensity distributions. Specifically, rather than relying solely on spatial losses, we introduce a histogram-based regularization term that constrains voxel intensity distributions during the generative process. The model combines subtype, spatial mask, and Hounsfield unit (HU) histogram conditioning with the differentiable feature-space histogram regularization term to better align lesion-level intensity distributions, improving the visual plausibility and subtype consistency of synthesized nodules. Extensive experiments on lung CT data demonstrate that our framework achieves strong visual realism, validated through both quantitative metrics and a visual Turing test. Furthermore, when used for data augmentation, the generated nodules improve performance in downstream clinical tasks, particularly for underrepresented nodule subtypes, and show a potential benefit for subtype-informed malignancy classification.

cs.CV

Language Game: Talking to Non-Human Systems

Language carries thought and coordination among humans but rarely reaches further along the spectrum of diverse intelligence. Yet non-neural systems -- from gene regulatory networks and microbial consortia to fungi -- are increasingly recognized as substrates of computation, decision-making and memory, making dialogue with non-human intelligence newly conceivable. Today such dialogue is attempted only by proxy: a large language model speaks on the system's behalf, so any intelligence on display originates from the model while the system itself remains silent. Here we ask whether the system can speak in its own voice. Following Wittgenstein, who located meaning in use, we treat communication as a game played with the system. Its internal dynamics are frozen as the nonlinear core of a reinforcement-learning policy, with only linear input and output interfaces trained. Through use and reward, the system's states and responses acquire meaning within the game, so playing becomes speaking. Because different architectures playing the same game optimize the same reward, their behaviors can all be read as pursuit of that reward; the game serves as a lingua franca across otherwise irreconcilable representations. Given a human prompt, a language model routes it to the game whose semantics best match it and designs an environmental state for which the desired action is the rational response, letting the system reply through its own behavior. Applied across diverse gene regulatory networks and reinforcement-learning tasks, the framework yields fluent dialogue without altering any system parameter, shows that well-trained agents of disparate origin converge on similar behavior, and reveals that specific GRN properties make a system easier or harder to talk with -- an inductive bias of the reservoir itself. Our framework opens a new route to conversing with any dynamical system on its own terms.

cs.LG

Ramsey numbers and Gallai--Ramsey numbers of disjoint unions of cherries

For graphs $G_1,\ldots,G_k$, the Ramsey number $R(G_1,\ldots,G_k)$ is the smallest positive integer $N$ such that every $k$-edge-coloring of $K_N$ contains a monochromatic copy of $G_i$ in color $i$ for some $i\in[k]$. The Gallai--Ramsey number $GR(G_1,\ldots,G_k)$ is defined analogously, with the colorings restricted to Gallai colorings (i.e., edge-colorings with no rainbow triangle). A copy of $P_3$ is called a cherry. Let $n_iP_3$ denote the disjoint union of $n_i$ cherries. Wu, Magnant, Nowbandegani, and Xia (Discrete Appl. Math., 2019) proposed two conjectures: \[ R(n_1P_3,\ldots,n_kP_3)=N\ \text{and}\ GR(n_1P_3,\ldots,n_kP_3)=N\,, \] where $N=2\max\{n_1,\ldots,n_k\}+\sum_{i=1}^kn_i-k+1$. We disprove the Ramsey conjecture and provide some sufficient conditions for determining the exact value of $R(n_1P_3,\ldots,n_kP_3)$. In contrast, we confirm the Gallai--Ramsey conjecture.

math.CO

New bounds for Ramsey numbers involving graphs with a center

Let $F_n$, $W_n$, and $\widehat{K}_n$ be the graphs obtained by joining a vertex to $n$ independent edges, a cycle and a path of order $n-1$, respectively. In this paper, we give new bounds for the Ramsey numbers $R(F_n,F_m)$ and $R(W_n,W_n)$, which improve those due to Chen, Yu, and Zhao [EJC, 2021] and Mao, Wang, Magnant, and Schiermeyer [G&C, 2022], respectively, and establish lower and upper bounds for $R(\widehat{K}_n,\widehat{K}_n)$. Moreover, we present a blow-up technique to establish some new lower bounds for the Ramsey numbers of wheels versus cliques.

math.CO

A Little Rank Goes a Long Way: Random Scaffolds with LoRA Adapters Are All You Need

How many of a neural network's parameters actually encode task-specific information? We investigate this question with LottaLoRA, a training paradigm in which every backbone weight is drawn at random and frozen; only low-rank LoRA adapters are trained. Across nine benchmarks spanning diverse architecture families from single-layer classifiers to 900M parameter Transformers low-rank adapters over frozen random backbones recover 96-100% of fully trained performance while training only 0.5-40% of the parameters. The task-specific signal therefore occupies a subspace orders of magnitude smaller than the full parameter count suggests. Three mechanistic findings underpin this result:(1) the frozen backbone is actively exploited when static the learned scaling~$\beta$ remains strictly positive across all architectures but when the scaffold is destabilized, the optimizer silences it and the LoRA factors absorb all task information; (2) the frozen backbone is preferable but interchangeable any random initialization works equally well, provided it remains fixed throughout training; and (3) the minimum LoRA rank at which performance saturates estimates the intrinsic dimensionality of the task, reminiscent of the number of components retained in Principal Component Analysis (PCA). The construction is formally analogous to Reservoir Computing unfolded along the depth axis of a feedforward network. Because the backbone is determined by a random seed alone, models can be distributed as adapters plus seed a footprint that grows with task complexity, not model size, so that storage and memory savings compound as architectures scale.

cs.LG

Scribe Verification in Chinese manuscripts using Siamese, Triplet, and Vision Transformer Neural Networks

The paper examines deep learning models for scribe verification in Chinese manuscripts. That is, to automatically determine whether two manuscript fragments were written by the same scribe using deep metric learning methods. Two datasets were used: the Tsinghua Bamboo Slips Dataset and a selected subset of the Multi-Attribute Chinese Calligraphy Dataset, focusing on the calligraphers with a large number of samples. Siamese and Triplet neural network architectures are implemented, including convolutional and Transformer-based models. The experimental results show that the MobileNetV3+ Custom Siamese model trained with contrastive loss achieves either the best or the second-best overall accuracy and area under the Receiver Operating Characteristic Curve on both datasets.

cs.LG

Specializing Foundation Models via Mixture of Low-Rank Experts for Comprehensive Head CT Analysis

Foundation models pre-trained on large-scale datasets demonstrate strong transfer learning capabilities; however, their adaptation to complex multi-label diagnostic tasks-such as comprehensive head CT finding detection-remains understudied. Standard parameter-efficient fine-tuning methods such as LoRA apply uniform adaptations across pathology types, which may limit performance for diverse medical findings. We propose a Mixture of Low-Rank Experts (MoLRE) framework that extends LoRA with multiple specialized low-rank adapters and unsupervised soft routing. This approach enables conditional feature adaptation with less than 0.5% additional parameters and without explicit pathology supervision. We present a comprehensive benchmark of MoLRE across six state-of-the-art medical imaging foundation models spanning 2D and 3D architectures, general-domain, medical-domain, and head CT-specific pretraining, and model sizes ranging from 7M to 431M parameters. Using over 70,000 non-contrast head CT scans with 75 annotated findings-including hemorrhage, infarction, trauma, mass lesions, structural abnormalities, and chronic changes-our experiments demonstrate consistent performance improvements across all models. Gains vary substantially: general-purpose and medical-domain models show the largest improvements (DINOv3-Base: +4.6%; MedGemma: +4.3%), whereas 3D CT-specialized or very large models show more modest gains (+0.2-1.3%). The combination of MoLRE and MedGemma achieves the highest average detection AUC of 0.917. These findings highlight the importance of systematic benchmarking on target clinical tasks, as pretraining domain, architecture, and model scale interact in non-obvious ways.

cs.CV

Online Ramsey numbers of the claw versus cycles

The online Ramsey number $\tilde r(G,H)$ is defined via a Builder--Painter game on an empty graph with countably many vertices. In each round, Builder reveals an edge, which Painter immediately colors either red or blue. Builder wins once a red copy of $G$ or a blue copy of $H$ appears, and $\tilde r(G,H)$ is the minimum number of edges Builder must reveal to force a win. For a long cycle $C_\ell$, the online Ramsey numbers $\tilde r(G,C_\ell)$ are known only for a few specific choices of $G$. In particular, exact values were determined for $G=P_3$ by Cyman, Dzido, Lapinskas, and Lo (Electron. J. Combin., 2015), while asymptotically tight results were obtained when $G$ is an even cycle by Adamski, Bednarska-Bzd\c{e}ga, and Bla\v{z}ej (SIAM J. Discrete Math., 2024). In this paper, we consider the case where $G$ is the claw $K_{1,3}$ and determine the exact value of $\tilde r(K_{1,3},C_\ell)$. We show that \[ \tilde r(K_{1,3},C_\ell)=\left\lfloor \frac{3(\ell+1)}{2} \right\rfloor \quad \text{for all } \ell \ge 13. \]

math.CO

Revisiting 2D Foundation Models for Scalable 3D Medical Image Classification

3D medical image classification is essential for modern clinical workflows. Medical foundation models (FMs) have emerged as a promising approach for scaling to new tasks, yet current research suffers from three critical pitfalls: data-regime bias, suboptimal adaptation, and insufficient task coverage. In this paper, we address these pitfalls and introduce AnyMC3D, a scalable 3D classifier adapted from 2D FMs. Our method scales efficiently to new tasks by adding only lightweight plugins (about 1M parameters per task) on top of a single frozen backbone. This versatile framework also supports multi-view inputs, auxiliary pixel-level supervision, and interpretable heatmap generation. We establish a comprehensive benchmark of 12 tasks covering diverse pathologies, anatomies, and modalities, and systematically analyze state-of-the-art 3D classification techniques. Our analysis reveals key insights: (1) effective adaptation is essential to unlock FM potential, (2) general-purpose FMs can match medical-specific FMs if properly adapted, and (3) 2D-based methods surpass 3D architectures for 3D classification. For the first time, we demonstrate the feasibility of achieving state-of-the-art performance across diverse applications using a single scalable framework (including 1st place in the VLM3D challenge), eliminating the need for separate task-specific models.

cs.CV

Senti-iFusion: An Integrity-centered Hierarchical Fusion Framework for Multimodal Sentiment Analysis under Uncertain Modality Missingness

Multimodal Sentiment Analysis (MSA) is critical for human-computer interaction but faces challenges when the modalities are incomplete or missing. Existing methods often assume pre-defined missing modalities or fixed missing rates, limiting their real-world applicability. To address this challenge, we propose Senti-iFusion, an integrity-centered hierarchical fusion framework capable of handling both inter- and intra-modality missingness simultaneously. It comprises three hierarchical components: Integrity Estimation, Integrity-weighted Completion, and Integrity-guided Fusion. First, the Integrity Estimation module predicts the completeness of each modality and mitigates the noise caused by incomplete data. Second, the Integrity-weighted Cross-modal Completion module employs a novel weighting mechanism to disentangle consistent semantic structures from modality-specific representations, enabling the precise recovery of sentiment-related features across language, acoustic, and visual modalities. To ensure consistency in reconstruction, a dual-depth validation with semantic- and feature-level losses ensures consistent reconstruction at both fine-grained (low-level) and semantic (high-level) scales. Finally, the Integrity-guided Adaptive Fusion mechanism dynamically selects the dominant modality for attention-based fusion, ensuring that the most reliable modality, based on completeness and quality, contributes more significantly to the final prediction. Senti-iFusion employs a progressive training approach to ensure stable convergence. Experimental results on popular MSA datasets demonstrate that Senti-iFusion outperforms existing methods, particularly in fine-grained sentiment analysis tasks. The code and our proposed Senti-iFusion model will be publicly available.

cs.HC

Three-color online Ramsey numbers $\tilde{r}(P_3,P_3,P_{\ell})$ and $\tilde{r}(P_3, P_3, C_{\ell})$

For given graphs $G_1, \ldots, G_k$, let $\tilde{r}(G_1, \ldots, G_k)$ denote their online Ramsey number. In an influential paper on the online Ramsey numbers for paths and cycles, Cyman, Dzido, Lapinskas, and Lo (Electron. J. Combin., 2015) determined the exact values of $\tilde{r}(P_3, P_{\ell})$ and $\tilde{r}(P_3, C_{\ell})$. They also conjectured the exact value of $\tilde{r}(P_4, P_{\ell})$ and the limit of $\tilde{r}(P_k, P_{\ell})/\ell$ as $\ell \to \infty$ for $k \ge 5$. The former conjecture was independently confirmed by Bednarska-Bzd\c{e}ga (European J. Combin., 2024) and Y.B. Zhang and Y.X. Zhang (arXiv:2302.13640), while the latter was disproved by Mond and Portier (European J. Combin., 2024). In this paper, we extend this line of research to the three-color setting and establish the exact value of $\tilde{r}(P_3,P_3,P_{\ell})$ for $\ell\ge 2$ and $\tilde{r}(P_3, P_3, C_{\ell})$ for $\ell \ge 16$.

math.CO

Equilibrium flow: From Snapshots to Dynamics

Scientific data, from cellular snapshots in biology to celestial distributions in cosmology, often consists of static patterns from underlying dynamical systems. These snapshots, while lacking temporal ordering, implicitly encode the processes that preserve them. This work investigates how strongly such a distribution constrains its underlying dynamics and how to recover them. We introduce the Equilibrium flow method, a framework that learns continuous dynamics that preserve a given pattern distribution. Our method successfully identifies plausible dynamics for 2-D systems and recovers the signature chaotic behavior of the Lorenz attractor. For high-dimensional Turing patterns from the Gray-Scott model, we develop an efficient, training-free variant that achieves high fidelity to the ground truth, validated both quantitatively and qualitatively. Our analysis reveals the solution space is constrained not only by the data but also by the learning model's inductive biases. This capability extends beyond recovering known systems, enabling a new paradigm of inverse design for Artificial Life. By specifying a target pattern distribution, we can discover the local interaction rules that preserve it, leading to the spontaneous emergence of complex behaviors, such as life-like flocking, attraction, and repulsion patterns, from simple, user-defined snapshots.

cs.LG

ZapGPT: Free-form Language Prompting for Simulated Cellular Control

Human language is one of the most expressive tools for conveying intent, yet most artificial or biological systems lack mechanisms to interpret or respond meaningfully to it. Bridging this gap could enable more natural forms of control over complex, decentralized systems. In AI and artificial life, recent work explores how language can specify high-level goals, but most systems still depend on engineered rewards, task-specific supervision, or rigid command sets, limiting generalization to novel instructions. Similar constraints apply in synthetic biology and bioengineering, where the locus of control is often genomic rather than environmental perturbation. A key open question is whether artificial or biological collectives can be guided by free-form natural language alone, without task-specific tuning or carefully designed evaluation metrics. We provide one possible answer here by showing, for the first time, that simple agents' collective behavior can be guided by free-form language prompts: one AI model transforms an imperative prompt into an intervention that is applied to simulated cells; a second AI model scores how well the prompt describes the resulting cellular dynamics; and the former AI model is evolved to improve the scores generated by the latter. Unlike previous work, our method does not require engineered fitness functions or domain-specific prompt design. We show that the evolved system generalizes to unseen prompts without retraining. By treating natural language as a control layer, the system suggests a future in which spoken or written prompts could direct computational, robotic, or biological systems to desired behaviors. This work provides a concrete step toward this vision of AI-biology partnerships, in which language replaces mathematical objective functions, fixed rules, and domain-specific programming.

cs.AI

Ramsey numbers of sparse graphs versus disjoint books

Let $B_k$ denote a book on $k+2$ vertices and $tB_k$ be $t$ vertex-disjoint $B_k$'s. Let $G$ be a connected graph with $n$ vertices and at most $n(1+\epsilon)$ edges, where $\epsilon$ is a constant depending on $k$ and $t$. In this paper, we show that the Ramsey number $$r(G,tB_k)=2n+t-2$$ provided $n\ge 111t^3k^3$. Our result extends the work of Erd\H{o}s, Faudree, Rousseau, and Schelp (1988), who established the corresponding result for $G$ being a tree and $t=1$.

math.CO