SearcharxivSearch

arXiv subjects

Yao Tong

Publications and source records attributed to Yao Tong.

At least 19 recordsLinked to original sources

Dimensions of surface repellers and attractors of non-linear planar IFSs

We establish that for a $C^r$-generic surface repeller ($1 < r \leq \infty$), its Hausdorff and box dimensions are exactly the unique zero of the sub-additive topological pressure function. As an application of our framework, we show that the attractor of a uniformly non-conformal and weakly irreducible planar non-linear iterated function system (IFS) satisfying the strong separation condition (SSC) attains its expected Hausdorff and box dimensions. Furthermore, as a direct consequence of our generic repeller theorem, we deduce that this dimension formula also holds for a generic $C^r$ planar IFS satisfying the SSC. Using this approach, we also extend the dichotomy result for graphs of Weierstrass-type functions of Ren and Shen (2021) by weakening their real-analytic requirement to arbitrary $C^r$ regularity for $r > 1$.

math.DS

A prime orbit theorem for smooth surface diffeomorphisms

We establish a sharp prime orbit theorem for every homoclinic class of a $C^\infty$ diffeomorphism on a closed surface with positive topological entropy. Let $\mathcal{H}$ be a homoclinic class with topological entropy $h > 0$. Then there exists a constant $\chi_2 < 0$ such that for any $\chi_1 \in (0, h)$, \[ \lim_{\substack{l(\mathcal{H}) \mid n \\ n\to\infty}} \frac{\sharp P_{\chi_1,\chi_2}(n)}{e^{nh}} = l(\mathcal{H}). \] Here $P_{\chi_1,\chi_2}(n)$ stands for the set of period-$n$ saddle points in $\mathcal{H}$ with Lyapunov exponents lying outside the interval $[\chi_2,\chi_1]$, and $l(\mathcal{H})$ denotes the period associated with the homoclinic class $\mathcal{H}$.

math.DS

SRB Measures for $C^{1+\mathrm{Dini}}$ Diffeomorphisms

For $C^{1+\mathrm{Dini}}$ diffeomorphisms, we prove a Ledrappier--Young-type characterization of SRB measures among invariant measures whose supports admit dominated splittings. When the dominating bundle has only positive Lyapunov exponents, absolute continuity of the conditional measures on the corresponding Pesin unstable manifolds is equivalent to the partial Pesin entropy formula. As an application, we obtain SRB measures for partially hyperbolic mostly expanding attractors in the $C^{1+\mathrm{Dini}}$ category. Counterexamples are provided to show that these conclusions fail in the $C^1$ category, even under uniform hyperbolicity.

math.DS

Ledrappier-Young entropy formula for $C^1$ diffeomorphisms with dominated splitting Part 2: Entropy formulas and measure dimension

In this paper, we partially extend the Ledrappier-Young entropy formula to invariant measures of $C^1$ diffeomorphisms with dominated splittings. For such measures, we show that whenever the $i$-th Lyapunov exponent has multiplicity one, the $i$-th transverse entropy equals the product of the $i$-th Lyapunov exponent and the corresponding transverse measure dimension. Furthermore, if all intermediate non-negative Lyapunov exponents have multiplicity one, then the Ledrappier-Young entropy formula holds. As applications, we derive $C^1$ versions of numerous results in measure dimension theory, including the famous works by Ledrappier-Young, Barreira-Pesin-Schmeling, and Ledrappier-Xie.

math.DS

Continuity properties of partial entropy

We establish a general criterion on the upper semi-continuity of partial entropy in all directions for $C^{1+\alpha}$ diffeomorphisms: it holds when the respective sums of Lyapunov exponents are continuous. This addresses, in arbitrary dimensions, the converse aspect of the entropic continuity of the Lyapunov exponents established by Buzzi, Crovisier, and Sarig. Consequently, the entropy (and all the partial entropies) is always upper semi-continuous at generic ergodic measures of every $C^{1+\alpha}$ diffeomorphism, which extends the $C^{\infty}$ result of Newhouse. Numerous applications and examples are provided, including topics related to measures with dominated splittings, SRB measures, average expanding diffeomorphisms, singular flows, standard maps, and symbolic codings for diffeomorphisms.

math.DS

Generalization in LLM Problem Solving: The Case of the Shortest Path

Whether language models can systematically generalize remains actively debated. Yet empirical performance is jointly shaped by multiple factors such as training data, training paradigms, and inference-time strategies, making failures difficult to interpret. We introduce a controlled synthetic environment based on shortest-path planning, a canonical composable sequential optimization problem. The setup enables clean separation of these factors and supports two orthogonal axes of generalization: spatial transfer to unseen maps and length scaling to longer-horizon problems. We find that models exhibit strong spatial transfer but consistently fail under length scaling due to recursive instability. We further analyze how distinct stages of the learning pipeline influence systematic problem-solving: for example, data coverage sets capability limits; reinforcement learning improves training stability but does not expand those limits; and inference-time scaling enhances performance but cannot rescue length-scaling failures.

cs.AI

Transformers Are Born Biased: Structural Inductive Biases at Random Initialization and Their Practical Consequences

Transformers underpin modern large language models (LLMs) and are commonly assumed to be behaviorally unstructured at random initialization, with all meaningful preferences emerging only through large-scale training. We challenge this assumption by showing that randomly initialized transformers already exhibit strong and systematic structural biases. In particular, untrained models display extreme token preferences: across random input sequences, certain tokens are predicted with probabilities orders of magnitude larger. We provide a mechanistic explanation for this phenomenon by dissecting the transformer architecture at initialization. We show that extreme token preference arises from a contraction of token representations along a random seed-dependent direction. This contraction is driven by two interacting forces: (i) asymmetric nonlinear activations in MLP sublayers induce global (inter-sequence) representation concentration, and (ii) self-attention further amplifies this effect through local (intra-sequence) aggregation. Together, these mechanisms align hidden representations along a direction determined solely by the random initialization, producing highly non-uniform next-token predictions. Beyond mechanistic insight, we demonstrate that these initialization-induced biases persist throughout training, forming a stable and intrinsic model identity. Leveraging this property, we introduce SeedPrint, a fingerprinting method that can reliably distinguish models that differ only in their random initialization, even after extensive training and under substantial distribution shift. Finally, we identify a fundamental positional discrepancy inherent to the attention mechanism's intra-sequence contraction that is causally linked to the attention-sink phenomenon. This discovery provides a principled explanation for the emergence of sinks and offers a pathway for their control.

stat.ML

Mapping the Vanishing and Transformation of Urban Villages in China

Urban villages (UVs), informal settlements embedded within China's urban fabric, have undergone widespread demolition and redevelopment in recent decades. However, there remains a lack of systematic evaluation of whether the demolished land has been effectively reused, raising concerns about the efficacy and sustainability of current redevelopment practices. To address the gap, this study proposes a deep learning-based framework to monitor the spatiotemporal changes of UVs in China. Specifically, semantic segmentation of multi-temporal remote sensing imagery is first used to map evolving UV boundaries, and then post-demolition land use is classified into six categories based on the "remained-demolished-redeveloped" phase: incomplete demolition, vacant land, construction sites, buildings, green spaces, and others. Four representative cities from China's four economic regions were selected as the study areas, i.e., Guangzhou (East), Zhengzhou (Central), Xi'an (West), and Harbin (Northeast). The results indicate: 1) UV redevelopment processes were frequently prolonged; 2) redevelopment transitions primarily occurred in peripheral areas, whereas urban cores remained relatively stable; and 3) three spatiotemporal transformation pathways, i.e., synchronized redevelopment, delayed redevelopment, and gradual optimization, were revealed. This study highlights the fragmented, complex and nonlinear nature of UV redevelopment, underscoring the need for tiered and context-sensitive planning strategies. By linking spatial dynamics with the context of redevelopment policies, the findings offer valuable empirical insights that support more inclusive, efficient, and sustainable urban renewal, while also contributing to a broader global understanding of informal settlement transformations.

cs.CV

SeedPrints: Fingerprints Can Even Tell Which Seed Your Large Language Model Was Trained From

Fingerprinting Large Language Models (LLMs)is essential for provenance verification and model attribution. Existing fingerprinting methods are primarily evaluated after fine-tuning, where models have already acquired stable signatures from training data, optimization dynamics, or hyperparameters. However, most of a model's capacity and knowledge are acquired during pretraining rather than downstream fine-tuning, making large-scale pretraining a more fundamental regime for lineage verification. We show that existing fingerprinting methods become unreliable in this regime, as they rely on post-hoc signatures that only emerge after substantial training. This limitation contradicts the classical Galton notion of a fingerprint as an intrinsic and persistent identity. In contrast, we propose a stronger and more intrinsic notion of LLM fingerprinting: SeedPrints, a method that leverages random initialization biases as persistent, seed-dependent identifiers present even before training begins. We show that untrained models exhibit reproducible prediction biases induced by their initialization seed, and that these weak signals remain statistically detectable throughout training, enabling high-confidence lineage verification. Unlike prior techniques that fail during early pretraining or degrade under distribution shifts, SeedPrints remains effective across all training stages, from initialization to large-scale pretraining and downstream adaptation. Experiments on LLaMA-style and Qwen-style models demonstrate seed-level distinguishability and enable birth-to-lifecycle identity verification. Evaluations on large-scale pretraining trajectories and real-world fingerprinting benchmarks further confirm its robustness under prolonged training, domain shifts, and parameter modifications.

cs.CR

From Harm to Help: Turning Reasoning In-Context Demos into Assets for Reasoning LMs

Recent reasoning LLMs (RLMs), especially those trained with verifier-based reinforcement learning, often perform worse with few-shot CoT than with direct answering. We revisit this paradox using high-quality reasoning traces from DeepSeek-R1 as demonstrations and find that adding more exemplars consistently degrades accuracy, even when demonstrations are optimal. A detailed analysis reveals two mechanisms behind this decline: (i) semantic misguidance, where high textual similarity leads the model to treat the target as the same as the exemplar and to copy intermediate steps verbatim; and (ii) strategy transfer failure, where the model struggles to extract useful reasoning strategies and apply them to target questions. Guided by these, we introduce Insight-to-Solve (I2S), a sequential test-time procedure that turns demonstrations into explicit, reusable insights and derives a target-specific reasoning trace; optionally, the reasoning is self-refined for coherence and correctness (I2S+). Extensive experiments on diverse benchmarks show that I2S and I2S+ consistently outperform both direct answering and test-time scaling baselines across open- and closed-source models. Even for GPT models, our method helps: on AIME'25, GPT-4.1 rises by +14.0%, and o1-mini improves by +2.7% on AIME and +1.7% on GPQA, indicating that in-context demonstrations can be harnessed effectively via insight-refine-solve framework.

cs.CL

Ledrappier-Young entropy formula for $C^1$ diffeomorphisms with dominated splitting Part 1: Unstable entropy formula and invariance principle

We study the unstable entropy of $C^1$ diffeomorphisms with dominated splittings. Our main result shows that when the zero Lyapunov exponent has multiplicity one, the center direction contributes no entropy, and the unstable entropy coincides with the metric entropy. This extends the celebrated work of Ledrappier-Young [18] for $C^2$ diffeomorphisms to the $C^1$ setting under these assumptions. In particular, our results apply to $C^1$ diffeomorphisms away from homoclinic tangencies due to [20]. As consequences, we obtain several applications at $C^1$ regularity. The Avila-Viana invariance principle [7, 33] holds when the center is one-dimensional. Results on measures of maximal entropy due to Hertz-Hertz-Tahzibi-Ures [25], Tahzibi-Yang [33], and Ures-Viana-Yang-Yang [34, 35] also remain valid for $C^1$ diffeomorphisms.

math.DS

Robust trapping of 2D excitons in an engineered 1D potential from proximal ferroelectric domain walls

We investigate the confinement of neutral excitons in a one-dimensional (1D) potential, engineered by proximizing hBN-encapsulated monolayer MoSe$_2$ to ferroelectric domain walls (DW) in periodically poled LiNbO$_3$. Our device exploits the nanometer scale in-plane electric field gradient at the DW to induce the dipolar exciton confinement via the Stark effect. Spatially resolved photoluminescence (PL) spectroscopy reveals the emergence of narrow emission lines redshifted from the MoSe$_2$ neutral exciton by up to $\sim100\,$meV, depending on the sample structure. The spatial distribution, excitation energy response and polarization properties of the emission is consistent with signatures of 1D-confined excitons. The large electric field gradients accessible via proximal ferroelectric systems open up new avenues for the creation of robust quantum-confined excitons in atomically thin materials and their heterostructures.

cond-mat.mes-hall

Cut the Deadwood Out: Backdoor Purification via Guided Module Substitution

Model NLP models are commonly trained (or fine-tuned) on datasets from untrusted platforms like HuggingFace, posing significant risks of data poisoning attacks. A practical yet underexplored challenge arises when such backdoors are discovered after model deployment, making retraining-required defenses less desirable due to computational costs and data constraints. In this work, we propose Guided Module Substitution (GMS), an effective retraining-free method based on guided merging of the victim model with just a single proxy model. Unlike prior ad-hoc merging defenses, GMS uses a guided trade-off signal between utility and backdoor to selectively replaces modules in the victim model. GMS offers four desirable properties: (1) robustness to the choice and trustworthiness of the proxy model, (2) applicability under inaccurate data knowledge, (3) stability across hyperparameters, and (4) transferability across different attacks. Extensive experiments on encoder models and decoder LLMs demonstrate the strong effectiveness of GMS. GMS significantly outperforms even the strongest defense baseline, particularly against challenging attacks like LWS.

cs.CL

The Stronger the Diffusion Model, the Easier the Backdoor: Data Poisoning to Induce Copyright Breaches Without Adjusting Finetuning Pipeline

The commercialization of text-to-image diffusion models (DMs) brings forth potential copyright concerns. Despite numerous attempts to protect DMs from copyright issues, the vulnerabilities of these solutions are underexplored. In this study, we formalized the Copyright Infringement Attack on generative AI models and proposed a backdoor attack method, SilentBadDiffusion, to induce copyright infringement without requiring access to or control over training processes. Our method strategically embeds connections between pieces of copyrighted information and text references in poisoning data while carefully dispersing that information, making the poisoning data inconspicuous when integrated into a clean dataset. Our experiments show the stealth and efficacy of the poisoning data. When given specific text prompts, DMs trained with a poisoning ratio of 0.20% can produce copyrighted images. Additionally, the results reveal that the more sophisticated the DMs are, the easier the success of the attack becomes. These findings underline potential pitfalls in the prevailing copyright protection strategies and underscore the necessity for increased scrutiny to prevent the misuse of DMs.

cs.CR

Towards Regulatable AI Systems: Technical Gaps and Policy Opportunities

There is increasing attention being given to how to regulate AI systems. As governing bodies grapple with what values to encapsulate into regulation, we consider the technical half of the question: To what extent can AI experts vet an AI system for adherence to regulatory requirements? We investigate this question through the lens of two public sector procurement checklists, identifying what we can do now, what should be possible with technical innovation, and what requirements need a more interdisciplinary approach.

cs.AI

Equivariant divergence formula for chaotic flows

We prove the equivariant divergence formula for the axiom A flow attractors, which is a recursive formula for perturbation of transfer operators of physical measures along center-unstable manifolds. Hence the linear response acquires an `ergodic theorem', which means that it can be sampled by recursively computing only $2u$ many vectors on one orbit, where $u$ is the unstable dimension.

math.DS

Recursive divergence formulas for perturbing unstable transfer operators and physical measures

We show that the derivative of the (measure) transfer operator with respect to the parameter of the map is a divergence. Then, for physical measures of discrete-time hyperbolic chaotic systems, we derive an equivariant divergence formula for the unstable perturbation of transfer operators along unstable manifolds. This formula and hence the linear response, the parameter-derivative of physical measures, can be sampled by recursively computing only $2u$ many vectors on one orbit, where $u$ is the unstable dimension. The numerical implementation of this formula in \cite{far} is neither cursed by dimensionality nor the sensitive dependence on initial conditions.

math.NA

Fusion: Efficient and Secure Inference Resilient to Malicious Servers

In secure machine learning inference, most of the schemes assume that the server is semi-honest (honestly following the protocol but attempting to infer additional information). However, the server may be malicious (e.g., using a low-quality model or deviating from the protocol) in the real world. Although a few studies have considered a malicious server that deviates from the protocol, they ignore the verification of model accuracy (where the malicious server uses a low-quality model) meanwhile preserving the privacy of both the server's model and the client's inputs. To address these issues, we propose \textit{Fusion}, where the client mixes the public samples (which have known query results) with their own samples to be queried as the inputs of multi-party computation to jointly perform the secure inference. Since a server that uses a low-quality model or deviates from the protocol can only produce results that can be easily identified by the client, \textit{Fusion} forces the server to behave honestly, thereby addressing all those aforementioned issues without leveraging expensive cryptographic techniques. Our evaluation indicates that \textit{Fusion} is 48.06$\times$ faster and uses 30.90$\times$ less communication than the existing maliciously secure inference protocol (which currently does not support the verification of the model accuracy). In addition, to show the scalability, we conduct ImageNet-scale inference on the practical ResNet50 model and it costs 8.678 minutes and 10.117 GiB of communication in a WAN setting, which is 1.18$\times$ faster and has 2.64$\times$ less communication than those of the semi-honest protocol.

cs.CR