SearcharxivSearch

arXiv subjects

Wentian Li

Publications and source records attributed to Wentian Li.

At least 19 recordsLinked to original sources

Non-Zipfian Distribution of Stopwords or Function Words and Subset Selection Models

Stopwords and function words are relatively less informative for the content of a language and more often play a structural role in a sentence. Stopwords are ubiquitous words and may contain verbs, adjectives and adverbs. On the other hand, function words are strictly prepositions, conjunctions, pronouns, determiners, qualifiers, articles, interrogatives, and a limited number of auxiliary verbs. In contrast to the well known Zipf's law for rank-frequency plot for all words, the rank-frequency plots for stopwords or function words are best fitted by the Beta Rank Function (BRF). On the other hand, the rank-frequency plots of non-stopwords or non-function-words also deviate from the Zipf's law, but are better described by a quadratic function of log-token-count over log-rank than by BRF. Based on the observed rank of stopwords or function words in the full word list, we propose a stopword/function word/subset selection model that the probability for being selected, as a function of the word's rank $r$, is a decreasing Hill's function ($1/(1+(r/r_{mid})^\gamma)$); whereas the probability for not being selected is the standard Hill's function ($1/(1+(r_{mid}/r)^\gamma)$). We validate this selection probability model by a direct estimation from an independent collection of texts. We also show analytically that this model leads to a BRF rank-frequency distribution for stopwords or function words when the original full word list follows the Zipf's law, as well as explaining the quadratic fitting function for the non-stopwords or non-function-words. A corollary of these results is that Zipf's law is not expected to be true for telegraphic speech in early childhood language learners or in agrammatism patients.

cs.CL

Modeling Two-Scale Rank Distributions via Redistribution Dynamics or an Analytic Derivation of the Beta Rank Function

Beta Rank Function (BRF) is a two-sided distribution characterized by a smooth peak and double powerlaw decay, widely used to model empirical data exhibiting deviations from pure power laws. In this paper, we introduce a novel two-step generative process that produces data exactly following the BRF distribution. The first step involves any mechanism generating a power-law distribution, while the second step applies a regressive redistribution process that reallocates resources from poorer to richer entities, thereby amplifying inequality. This approach represents the first analytic derivation of an exact BRF distribution from a generative mechanism. We validate the model through applications to income and urban population distributions. Beyond exact generation, this framework offers new insights into the systemic origins of deviations from power laws frequently observed in complex systems, linking rank distributions to underlying feedback and redistribution dynamics.

physics.soc-ph

Quadratic Term Correction on Heaps' Law

Heaps' or Herdan's law characterizes the word-type vs. word-token relation by a power-law function, which is concave in linear-linear scale but a straight line in log-log scale. However, it has been observed that even in log-log scale, the type-token curve is still slightly concave, invalidating the power-law relation. At the next-order approximation, we have shown, by twenty English novels or writings (some are translated from another language to English), that quadratic functions in log-log scale fit the type-token data perfectly. Regression analyses of log(type)-log(token) data with both a linear and quadratic term consistently lead to a linear coefficient of slightly larger than 1, and a quadratic coefficient around -0.02. Using the ``random drawing colored ball from the bag with replacement" model, we have shown that the curvature of the log-log scale is identical to a ``pseudo-variance" which is negative. Although a pseudo-variance calculation may encounter numeric instability when the number of tokens is large, due to the large values of pseudo-weights, this formalism provides a rough estimation of the curvature when the number of tokens is small.

cs.CL

AesTest: Measuring Aesthetic Intelligence from Perception to Production

Perceiving and producing aesthetic judgments is a fundamental yet underexplored capability for multimodal large language models (MLLMs). However, existing benchmarks for image aesthetic assessment (IAA) are narrow in perception scope or lack the diversity needed to evaluate systematic aesthetic production. To address this gap, we introduce AesTest, a comprehensive benchmark for multimodal aesthetic perception and production, distinguished by the following features: 1) It consists of curated multiple-choice questions spanning ten tasks, covering perception, appreciation, creation, and photography. These tasks are grounded in psychological theories of generative learning. 2) It integrates data from diverse sources, including professional editing workflows, photographic composition tutorials, and crowdsourced preferences. It ensures coverage of both expert-level principles and real-world variation. 3) It supports various aesthetic query types, such as attribute-based analysis, emotional resonance, compositional choice, and stylistic reasoning. We evaluate both instruction-tuned IAA MLLMs and general MLLMs on AesTest, revealing significant challenges in building aesthetic intelligence. We will publicly release AesTest to support future research in this area.

cs.CV

A Multidisciplinary Design and Optimization (MDO) Agent Driven by Large Language Models

To accelerate mechanical design and enhance design quality and innovation, we present a Multidisciplinary Design and Optimization (MDO) Agent driven by Large Language Models (LLMs). The agent semi-automates the end-to-end workflow by orchestrating three core capabilities: (i) natural-language-driven parametric modeling, (ii) retrieval-augmented generation (RAG) for knowledge-grounded conceptualization, and (iii) intelligent orchestration of engineering software for performance verification and optimization. Working in tandem, these capabilities interpret high-level, unstructured intent, translate it into structured design representations, automatically construct parametric 3D CAD models, generate reliable concept variants using external knowledge bases, and conduct evaluation with iterative optimization via tool calls such as finite-element analysis (FEA). Validation on three representative cases - a gas-turbine blade, a machine-tool column, and a fractal heat sink - shows that the agent completes the pipeline from natural-language intent to verified and optimized designs with reduced manual scripting and setup effort, while promoting innovative design exploration. This work points to a practical path toward human-AI collaborative mechanical engineering and lays a foundation for more dependable, vertically customized MDO systems.

cs.HC

Multistable Synaptic Plasticity induces Memory Effects and Cohabitation of Chimera and Bump States in Leaky Integrate-and-Fire Networks

Chimera states and bump states are collective synchronization phenomena observed independently (at different parameter regions) in networks of coupled nonlinear oscillators. And while chimera states are characterized by coexistence of coherent and incoherent domains, bump states consist of active domains operating on a silent background. Multistable plasticity in the network connections originates from brain dynamics and is based on the idea that neural cells may transmit inhibitory or excitatory signals depending on various factors, such as local connectivity, influence of neighboring cells etc. During the system/network integration, the link weights adapt and, in the case of multistability, they may organize in coexisting excitatory and/or inhibitory domains. Here, we explore the influence of bistable plasticity on collective synchronization states and we numerically demonstrate that the dynamics of the linking may give rise to co-existence of bump-like and chimera-like states simultaneously in the network. In the case of bump and chimera co-existence, confinement effects are developed: the different domains stay localized and do not travel around the network. Memory effects are also reported in the sense that the final spatial arrangement of the coupling strengths reflects some of the local properties of the initial link distribution. For the quantification of the system's spatial and temporal features, the global and local entropy functions are employed as measures of the network organization, while the average firing rates account for the network evolution and dynamics.

nlin.CD

Extending 1089 attractor to any number of digits and any number of steps

The well-known 1089 trick reflects an amazing trait of digital reversal process and reminisces of a limiting attractor in dynamical systems even though it takes only two steps. It is natural to consider the situations when the number of digits is beyond three as in the original 1089 trick, as well as situations when the number of steps is beyond two. The first part has been mostly done by Webster which we will reproduce. After two steps, the resulting integers are called Papadakis-Webster integers (PWI), which is always divisible by 99, and the resulting quotients consist of only 0's and 1's, which we name Papadakis-Webster binary strings (PWBS). Not all binary strings could be PWBS, and we define the hairpin pairing rule to determine if a binary string is a PWBS. For the second part, we propose a two-option iteration system named iterative digital reversal (IDR) suitably interweaving additions and subtractions. The simplest limiting behavior of IDR is 2-cycles. The elements in an IDR 2-cycle are all composed of repetitions of the 10(9)$_L$89 (L>=0) motif, and are all PWIs. The lower 2-cycle elements after division of 99 belong to the subset of PWBS that are palindromic and consist of 0- and 1-blocks with a minimal length of two. IDR also has higher p-cycles (p=10,12,71) whose elements seem to contain at least one PWI. Another interesting finding about IDR is that it contains non-periodic and diverging trajectories, as the integer values grow to infinity. In these diverging trajectories, while the number of flanking digits around the middle point increases by the iteration, the middle part has an 8-cycle rhythm or signature which has been found in all diverging trajectories. Overall, the generalization of the original 1089 trick in both space and time leads to new patterns in integers and new phenomenology in dynamics.

nlin.CD

Rich dynamical behaviors from a digital reversal operation

Repeatedly adding or subtracting the digital reversal to or from an integer, depending on which one is larger, can be treated as a dynamical system. On one hand, a three-digit version of this map running only two steps is the 1089 mathematical trick problem; on the other hand, this mapping can be compared to John Conway's reverse-add-then-sort (RATS) iteration, as well as the 3x+1 problem, also known as Collatz's map. We numerically run this map and find interesting dynamics, including limiting cycles with unusual periodicity and length-8 diverging trajectories.

nlin.CD

An Extensive Study of Two-Node McCulloch-Pitts Networks

Networks with two nodes are previously grouped into either two classes (mutually interactive, master-slave) or five classes (mutualism, competition, predator-prey, commensalism, amensalism). By allowing self-loops, the number of signed regulatory graphs increases to 39. We provide a complete summary of dynamical behaviors of the 39 two-node McCulloch-Pitts models when the link weights are constrained to three values [$-1$,0,$+1$] and Boolean node variables. Depending on whether the Boolean values are [$-1,1$] (bipolar) or [0,1] (binary), we show that the dynamics could also be different with the same signed regulatory graphs. We demonstrate that slight variations in the McCulloch-Pitts model (called variants) may lead to fundamentally different dynamics. We study the full model space and three kinds of robustness or stability: a) of a rule against parameter change on its overall dynamics, b) for a given state against parameter change on its final state, and c) against an initial state change on its final state. All these stability properties are loosely related to a model's limiting dynamics, with the fixed-point rules to be more stable in the first two types of robustness, but less stable in the third robustness type. These analyses pave the way towards a better understanding of a minimum complex system.

cs.SI

Range-Limited Heaps' Law for Functional DNA Words in the Human Genome

Heaps' or Herdan's law is a linguistic law describing the relationship between the vocabulary/dictionary size (type) and word counts (token) to be a power-law function. Its existence in genomes with certain definition of DNA words is unclear partly because the dictionary size in genome could be much smaller than that in a human language. We define a DNA word as a coding region in a genome that codes for a protein domain. Using human chromosomes and chromosome arms as individual samples, we establish the existence of Heaps' law in the human genome within limited range. Our definition of words in a genomic or proteomic context is different from other definitions such as over-represented k-mers which are much shorter in length. Although an approximate power-law distribution of protein domain sizes due to gene duplication and the related Zipf's law is well known, their translation to the Heaps' law in DNA words is not automatic. Several other animal genomes are shown herein also to exhibit range-limited Heaps' law with our definition of DNA words, though with various exponents. When tokens were randomly sampled and sample sizes reach to the maximum level, a deviation from the Heaps' law was observed, but a quadratic regression in log-log type-token plot fits the data perfectly. Investigation of type-token plot and its regression coefficients could provide an alternative narrative of reusage and redundancy of protein domains as well as creation of new protein domains from a linguistic perspective.

q-bio.GN

Human mobility patterns in Mexico City and their links with socioeconomic variables during the COVID-19 pandemic

The availability of cellphone geolocation data provides a remarkable opportunity to study human mobility patterns and how these patterns are affected by the recent pandemic. Two simple centrality metrics allow us to measure two different aspects of mobility in origin-destination networks constructed with this type of data: variety of places connected to a certain node (degree) and number of people that travel to or from a given node (strength). In this contribution, we present an analysis of node degree and strength in daily origin-destination networks for Greater Mexico City during 2020. Unlike what is observed in many complex networks, these origin-destination networks are not scale free. Instead, there is a characteristic scale defined by the distribution peak; centrality distributions exhibit a skewed two-tail distribution with power law decay on each side of the peak. We found that high mobility areas tend to be closer to the city center, have higher population and better socioeconomic conditions. Areas with anomalous behavior are almost always on the periphery of the city, where we can also observe qualitative difference in mobility patterns between east and west. Finally, we study the effect of mobility restrictions due to the outbreak of the COVID-19 pandemics on these mobility patterns.

cs.SI

Revisiting the Neutral Dynamics Derived Limiting Guanine-Cytosine Content Using the Human De Novo Point Mutation Data

We revisit the topic of human genome guanine-cytosine content under neutral evolution. For this study, the de novo mutation data within human is used to estimate mutational rate instead of using base substitution data between related species. We then define a new measure of mutation bias which separate the de novo mutation counts from the background guanine-cytosine content itself, making comparison between different datasets easier. We derive a new formula for calculating limiting guanine-cytosine content by separating CpG-involved mutational events as an independent variable. Using the formula when CpG-involved mutations are considered, the guanine-cytosine content drops less severely in the limit of neutral dynamics. We provide evidence, under certain assumptions, that an isochore-like structure might remain as a limiting configuration of the neutral mutational dynamics.

q-bio.GN

Beta Rank Function: A Smooth Double-Pareto-Like Distribution

The Beta Rank Function (BRF) $x(u) =A(1-u)^b/u^a$, where $u$ is the normalized and continuous rank of an observation $x$, has wide applications in fitting real-world data from social science to biological phenomena. The underlying probability density function (pdf) $f_X(x)$ does not usually have a closed expression except for specific parameter values. We show however that it is approximately a unimodal skewed and asymmetric two-sided power law/double Pareto/log-Laplacian distribution. The BRF pdf has simple properties when the independent variable is log-transformed: $f_{Z=\log(X)}(z)$ . At the peak it makes a smooth turn and it does not diverge, lacking the sharp angle observed in the double Pareto or Laplace distribution. The peak position of $f_Z(z)$ is $z_0=\log A+(a-b)\log(\sqrt{a}+\sqrt{b})-(a\log(a)-b\log(b))/2 $; the probability is partitioned by the peak to the proportion of $\sqrt{b}/(\sqrt{a}+\sqrt{b})$ (left) and $\sqrt{a}/(\sqrt{a}+\sqrt{b})$ (right); the functional form near the peak is controlled by the cubic term in the Taylor expansion when $a\ne b$; the mean of $Z$ is $E[Z]=\log A+a-b$; the decay on left and right sides of the peak is approximately exponential with forms $e^{\frac{z-\log A}{b} }/b$ and $e^{ -\frac{z-\log A}{a}}/a$. These results are confirmed by numerical simulations. Properties of $f_X(x)$ without log-transforming the variable are much more complex, though the approximate double Pareto behavior, $(x/A)^{1/b}/(bx)$ (for $x A$) is simple. Our results elucidate the relationship between BRF and log-normal distributions when $a=b$ and explain why the BRF is ubiquitous and versatile. Based on the pdf, we suggest a quick way to elucidate if a real data set follows a one-sided power-law, a log-normal, a two-sided power-law or a BRF. We illustrate our results with two examples: urban populations and financial returns.

stat.ME

Quantifying Local Randomness in Human DNA and RNA Sequences Using Erdos Motifs

In 1932, Paul Erdos asked whether a random walk constructed from a binary sequence can achieve the lowest possible deviation (lowest discrepancy), for the sequence itself and for all its subsequences formed by homogeneous arithmetic progressions. Although avoiding low discrepancy is impossible for infinite sequences, as recently proven by Terence Tao, attempts were made to construct such sequences with finite lengths. We recognize that such constructed sequences (we call these "Erdos sequences") exhibit certain hallmarks of randomness at the local level: they show roughly equal frequencies of subsequences, and at the same time exclude the trivial periodic patterns. For the human DNA we examine the frequency of a set of Erdos motifs of length-10 using three nucleotides-to-binary mappings. The particular length-10 Erdos sequence is derived by the length-11 Mathias sequence and is identical with the first 10 digits of the Thue-Morse sequence, underscoring the fact that both are deficient in periodicities. Our calculations indicate that: (1) the purine (A and G)/pyridimine (C and T) based Erdos motifs are greatly underrepresented in the human genome, (2) the strong(G and C)/weak(A and T) based Erdos motifs are slightly overrepresented, (3) the densities of the two are negatively correlated, (4) the Erdos motifs based on all three mappings being combined are slightly underrepresented, and (5) the strong/weak based Erdos motifs are greatly overrepresented in the human messenger RNA sequences.

q-bio.GN

Population patterns in World's administrative units

While there has been an extended discussion concerning city population distribution, little has been said about administrative units. Even though there might be a correspondence between cities and administrative divisions, they are conceptually different entities and the correspondence breaks as artificial divisions form and evolve. In this work we investigate the population distribution of second level administrative units for 150 countries and propose the Discrete Generalized Beta Distribution (DGBD) rank-size function to describe the data. After testing the goodness of fit of this two parameter function against power law, which is the most common model for city population, DGBD is a good statistical model for 73% of our data sets and better than power law in almost every case. Particularly, DGBD is better than power law for fitting country population data. The fitted parameters of this function allow us to construct a phenomenological characterization of countries according to the way in which people are distributed inside them. We present a computational model to simulate the formation of administrative divisions and give numerical evidence that DGBD arises from it. This model along with the DGBD function prove adequate to reproduce and describe local unit evolution and its effect on population distribution.

stat.AP

Beyond Zipf's Law: The Lavalette Rank Function and its Properties

Although Zipf's law is widespread in natural and social data, one often encounters situations where one or both ends of the ranked data deviate from the power-law function. Previously we proposed the Beta rank function to improve the fitting of data which does not follow a perfect Zipf's law. Here we show that when the two parameters in the Beta rank function have the same value, the Lavalette rank function, the probability density function can be derived analytically. We also show both computationally and analytically that Lavalette distribution is approximately equal, though not identical, to the lognormal distribution. We illustrate the utility of Lavalette rank function in several datasets. We also address three analysis issues on the statistical testing of Lavalette fitting function, comparison between Zipf's law and lognormal distribution through Lavalette function, and comparison between lognormal distribution and Lavalette distribution.

physics.data-an

Using Volcano Plots and Regularized-Chi Statistics in Genetic Association Studies

Labor intensive experiments are typically required to identify the causal disease variants from a list of disease associated variants in the genome. For designing such experiments, candidate variants are ranked by their strength of genetic association with the disease. However, the two commonly used measures of genetic association, the odds-ratio (OR) and p-value, may rank variants in different order. To integrate these two measures into a single analysis, here we transfer the volcano plot methodology from gene expression analysis to genetic association studies. In its original setting, volcano plots are scatter plots of fold-change and t-test statistic (or -log of the p-value), with the latter being more sensitive to sample size. In genetic association studies, the OR and Pearson's chi-square statistic (or equivalently its square root, chi; or the standardized log(OR)) can be analogously used in a volcano plot, allowing for their visual inspection. Moreover, the geometric interpretation of these plots leads to an intuitive method for filtering results by a combination of both OR and chi-square statistic, which we term "regularized-chi". This method selects associated markers by a smooth curve in the volcano plot instead of the right-angled lines which corresponds to independent cutoffs for OR and chi-square statistic. The regularized-chi incorporates relatively more signals from variants with lower minor-allele-frequencies than chi-square test statistic. As rare variants tend to have stronger functional effects, regularized-chi is better suited to the task of prioritization of candidate genes.

q-bio.QM

Application of Volcano Plots in Analyses of mRNA Differential Expressions with Microarrays

Volcano plot displays unstandardized signal (e.g. log-fold-change) against noise-adjusted/standardized signal (e.g. t-statistic or -log10(p-value) from the t test). We review the basic and an interactive use of the volcano plot, and its crucial role in understanding the regularized t-statistic. The joint filtering gene selection criterion based on regularized statistics has a curved discriminant line in the volcano plot, as compared to the two perpendicular lines for the "double filtering" criterion. This review attempts to provide an unifying framework for discussions on alternative measures of differential expression, improved methods for estimating variance, and visual display of a microarray analysis result. We also discuss the possibility to apply volcano plots to other fields beyond microarray.

q-bio.QM