SearcharxivSearch

arXiv subjects

Jichang Zhao

Publications and source records attributed to Jichang Zhao.

At least 19 recordsLinked to original sources

Can Machines Really See Objects in Images? A Study Based on Syntactic Distance and Visual Self-Referential Instances

Can a vision model truly see an object, or does it only fit surface-level visual cues? Following Wittgenstein's view that the limits of language are the limits of the world, we view a model's recognition ability as bounded by the descriptive system it has learned. In current vision models, this system is often realized through learned feature representations that exploit local statistical cues. We therefore ask whether a model can still classify correctly when such local cues provide no stable basis for distinction. We formalize this question with syntactic distance, which measures class separability through the symmetry of the operations mapping one class to the other: positive distance exposes exploitable local features, whereas zero distance requires global semantics rather than local rules. We construct a visual self-referential task in maximum-variance binary noise: positive samples contain a closed square, while negative samples contain an otherwise identical square with one flipped boundary pixel. The two classes differ in global semantics but have zero syntactic distance, making local statistical shortcuts unreliable. Experiments on ResNets and Vision Transformers reveal a consistent phase-transition phenomenon, with accuracy collapsing to random guessing once the image scale crosses a critical point and does not recover within the tested range. Larger training sets and models only delay this collapse, while globally attentive ViTs reach it earlier. These results reveal a structural capability boundary of current architectures on global-concept tasks, suggesting that general intelligence may require creating new language, not reusing an existing one.

cs.CV

From News Source Sharers to Post Viewers: How Topic Diversity and Conspiracy Theories Shape Engagement With Misinformation During a Health Crisis

Online engagement with misinformation threatens societal well-being, particularly during health crises when susceptibility to misinformation is heightened in a multi-topic context. Here, we focus on the COVID-19 pandemic and address a critical gap in understanding engagement with multi-topic misinformation on social media at two user levels: news source sharers (who post news items) and post viewers (who engage with news posts). To this end, we analyze 7273 fact-checked source news items and their associated posts on X through the lens of topic diversity and conspiracy theories. We find that false news, especially those containing conspiracy theories, exhibits higher topic diversity than true news. At news source sharer level, false news has a longer lifetime and receives more posts on X than true news, with conspiracy theories further extending its longevity. However, topic diversity does not significantly influence news source sharers' engagement. At post viewer level, contrary to news source sharer level, posts characterized by heightened topic diversity receive more reposts, likes, and replies. Notably, post viewers tend to engage more with misinformation containing conspiracy narratives: false news posts that contain conspiracy theories, on average, receive 40.8% more reposts, 45.2% more likes, and 44.1% more replies compared to those without conspiracy theories. Our findings suggest that news source sharers and post viewers exhibit distinct engagement patterns on X, offering valuable insights into refining misinformation interventions at these two user levels.

cs.SI

Is Fact-Checking Politically Neutral? Asymmetries in How U.S. Fact-Checking Organizations Pick Up False Statements Mentioning Political Elites

Political elites play an important role in the proliferation of online misinformation. However, an understanding of how fact-checking platforms pick up politicized misinformation for fact-checking is still in its infancy. Here, we conduct an empirical analysis of mentions of U.S. political elites within fact-checked statements. For this purpose, we collect a comprehensive dataset consisting of 35,014 true and false statements that have been fact-checked by two major fact-checking organizations (Snopes, PolitiFact) in the U.S. between 2008 and 2023, i.e., within an observation period of 15 years. Subsequently, we perform content analysis and explanatory regression modeling to analyze how veracity is linked to mentions of U.S. political elites in fact-checked statements. Our analysis yields the following main findings: (i) Fact-checked false statements are, on average, 20% more likely to mention political elites than true fact-checked statements. (ii) There is a partisan asymmetry such that fact-checked false statements are 88.1% more likely to mention Democrats, but 26.5% less likely to mention Republicans, compared to fact-checked true statements. (iii) Mentions of political elites in fact-checked false statements reach the highest level during the months preceding elections. (iv) Fact-checked false statements that mention political elites carry stronger other-condemning emotions and are more likely to be pro-Republican, compared to fact-checked true statements. In sum, our study offers new insights into understanding mentions of political elites in false statements on U.S. fact-checking platforms, and bridges important findings at the intersection between misinformation and politicization.

cs.SI

A Transparent and Nonlinear Method for Variable Selection

Variable selection is a procedure to attain the truly important predictors from inputs. Complex nonlinear dependencies and strong coupling pose great challenges for variable selection in high-dimensional data. In addition, real-world applications have increased demands for interpretability of the selection process. A pragmatic approach should not only attain the most predictive covariates, but also provide ample and easy-to-understand grounds for removing certain covariates. In view of these requirements, this paper puts forward an approach for transparent and nonlinear variable selection. In order to transparently decouple information within the input predictors, a three-step heuristic search is designed, via which the input predictors are grouped into four subsets: the relevant to be selected, and the uninformative, redundant, and conditionally independent to be removed. A nonlinear partial correlation coefficient is introduced to better identify the predictors which have nonlinear functional dependence with the response. The proposed method is model-free and the selected subset can be competent input for commonly used predictive models. Experiments demonstrate the superior performance of the proposed method against the state-of-the-art baselines in terms of prediction accuracy and model interpretability.

stat.ME

Space-Invariant Projection in Streaming Network Embedding

Newly arriving nodes in dynamics networks would gradually make the node embedding space drifted and the retraining of node embedding and downstream models indispensable. An exact threshold size of these new nodes, below which the node embedding space will be predicatively maintained, however, is rarely considered in either theory or experiment. From the view of matrix perturbation theory, a threshold of the maximum number of new nodes that keep the node embedding space approximately equivalent is analytically provided and empirically validated. It is therefore theoretically guaranteed that as the size of newly arriving nodes is below this threshold, embeddings of these new nodes can be quickly derived from embeddings of original nodes. A generation framework, Space-Invariant Projection (SIP), is accordingly proposed to enables arbitrary static MF-based embedding schemes to embed new nodes in dynamics networks fast. The time complexity of SIP is linear with the network size. By combining SIP with four state-of-the-art MF-based schemes, we show that SIP exhibits not only wide adaptability but also strong empirical performance in terms of efficiency and efficacy on the node classification task in three real datasets.

cs.SI

PATE: Property, Amenities, Traffic and Emotions Coming Together for Real Estate Price Prediction

Real estate prices have a significant impact on individuals, families, businesses, and governments. The general objective of real estate price prediction is to identify and exploit socioeconomic patterns arising from real estate transactions over multiple aspects, ranging from the property itself to other contributing factors. However, price prediction is a challenging multidimensional problem that involves estimating many characteristics beyond the property itself. In this paper, we use multiple sources of data to evaluate the economic contribution of different socioeconomic characteristics such as surrounding amenities, traffic conditions and social emotions. Our experiments were conducted on 28,550 houses in Beijing, China and we rank each characteristic by its importance. Since the use of multi-source information improves the accuracy of predictions, the aforementioned characteristics can be an invaluable resource to assess the economic and social value of real estate. Code and data are available at: https://github.com/IndigoPurple/PATE

cs.CY

H4M: Heterogeneous, Multi-source, Multi-modal, Multi-view and Multi-distributional Dataset for Socioeconomic Analytics in the Case of Beijing

The study of socioeconomic status has been reformed by the availability of digital records containing data on real estate, points of interest, traffic and social media trends such as micro-blogging. In this paper, we describe a heterogeneous, multi-source, multi-modal, multi-view and multi-distributional dataset named "H4M". The mixed dataset contains data on real estate transactions, points of interest, traffic patterns and micro-blogging trends from Beijing, China. The unique composition of H4M makes it an ideal test bed for methodologies and approaches aimed at studying and solving problems related to real estate, traffic, urban mobility planning, social sentiment analysis etc. The dataset is available at: https://indigopurple.github.io/H4M/index.html

cs.CY

Enhancing the SVD Compression

Orthonormality is the foundation of matrix decomposition. For example, Singular Value Decomposition (SVD) implements the compression by factoring a matrix with orthonormal parts and is pervasively utilized in various fields. Orthonormality, however, inherently includes constraints that would induce redundant degrees of freedom, preventing SVD from deeper compression and even making it frustrated as the data fidelity is strictly required. In this paper, we theoretically prove that these redundancies resulted by orthonormality can be completely eliminated in a lossless manner. An enhanced version of SVD, namely E-SVD, is accordingly established to losslessly and quickly release constraints and recover the orthonormal parts in SVD by avoiding multiple matrix multiplications. According to the theory, advantages of E-SVD over SVD become increasingly evident with the rising requirement of data fidelity. In particular, E-SVD will reduce 25% storage units as SVD reaches its limitation and fails to compress data. Empirical evidences from typical scenarios of remote sensing and Internet of things further justify our theory and consistently demonstrate the superiority of E-SVD in compression. The presented theory sheds insightful lights on the constraint solution in orthonormal matrices and E-SVD, guaranteed by which will profoundly enhance the SVD-based compression in the context of explosive growth in both data acquisition and fidelity levels.

cs.DS

The illiquidity network of stocks in China's market crash

The Chinese stock market experienced an abrupt crash in 2015, and over one-third of its market value evaporated. Given its associations with fear and the fine resolution with respect to frequency, the illiquidity of stocks may offer a promising perspective for understanding and even signaling a market crash. In this study, by connecting stocks with illiquidity comovements, an illiquidity network is established to model the market. Compared to noncrash days, on crash days, the market is more densely connected due to heavier but more homogeneous illiquidity dependencies that facilitate abrupt collapses. Critical stocks in the illiquidity network, particularly those in the finance sector, are targeted for inspection because of their crucial roles in accumulating and passing on illiquidity losses. The cascading failures of stocks in market crashes are profiled as disseminating from small degrees to high degrees that are usually located in the core of the illiquidity network and then back to the periphery. By counting the days with random failures in the previous five days, an early signal is implemented to successfully predict more than half of the crash days, especially consecutive days in the early phase. Additional evidence from both the Granger causality network and the random network further testifies to the robustness of the signal. Our results could help market practitioners such as regulators detect and prevent the risk of crashes in advance.

q-fin.CP

Price graphs: Utilizing the structural information of financial time series for stock prediction

Great research efforts have been devoted to exploiting deep neural networks in stock prediction. While long-range dependencies and chaotic property are still two major issues that lower the performance of state-of-the-art deep learning models in forecasting future price trends. In this study, we propose a novel framework to address both issues. Specifically, in terms of transforming time series into complex networks, we convert market price series into graphs. Then, structural information, referring to associations among temporal points and the node weights, is extracted from the mapped graphs to resolve the problems regarding long-range dependencies and the chaotic property. We take graph embeddings to represent the associations among temporal points as the prediction model inputs. Node weights are used as a priori knowledge to enhance the learning of temporal attention. The effectiveness of our proposed framework is validated using real-world stock data, and our approach obtains the best performance among several state-of-the-art benchmarks. Moreover, in the conducted trading simulations, our framework further obtains the highest cumulative profits. Our results supplement the existing applications of complex network methods in the financial realm and provide insightful implications for investment applications regarding decision support in financial markets.

q-fin.ST

Anger makes fake news viral online

Fake news that manipulates political elections, strikes financial systems, and even incites riots is more viral than real news online, resulting in unstable societies and buffeted democracy. The easier contagion of fake news online can be causally explained by the greater anger it carries. The same results in Twitter and Weibo indicate that this mechanism is independent of the platform. Moreover, mutations in emotions like increasing anger will progressively speed up the information spread. Specifically, increasing the occupation of anger by 0.1 and reducing that of joy by 0.1 will produce nearly 6 more retweets in the Weibo dataset. Offline questionnaires reveal that anger leads to more incentivized audiences in terms of anxiety management and information sharing and accordingly makes fake news more contagious than real news online. Cures such as tagging anger in social media could be implemented to slow or prevent the contagion of fake news at the source.

cs.SI

Positive emotions help rank negative reviews in e-commerce

Negative reviews, the poor ratings in postpurchase evaluation, play an indispensable role in e-commerce, especially in shaping future sales and firm equities. However, extant studies seldom examine their potential value for sellers and producers in enhancing capabilities of providing better services and products. For those who exploited the helpfulness of reviews in the view of e-commerce keepers, the ranking approaches were developed for customers instead. To fill this gap, in terms of combining description texts and emotion polarities, the aim of the ranking method in this study is to provide the most helpful negative reviews under a certain product attribute for online sellers and producers. By applying a more reasonable evaluating procedure, experts with related backgrounds are hired to vote for the ranking approaches. Our ranking method turns out to be more reliable for ranking negative reviews for sellers and producers, demonstrating a better performance than the baselines like BM25 with a result of 8% higher. In this paper, we also enrich the previous understandings of emotions in valuing reviews. Specifically, it is surprisingly found that positive emotions are more helpful rather than negative emotions in ranking negative reviews. The unexpected strengthening from positive emotions in ranking suggests that less polarized reviews on negative experience in fact offer more rational feedbacks and thus more helpfulness to the sellers and producers. The presented ranking method could provide e-commerce practitioners with an efficient and effective way to leverage negative reviews from online consumers.

cs.CL

Weak ties strengthen anger contagion in social media

Increasing evidence suggests that, similar to face-to-face communications, human emotions also spread in online social media. However, the mechanisms underlying this emotion contagion, for example, whether different feelings spread in unlikely ways or how the spread of emotions relates to the social network, is rarely investigated. Indeed, because of high costs and spatio-temporal limitations, explorations of this topic are challenging using conventional questionnaires or controlled experiments. Because they are collection points for natural affective responses of massive individuals, online social media sites offer an ideal proxy for tackling this issue from the perspective of computational social science. In this paper, based on the analysis of millions of tweets in Weibo, surprisingly, we find that anger travels easily along weaker ties than joy, meaning that it can infiltrate different communities and break free of local traps because strangers share such content more often. Through a simple diffusion model, we reveal that weaker ties speed up anger by applying both propagation velocity and coverage metrics. To the best of our knowledge, this is the first time that quantitative long-term evidence has been presented that reveals a difference in the mechanism by which joy and anger are disseminated. With the extensive proliferation of weak ties in booming social media, our results imply that the contagion of anger could be profoundly strengthened to globalize its negative impact.

cs.SI

How do online consumers review negatively?

Negative reviews on e-commerce platforms, mainly in the form of texts, are posted by online consumers to express complaints about unsatisfactory experiences, providing a proxy of big data for sellers to consider improvements. However, the exact knowledge that lies beyond the negative reviewing still remains unknown. Aimed at a systemic understanding of how online consumers post negative reviews, using 1, 450, 000 negative reviews from JD.com, the largest B2C platform in China, the behavioral patterns from temporal, perceptional and emotional perspectives are comprehensively explored in the present study. Massive consumers behind these reviews across four sectors in the most recent 10 years are further split into five levels to reveal group discriminations at a fine resolution. Circadian rhythms of negative reviewing after making purchases were found, and the periodic intervals suggest stable habits in online consumption and that consumers tend to negatively review at the same hour of the purchase. Consumers from lower levels express more intensive negative feelings, especially on product pricing and seller attitudes, while those from upper levels demonstrate a stronger momentum of negative emotion. The value of negative reviews from higher-level consumers is thus unexpectedly highlighted because of less emotionalization and less biased narration, while the longer-lasting characteristic of these consumers' negative responses also stresses the need for more attention from sellers. Our results shed light on implementing distinguished proactive strategies in different buyer groups to help mitigate the negative impact due to negative reviews.

econ.GN

Behavior variations and their implications for popularity promotions: From elites to mass in Weibo

The boom in social media with regard to producing and consuming information simultaneously implies the crucial role of online user influence in determining content popularity. In particular, understanding behavior variations between the influential elites and the mass grassroots is an important issue in communication. However, how their behavior varies across user categories and content domains, and how these differences influence content popularity are rarely addressed. From a novel view of seven content-domains, a detailed picture of behavior variations among five user groups, from both views of elites and mass, is drawn in Weibo, one of the most popular Twitter-like services in China. Interestingly, elites post more diverse contents with video links while the mass possess retweeters of higher loyalty. According to these variations, user-oriented actions of enhancing content popularity are discussed and testified. The most surprising finding is that the diversity of contents do not always bring more retweets, and the mass and elites should promote content popularity by increasing their retweeter counts and loyalty, respectively. Our results for the first time demonstrate the possibility of highly individualized strategies of popularity promotions in social media, instead of a universal principle.

cs.SI

The emergence of critical stocks in market crash

In complex systems like financial market, risk tolerance of individuals is crucial for system resilience.The single-security price limit, designed as risk tolerance to protect investors by avoiding sharp price fluctuation, is blamed for feeding market panic in times of crash.The relationship between the critical market confidence which stabilizes the whole system and the price limit is therefore an important aspect of system resilience. Using a simplified dynamic model on networks of investors and stocks, an unexpected linear association between price limit and critical market confidence is theoretically derived and empirically verified in this paper. Our results highlight the importance of relatively `small' but critical stocks that drive the system to collapse by passing the failure from periphery to core. These small stocks, largely originating from homogeneous investment strategies across the market, has unintentionally suppressed system resilience with the exclusive increment of individual risk tolerance. Imposing random investment requirements to mitigate herding behavior can thus improve the market resilience.

q-fin.GN

Same Influenza, Different Responses: Social Media Can Sense a Regional Spectrum of Symptoms

Influenza is an acute respiratory infection caused by a virus. It is highly contagious and rapidly mutative. However, its epidemiological characteristics are conventionally collected in terms of outpatient records. In fact, the subjective bias of the doctor emphasizes exterior signs, and the necessity of face-to-face inquiry results in an inaccurate and time-consuming manner of data collection and aggregation. Accordingly, the inferred spectrum of syndromes can be incomplete and lagged. With a massive number of users being sensors, online social media can indeed provide an alternative approach. Voluntary reports in Twitter and its variants can deliver not only exterior signs but also interior feelings such as emotions. These sophisticated signals can further be efficiently collected and aggregated in a real-time manner, and a comprehensive spectrum of syndromes could thus be inferred. Taking Weibo as an example, it is confirmed that a regional spectrum of symptoms can be credibly sensed. Aside from the differences in symptoms and treatment incentives between northern and southern China, it is also surprising that patients in the south are more optimistic, while those in the north demonstrate more intense emotions. The differences sensed from Weibo can even help improve the performance of regressions in monitoring influenza. Our results suggest that self-reports from social media can be profound supplements to the existing clinic-based systems for influenza surveillance.

cs.SI

Online reviews can predict long-term returns of individual stocks

Online reviews are feedback voluntarily posted by consumers about their consumption experiences. This feedback indicates customer attitudes such as affection, awareness and faith towards a brand or a firm and demonstrates inherent connections with a company's future sales, cash flow and stock pricing. However, the predicting power of online reviews for long-term returns on stocks, especially at the individual level, has received little research attention, making a comprehensive exploration necessary to resolve existing debates. In this paper, which is based exclusively on online reviews, a methodology framework for predicting long-term returns of individual stocks with competent performance is established. Specifically, 6,246 features of 13 categories inferred from more than 18 million product reviews are selected to build the prediction models. With the best classifier selected from cross-validation tests, a satisfactory increase in accuracy, 13.94%, was achieved compared to the cutting-edge solution with 10 technical indicators being features, representing an 18.28% improvement relative to the random value. The robustness of our model is further evaluated and testified in realistic scenarios. It is thus confirmed for the first time that long-term returns of individual stocks can be predicted by online reviews. This study provides new opportunities for investors with respect to long-term investments in individual stocks.

econ.GN