SearcharxivSearch

arXiv subjects

Yizhao Yang

Publications and source records attributed to Yizhao Yang.

7 recordsLinked to original sources

Benchmarking Multi-Modal Graph-based Social Media Popularity Prediction

Social media popularity prediction aims to forecast the future reach or influence of online content from early-stage observations. Accurate prediction enables key downstream applications, such as advertising optimization and strategic content planning by users, creators, and platforms. Despite substantial progress, existing popularity prediction works often fail to jointly consider multimodal content and temporal social interaction signals. Moreover, the literature remains highly fragmented across datasets, modalities, observation windows, prediction targets, and evaluation protocols. This fragmentation prevents fair comparison and obscures a systematic understanding of how textual, visual, temporal, and interaction-based signals jointly shape popularity dynamics. To address these challenges, we introduce MMG-Pop, a Multi-modal Graph-based Popularity Prediction benchmark, which unifies datasets, modalities, temporal interaction signals, and representative baselines under a standardized evaluation protocol. Furthermore, we propose MMG-PopNet, a unified multi-modal graph-based network that jointly models the aforementioned multi-modal signals and graph-structured social interactions. Extensive experiments on MMG-Pop, comprising four datasets across Bluesky and Reddit platforms, demonstrate the superior performance of MMG-PopNet and yield new insights into cross-platform training generalization, multi-task prediction benefits, multi-modality contributions, and LLM prediction limitation. These findings establish a unified foundation for future research on social dynamics modeling and intervention under heterogeneous modalities and socially-aware agentic ecosystem paradigms.

cs.SI

Escaping the Context Bottleneck: Active Context Curation for LLM Agents via Reinforcement Learning

Large Language Models (LLMs) struggle with long-horizon tasks due to the "context bottleneck" and the "lost-in-the-middle" phenomenon, where accumulated noise from verbose environments degrades reasoning over multi-turn interactions. To address this issue, we introduce a symbiotic framework that decouples context management from task execution. Our architecture pairs a lightweight, specialized policy model, ContextCurator, with a powerful frozen foundation model, TaskExecutor. Trained via reinforcement learning, ContextCurator actively reduces information entropy in the working memory. It aggressively prunes environmental noise while preserving reasoning anchors, that is, sparse data points that are critical for future deductions. On WebArena, our framework improves the success rate of Gemini-3.0-flash from 36.4% to 41.2% while reducing token consumption by 8.8% (from 47.4K to 43.3K). On DeepSearch, it achieves a 57.1% success rate, compared with 53.9%, while reducing token consumption by a factor of 8. Remarkably, a 7B ContextCurator matches the context management performance of GPT-4o, providing a scalable and computationally efficient paradigm for autonomous long-horizon agents.

cs.AI

Using Global Gravitational Potential Weighted Correlation Function to Constrain Modified Gravity Models

We propose a new marked two-point correlation function weighted by the global gravitational potential as a probe for testing gravity models. Using the LCDM model based on general relativity (GR) as a reference, we investigate two representative modified gravity (MG) scenarios: f(R) gravity and nDGP. The mark used in this work, the global gravitational potential that is reconstructed from the galaxy distribution via the Poisson equation, is in contrast to the local property based mark (e.g., local galaxy number density or gravitational potential of host halo) used in previous studies. By applying two weighting schemes to quantify environment-dependent clustering, we find that this statistic is able to distinguish MG models from GR, with the signal being enhanced in regions corresponding to particular ranges of gravitational potential. These results indicate that the proposed statistic can serve as a useful complement to conventional clustering probes in future surveys, once observational effects and modeling uncertainties are properly taken into account.

astro-ph.CO

COINBench: Moving Beyond Individual Perspectives to Collective Intent Understanding

Understanding human intent is a high-level cognitive challenge for Large Language Models (LLMs), requiring sophisticated reasoning over noisy, conflicting, and non-linear discourse. While LLMs excel at following individual instructions, their ability to distill Collective Intent - the process of extracting consensus, resolving contradictions, and inferring latent trends from multi-source public discussions - remains largely unexplored. To bridge this gap, we introduce COIN-BENCH, a dynamic, real-world, live-updating benchmark specifically designed to evaluate LLMs on collective intent understanding within the consumer domain. Unlike traditional benchmarks that focus on transactional outcomes, COIN-BENCH operationalizes intent as a hierarchical cognitive structure, ranging from explicit scenarios to deep causal reasoning. We implement a robust evaluation pipeline that combines a rule-based method with an LLM-as-the-Judge approach. This framework incorporates COIN-TREE for hierarchical cognitive structuring and retrieval-augmented verification (COIN-RAG) to ensure expert-level precision in analyzing raw, collective human discussions. An extensive evaluation of 20 state-of-the-art LLMs across four dimensions - depth, breadth, informativeness, and correctness - reveals that while current models can handle surface-level aggregation, they still struggle with the analytical depth required for complex intent synthesis. COIN-BENCH establishes a new standard for advancing LLMs from passive instruction followers to expert-level analytical agents capable of deciphering the collective voice of the real world. See our project page on COIN-BENCH.

cs.IR

ConsintBench: Evaluating Language Models on Real-World Consumer Intent Understanding

Understanding human intent is a complex, high-level task for large language models (LLMs), requiring analytical reasoning, contextual interpretation, dynamic information aggregation, and decision-making under uncertainty. Real-world public discussions, such as consumer product discussions, are rarely linear or involve a single user. Instead, they are characterized by interwoven and often conflicting perspectives, divergent concerns, goals, emotional tendencies, as well as implicit assumptions and background knowledge about usage scenarios. To accurately understand such explicit public intent, an LLM must go beyond parsing individual sentences; it must integrate multi-source signals, reason over inconsistencies, and adapt to evolving discourse, similar to how experts in fields like politics, economics, or finance approach complex, uncertain environments. Despite the importance of this capability, no large-scale benchmark currently exists for evaluating LLMs on real-world human intent understanding, primarily due to the challenges of collecting real-world public discussion data and constructing a robust evaluation pipeline. To bridge this gap, we introduce \bench, the first dynamic, live evaluation benchmark specifically designed for intent understanding, particularly in the consumer domain. \bench is the largest and most diverse benchmark of its kind, supporting real-time updates while preventing data contamination through an automated curation pipeline.

cs.CL

Cosmological constraints from the density gradient weighted correlation function

The mark weighted correlation function (MCF) $W(s,μ)$ is a computationally efficient statistical measure which can probe clustering information beyond that of the conventional 2-point statistics. In this work, we extend the traditional mark weighted statistics by using powers of the density field gradient $|\nabla ρ/ρ|^α$ as the weight, and use the angular dependence of the scale-averaged MCFs to constrain cosmological parameters. The analysis shows that the gradient based weighting scheme is statistically more powerful than the density based weighting scheme, while combining the two schemes together is more powerful than separately using either of them. Utilising the density weighted or the gradient weighted MCFs with $α=0.5,\ 1$, we can strengthen the constraint on $Ω_m$ by factors of 2 or 4, respectively, compared with the standard 2-point correlation function, while simultaneously using the MCFs of the two weighting schemes together can be $1.25$ times more statistically powerful than using the gradient weighting scheme alone. The mark weighted statistics may play an important role in cosmological analysis of future large-scale surveys. Many issues, including the possibility of using other types of weights, the influence of the bias on this statistics, as well as the usage of MCFs in the tomographic Alcock-Paczynski method, are worth further investigations.

astro-ph.CO

Using the Mark Weighted Correlation Functions to Improve the Constraints on Cosmological Parameters

We used the mark weighted correlation functions (MCFs), $W(s)$, to study the large scale structure of the Universe. We studied five types of MCFs with the weighting scheme $ρ^α$, where $ρ$ is the local density, and $α$ is taken as $-1,\ -0.5,\ 0,\ 0.5$, and 1. We found that different MCFs have very different amplitudes and scale-dependence. Some of the MCFs exhibit distinctive peaks and valleys that do not exist in the standard correlation functions. Their locations are robust against the redshifts and the background geometry, however it is unlikely that they can be used as ``standard rulers'' to probe the cosmic expansion history. Nonetheless we find that these features may be used to probe parameters related with the structure formation history, such as the values of $σ_8$ and the galaxy bias. Finally, after conducting a comprehensive analysis using the full shapes of the $W(s)$s and $W_{Δs}(μ)$s, we found that, combining different types of MCFs can significantly improve the cosmological parameter constraints. Compared with using only the standard correlation function, the combinations of MCFs with $α=0,\ 0.5,\ 1$ and $α=0,\ -1,\ -0.5,\ 0.5,\ 1$ can improve the constraints on $Ω_m$ and $w$ by $\approx30\%$ and $50\%$, respectively. We find highly significant evidence that MCFs can improve cosmological parameter constraints.

astro-ph.CO