SearcharxivSearch

arXiv subjects

Jiang Zhang

Publications and source records attributed to Jiang Zhang.

At least 19 recordsLinked to original sources

Save 2050: A Planetary-Scale Collective Prediction System for the Singularity Crisis

The Anthropocene mode of development is driving civilization toward a "singularity crisis": artificial intelligence (AI) is rapidly approaching general intelligence (AGI) with escalating risks of losing control, while unrestrained economic growth generates super-exponential growth of entropy production that pushes the Earth system toward its tipping points. This paper proposes the "Save 2050" initiative: a distributed planetary-scale collective prediction system that aggregates judgments about the future from humans and AI through open registration, crowdsourcing, and incentive mechanisms, and integrates them via large-scale simulation into inspectable "predicted worlds," enabling humanity to systematically \emph{see} the future for the first time. We argue for the initiative's feasibility along four dimensions---the maturation of AI forecasting, human collective intelligence and the institutional environment, supporting progress in related fields, and societal demand. We then identify three key enabling technologies: long-horizon automated resolution, simulation-based prediction aggregation, and reflexivity governance. We analyze potential risks---including reflexivity, cognitive monoculture, narrative capture, and regulatory and ethical concerns---together with mitigation strategies, and we outline a phased roadmap with open problems. The initiative's primary goal is not to intervene in the future, but to make the future visible, discussable, and co-writable.

physics.soc-ph

SQL-RewriteBench: A Correctness-Gated, Full-Denominator Benchmark for Statement-Level SQL Rewriting [Experiment,Analysis & Benchmark]

Statement-level SQL rewriting can improve query performance and maintainability without changing the DBMS kernel, but existing benchmarks do not evaluate rewrite methods as deployable systems. They typically focus on DBMS performance, rule regression, query equivalence, or dialect translation, while missing the full path from accepting an input query to producing an executable, result-consistent, and operationally useful rewrite. We present SQL-RewriteBench, a benchmark for statement-level SQL rewriting that applies correctness gating and full-denominator accounting. Its metric suite explicitly separates Source Acceptance, Generation Rate, Execution Coverage, Result Consistency, UnsafeRewrite Rate, and speedup distribution. It also defines SCS, a deterministic index of static SQL structure, and CGOQ, a correctness-gated optimization-quality score that gives optimization credit only after the case-specific Checker Contract is satisfied. CGOQ combines runtime improvement with structural simplification through a continuous scoring function, making it suitable for deployment-oriented rewrite assessment. As an artifact, SQL-RewriteBench provides 180 executable Benchmark Instances organized into EQUIV, PERF, ROBUST, and DIALECT pools, each packaged with SQL, schema metadata, provenance, evidence, and rewrite-opportunity documentation. Across seven representative academic and LLM-based methods, every full-benchmark CGOQ is negative. Existing methods often fail before rewriting, fail result checks, or return correct rewrites that are slower or no better than the input. These results show that deployable SQL rewrite requires broader input handling, result validation, and benefit-aware rewrite decisions.

cs.DB

KAN-LSTM-Transformer Neural Networks, MFV and Cosmological Parameters

Reconstructing the cosmic distance ladder directly from observations is a crucial issue in cosmology. In this paper, we present a novel method for modeling the cosmic distance ladder and estimating cosmological parameters through the use of Kolmogorov-Arnold networks (KAN), Long Short-Term Memory (LSTM), and Transformer networks (collectively referred to as KLT-Net), based on the apparent magnitude data from the Pantheon SN Ia compilation. As a data-driven, non-parametric method for reconstructing the distance modulus $\mu(z)$, KLT-Net is shown to be highly effective in capturing the intricate, nonlinear measurement distributions. After validating against various statistical and machine learning models, we have identified it as the most effective choice among the considered alternatives and ablation experiments. Subsequently, the statistical inference of $H_0$ and $\Omega_{\rm m}$ adopts the flat $\Lambda$CDM framework. Moreover, we introduce the Most Frequent Value (MFV) approach to evaluate the absolute magnitude, $M_B$, from existing literature data. In addition, we employ the Hessian matrix to validate the Bayesian method, demonstrating that the Hubble constant can be precisely constrained from the KLT-Net predictions within the flat $\Lambda$CDM framework. The integration of KLT-Net, the MFV approach, and Bayesian statistics establishes a robust framework for inferring cosmological parameters. This methodology facilitates future cosmological research, particularly in the analysis of complex datasets and the exploration of high-dimensional parameter spaces.

astro-ph.CO

Self-Reference in Large Language Models: The Introspection Threshold for Recursive Self-Improvement

The pursuit of self-evolving AI raises a critical question: when is autonomous self-improvement sustainable rather than degenerative? Drawing an analogy to von Neumann's complexity threshold for self-reproducing automata, we argue that sustainable recursive self-improvement in Large Language Models (LLMs) requires a functional analogue: introspection -- the system's capacity to simulate its own operations and target modifications. Grounded in Kleene's Second Recursion Theorem, we demonstrate the theoretical existence of such introspective programs. However, an empirical review reveals that while current LLMs exhibit quasi-introspection (e.g., partial metacognition), they fall short of true introspection due to structural bottlenecks: a lack of complete self-access, the feedforward nature of the Transformer, and computational class constraints that prevent fixed-point iteration. We conclude by outlining architectural paths to cross this complexity threshold and discussing the associated safety implications.

physics.soc-ph

Ray Antenna Array Enhanced Low-Altitude ISAC: Performance Analysis and Beamforming Design

The low-altitude economy (LAE) heavily relies on aerial vehicles, yet these platforms remain vulnerable to environmental and security risks, necessitating robust airspace monitoring. Integrated sensing and communication (ISAC) as one of the key technologies of 6G provides potential solutions for safe LAE. However, conventional antenna arrays face limitations in cost, scalability, and coverage, especially directly above the base station, due to hardware complexity and degraded angular resolution. By exploiting the recently proposed ray antenna array (RAA), this paper considers a RAA-enhanced low-altitude ISAC system. RAA architecture employs multiple ray-arranged arrays directly connected without phase shifters, significantly reducing hardware costs while supporting flexible beamforming via dynamic ray selection. Moreover, RAA can provide uniform angular resolution and eliminates coverage holes, making it particularly suitable for low-altitude ISAC. In this paper, we formulate an optimization problem for joint ray selection and beamforming to enhance sensing coverage under communication constraints. An efficient alternating optimization algorithm is proposed to solve this problem. Analytical and simulation results demonstrate that RAA achieves higher sensing signal-to-noise ratio compared to traditional arrays, offering a cost-effective and high-performance solution for achieving low-altitude ISAC.

cs.IT

Partial Effective Information Decomposition for Synergistic Causality

Causality is a central topic in scientific inquiry, yet for complex systems, the identification and analysis of synergistic causation remain a challenging and fundamental problem. In the context of causal relations among multivariate variables, a decomposition framework grounded in interventionist causation is still lacking. To address this gap, this paper proposes Partial Effective Information Decomposition (PEID), a framework that decomposes the influence of multiple source variables on a target variable under maximum-entropy interventions into unique and synergistic information, thereby providing a unified and computable characterization of synergistic causal relations. Theoretically, in the three-variable case, the proposed framework is compatible with the major axioms of Partial Information Decomposition (PID). Empirically, under maximum-entropy interventions, correlations among input variables are removed, causing redundancy to vanish and thereby enabling PEID to compute synergistic relations. Furthermore, based on this framework, it is possible to define causal graphs containing hyperedges as well as downward causation, thus offering a unified toolkit for analyzing cross-scale and multivariate causal mechanisms in complex systems. Finally, applying the framework to a machine-learning-based air quality forecasting task on KnowAir-V2, we demonstrate that PEID can extract interpretable inter-station causal structures from a learned dynamical model. These results suggest that PEID provides a general interventionist information-theoretic tool for analyzing multivariate and synergistic causal mechanisms in complex systems.

stat.ML

Tianwen-2 target asteroid (469219) Kamo'oalewa probably develops an Itokawa-compositional but ultra-highly space-weathered surface

China's Tianwen-2 mission plans to return samples from a small, rapidly spinning Earth quasi-satellite (469219) Kamo'oalewa. Previous studies linked Kamo'oalewa to lunar composition and origin. Here, we propose another scenario. We reanalyzed the reflectance spectrum of Kamo'oalewa and obtained an absorption band center at 1.001+-0.028 um (error is 1sigma), consistent with LL chondrites. We then conducted space weathering (SW) experiments on meteorites and found that highly space-weathered LL chondrite powder (but not slab) successfully reproduced the reflectance spectrum of Kamo'oalewa. We further traced the dynamical origin of Kamo'oalewa and found that it probably originated from the v6 secular resonance, and more specifically, the Flora family. Kamo'oalewa exhibits a similar composition to Itokawa and 7 objects in the Flora family, but with a higher degree of space weathering. We, therefore, proposed that Kamo'oalewa probably originated from the Flora family and developed an Itokawa-compositional, highly space-weathered, fine-regolith-dominated surface.

astro-ph.EP

Verifiable Reasoning for LLM-based Generative Recommendation

Reasoning in Large Language Models (LLMs) has recently shown strong potential in enhancing generative recommendation through deep understanding of complex user preference. Existing approaches follow a {reason-then-recommend} paradigm, where LLMs perform step-by-step reasoning before item generation. However, this paradigm inevitably suffers from reasoning degradation (i.e., homogeneous or error-accumulated reasoning) due to the lack of intermediate verification, thus undermining the recommendation. To bridge this gap, we propose a novel \textbf{\textit{reason-verify-recommend}} paradigm, which interleaves reasoning with verification to provide reliable feedback, guiding the reasoning process toward more faithful user preference understanding. To enable effective verification, we establish two key principles for verifier design: 1) reliability ensures accurate evaluation of reasoning correctness and informative guidance generation; and 2) multi-dimensionality emphasizes comprehensive verification across multi-dimensional user preferences. Accordingly, we propose an effective implementation called VRec. It employs a mixture of verifiers to ensure multi-dimensionality, while leveraging a proxy prediction objective to pursue reliability. Experiments on four real-world datasets demonstrate that VRec substantially enhances recommendation effectiveness and scalability without compromising efficiency. The codes can be found at https://github.com/Linxyhaha/Verifiable-Rec.

cs.IR

Rethinking ANN-based Retrieval: Multifaceted Learnable Index for Large-scale Recommendation System

Approximate nearest neighbor (ANN) search is widely used in the retrieval stage of large-scale recommendation systems. In this stage, candidate items are indexed using their learned embedding vectors, and ANN search is executed for each user (or item) query to retrieve a set of relevant items. However, ANN-based retrieval has two key limitations. First, item embeddings and their indices are typically learned in separate stages: indexing is often performed offline after embeddings are trained, which can yield suboptimal retrieval quality-especially for newly created items. Second, although ANN offers sublinear query time, it must still be run for every request, incurring substantial computation cost at industry scale. In this paper, we propose MultiFaceted Learnable Index (MFLI), a scalable, real-time retrieval paradigm that learns multifaceted item embeddings and indices within a unified framework and eliminates ANN search at serving time. Specifically, we construct a multifaceted hierarchical codebook via residual quantization of item embeddings and co-train the codebook with the embeddings. We further introduce an efficient multifaceted indexing structure and mechanisms that support real-time updates. At serving time, the learned hierarchical indices are used directly to identify relevant items, avoiding ANN search altogether. Extensive experiments on real-world data with billions of users show that MFLI improves recall on engagement tasks by up to 11.8\%, cold-content delivery by up to 57.29\%, and semantic relevance by 13.5\% compared with prior state-of-the-art methods. We also deploy MFLI in the system and report online experimental results demonstrating improved engagement, less popularity bias, and higher serving efficiency.

cs.IR

Composition and Space Weathering Characteristics of Tianwen-2 Mission's First Target Near-Earth Asteroid (469219) Kamo`oalewa

The near-Earth asteroid Kamo`oalewa, a quasi-satellite of the Earth and the target for sample return by China's Tianwen-2 mission, exhibits distinctive spectral characteristics. This study re-analyzes the visible and near-infrared reflectance spectrum of Kamo`oalewa published by B. N. L. Sharkey et al. (2021), obtained using the Large Binocular Telescope, to infer its mineral composition and space weathering characteristics. Spectral similarity analysis is performed by comparing the spectrum of Kamo`oalewa to the mean spectra of various types in the Bus-DeMeo taxonomy to make a preliminary constraint on the combined characteristics of surface mineralogy and space weathering effects. To further characterize the mineral composition, a detailed analysis of the 1 {\mu}m band center is conducted based on spectral data below 1.25 {\mu}m that have higher signal-to-noise ratios. Empirical models for normalized spectra are developed to estimate the Is/FeO content. The results suggest that asteroid Kamo`oalewa has higher olivine abundance than that of typical S-type asteroids and the Moon, exhibiting an immature to submature degree of space weathering. These findings enhance our understanding of the evolution of similar quasi-satellites and provide important implication for the future exploration of Tianwen-2 mission.

astro-ph.EP

Generative Multi-modal Feedback for Singing Voice Synthesis Evaluation

Singing voice synthesis (SVS) has advanced significantly, enabling models to generate vocals with accurate pitch and consistent style. As these capabilities improve, the need for reliable evaluation and optimization becomes increasingly critical. However, current methods like reward systems often rely on single numerical scores, struggle to capture various dimensions such as phrasing or expressiveness, and require costly annotations, limiting interpretability and generalization. To address these issues, we propose a generative feedback (i.e., reward model) framework that provides multi-dimensional language and audio feedback for SVS assessment. Our approach leverages an audio-language model to generate text and audio critiques-covering aspects such as melody, content, and auditory quality. The model is fine-tuned on a hybrid dataset combining human music reactions and synthetic critiques from a MLLMs, enhancing diversity and linguistic richness. Quantitative experiments validate the effectiveness of the proposed dataset and training strategy, demonstrating that the framework produces musically accurate and interpretable evaluations suitable for guiding generative model improvement. The code is at [https://github.com/opendilab/VocalCritic](https://github.com/opendilab/VocalCritic)

cs.SD

Physics-Guided Inductive Spatiotemporal Kriging for PM2.5 with Satellite Gradient Constraints

High-resolution mapping of fine particulate matter (PM2.5) is a cornerstone of sustainable urbanism but remains critically hindered by the spatial sparsity of ground monitoring networks. While traditional data-driven methods attempt to bridge this gap using satellite Aerosol Optical Depth (AOD), they often suffer from severe, non-random data missingness (e.g., due to cloud cover or nighttime) and inversion biases. To overcome these limitations, this study proposes the Spatiotemporal Physics-Guided Inference Network (SPIN), a novel framework designed for inductive spatiotemporal kriging. Unlike conventional approaches, SPIN synergistically integrates domain knowledge into deep learning by explicitly modeling physical advection and diffusion processes via parallel graph kernels. Crucially, we introduce a paradigm-shifting training strategy: rather than using error-prone AOD as a direct input, we repurpose it as a spatial gradient constraint within the loss function. This allows the model to learn structural pollution patterns from satellite data while remaining robust to data voids. Validated in the highly polluted Beijing-Tianjin-Hebei and Surrounding Areas (BTHSA), SPIN achieves a new state-of-the-art with a Mean Absolute Error (MAE) of 9.52 ug/m^3, effectively generating continuous, physically plausible pollution fields even in unmonitored areas. This work provides a robust, low-cost, and all-weather solution for fine-grained environmental management.

cs.LG

Causal Emergence of Consciousness through Learned Multiscale Neural Dynamics in Mice

Consciousness spans macroscopic experience and microscopic neuronal activity, yet linking these scales remains challenging. Prevailing theories, such as Integrated Information Theory, focus on a single scale, overlooking how causal power and its dynamics unfold across scales. Progress is constrained by scarce cross-scale data and difficulties in quantifying multiscale causality and dynamics. Here, we present a machine learning framework that infers multiscale causal variables and their dynamics from near-cellular-resolution calcium imaging in the mouse dorsal cortex. At lower levels, variables primarily aggregate input-driven information, whereas at higher levels they realize causality through metastable or saddle-point dynamics during wakefulness, collapsing into localized, stochastic dynamics under anesthesia. A one-dimensional top-level conscious variable captures the majority of causal power, yet variables across other scales also contribute substantially, giving rise to high emergent complexity in the conscious state. Together, these findings provide a multiscale causal framework that links neural activity to conscious states.

q-bio.NC

Accelerating quantum adiabatic evolution with $\pi$-pulse sequences

In quantum information processing, the development of fast and robust control schemes remains a central challenge. Although quantum adiabatic evolution is inherently robust against control errors, it typically demands long evolution times. In this work, we propose to achieve rapid adiabatic evolution, in which nonadiabatic transitions induced by fast changes in the system Hamiltonian are mitigated by flipping the nonadiabatic transition matrix using $\pi$ pulses. This enables a faster realization of adiabatic evolution while preserving its robustness. We demonstrate the effectiveness of our scheme in both two-level and three-level systems. Numerical simulations show that, for the same evolution duration, our scheme achieves higher fidelity and significantly suppresses nonadiabatic transitions compared to the traditional STIRAP protocol.

quant-ph

Speeding up adiabatic holonomic quantum gates via $\pi$-pulse modulation

Holonomic quantum computation (HQC) offers an inherently robust approach to quantum gate implementation by exploiting quantum holonomies. While adiabatic HQC benefits from robustness against certain control errors, its long runtime limits practical utility due to increased exposure to environmental noise. Nonadiabatic HQC addresses this issue by enabling faster gate operations but compromises robustness. In this work, we propose a scheme for fast holonomic quantum gates based on the $\pi$-pulse method, which accelerates adiabatic evolution while preserving its robustness. By guiding the system Hamiltonian along geodesic paths in the parameter space and applying phase-modulating $\pi$ pulses at discrete points, we realize a universal set of holonomic gates beyond the conventional adiabatic limit. Our scheme allows for arbitrary single-qubit and two-qubit controlled gates within a single-loop evolution and provides additional tunable parameters for flexible gate design. These results demonstrate a promising path toward high-fidelity, fast, and robust quantum computation.

quant-ph

Optimizing Recall or Relevance? A Multi-Task Multi-Head Approach for Item-to-Item Retrieval in Recommendation

The task of item-to-item (I2I) retrieval is to identify a set of relevant and highly engaging items based on a given trigger item. It is a crucial component in modern recommendation systems, where users' previously engaged items serve as trigger items to retrieve relevant content for future engagement. However, existing I2I retrieval models in industry are primarily built on co-engagement data and optimized using the recall measure, which overly emphasizes co-engagement patterns while failing to capture semantic relevance. This often leads to overfitting short-term co-engagement trends at the expense of long-term benefits such as discovering novel interests and promoting content diversity. To address this challenge, we propose MTMH, a Multi-Task and Multi-Head I2I retrieval model that achieves both high recall and semantic relevance. Our model consists of two key components: 1) a multi-task learning loss for formally optimizing the trade-off between recall and semantic relevance, and 2) a multi-head I2I retrieval architecture for retrieving both highly co-engaged and semantically relevant items. We evaluate MTMH using proprietary data from a commercial platform serving billions of users and demonstrate that it can improve recall by up to 14.4% and semantic relevance by up to 56.6% compared with prior state-of-the-art models. We also conduct live experiments to verify that MTMH can enhance both short-term consumption metrics and long-term user-experience-related metrics. Our work provides a principled approach for jointly optimizing I2I recall and semantic relevance, which has significant implications for improving the overall performance of recommendation systems.

cs.IR

PCDCNet: A Surrogate Model for Air Quality Forecasting with Physical-Chemical Dynamics and Constraints

Air quality forecasting (AQF) is critical for public health and environmental management, yet remains challenging due to the complex interplay of emissions, meteorology, and chemical transformations. Traditional numerical models, such as CMAQ and WRF-Chem, provide physically grounded simulations but are computationally expensive and rely on uncertain emission inventories. Deep learning models, while computationally efficient, often struggle with generalization due to their lack of physical constraints. To bridge this gap, we propose PCDCNet, a surrogate model that integrates numerical modeling principles with deep learning. PCDCNet explicitly incorporates emissions, meteorological influences, and domain-informed constraints to model pollutant formation, transport, and dissipation. By combining graph-based spatial transport modeling, recurrent structures for temporal accumulation, and representation enhancement for local interactions, PCDCNet achieves state-of-the-art (SOTA) performance in 72-hour station-level PM2.5 and O3 forecasting while significantly reducing computational costs. Furthermore, our model is deployed in an online platform, providing free, real-time air quality forecasts, demonstrating its scalability and societal impact. By aligning deep learning with physical consistency, PCDCNet offers a practical and interpretable solution for AQF, enabling informed decision-making for both personal and regulatory applications.

cs.LG

Preserving Privacy and Utility in LLM-Based Product Recommendations

Large Language Model (LLM)-based recommendation systems leverage powerful language models to generate personalized suggestions by processing user interactions and preferences. Unlike traditional recommendation systems that rely on structured data and collaborative filtering, LLM-based models process textual and contextual information, often using cloud-based infrastructure. This raises privacy concerns, as user data is transmitted to remote servers, increasing the risk of exposure and reducing control over personal information. To address this, we propose a hybrid privacy-preserving recommendation framework which separates sensitive from nonsensitive data and only shares the latter with the cloud to harness LLM-powered recommendations. To restore lost recommendations related to obfuscated sensitive data, we design a de-obfuscation module that reconstructs sensitive recommendations locally. Experiments on real-world e-commerce datasets show that our framework achieves almost the same recommendation utility with a system which shares all data with an LLM, while preserving privacy to a large extend. Compared to obfuscation-only techniques, our approach improves HR@10 scores and category distribution alignment, offering a better balance between privacy and recommendation quality. Furthermore, our method runs efficiently on consumer-grade hardware, making privacy-aware LLM-based recommendation systems practical for real-world use.

cs.IR