SearcharxivSearch

arXiv subjects

Xinyang Zhao

Publications and source records attributed to Xinyang Zhao.

6 recordsLinked to original sources

Modeling Multi-Dimensional Cognitive States in Large Language Models under Cognitive Crowding

Modeling human cognitive states is essential for advanced artificial intelligence. Existing Large Language Models (LLMs) mainly address isolated tasks such as emotion analysis or stance detection, and fail to capture interactions among cognitive dimensions defined in psychology, including emotion, thinking style, stance, and intention. To bridge this gap, we construct CognitiveBench, the first benchmark with unified annotations across the above four dimensions. Experiments on CognitiveBench show that although LLMs perform well on single dimension tasks, their performance drops sharply in joint multi-dimensional modeling. Using Gromov $\delta$-hyperbolicity analysis, we find that CognitiveBench exhibits a strong hierarchical structure. We attribute the performance bottleneck to ``Cognitive Crowding'', where hierarchical cognitive states require exponential representational space, while the Euclidean space of LLMs grows only polynomially, causing representation overlap and degraded performance. To address this mismatch, we propose HyCoLLM, which models cognitive states in hyperbolic space and aligns LLM representations via Hyperbolic Guided Alignment Tuning. Results show that HyCoLLM substantially improves multi-dimensional cognitive understanding, allowing 8B parameter model to outperform strong baselines, including GPT-4o.

cs.CL

Data Agent: A Holistic Architecture for Orchestrating Data+AI Ecosystems

Traditional Data+AI systems utilize data-driven techniques to optimize performance, but they rely heavily on human experts to orchestrate system pipelines, enabling them to adapt to changes in data, queries, tasks, and environments. For instance, while there are numerous data science tools available, developing a pipeline planning system to coordinate these tools remains challenging. This difficulty arises because existing Data+AI systems have limited capabilities in semantic understanding, reasoning, and planning. Fortunately, we have witnessed the success of large language models (LLMs) in enhancing semantic understanding, reasoning, and planning abilities. It is crucial to incorporate LLM techniques to revolutionize data systems for orchestrating Data+AI applications effectively. To achieve this, we propose the concept of a 'Data Agent' - a comprehensive architecture designed to orchestrate Data+AI ecosystems, which focuses on tackling data-related tasks by integrating knowledge comprehension, reasoning, and planning capabilities. We delve into the challenges involved in designing data agents, such as understanding data/queries/environments/tools, orchestrating pipelines/workflows, optimizing and executing pipelines, and fostering pipeline self-reflection. Furthermore, we present examples of data agent systems, including a data science agent, data analytics agents (such as unstructured data analytics agent, semantic structured data analytics agent, data lake analytics agent, and multi-modal data analytics agent), and a database administrator (DBA) agent. We also outline several open challenges associated with designing data agent systems.

cs.DB

CRAFTS for HI cosmology: I. data processing pipeline and validation tests

We present the calibration procedures and validation of source measurement with the data of the Commensal Radio Astronomy FAST Survey (CRAFTS) for \HI intensity mapping by the Five-hundred-meter Aperture Spherical Radio Telescope (FAST). Using 70-hour drift-scan observation with the L-band (1.05-1.45GHz) 19-beam receiver, we obtain the data covering $270\,\rm deg^2$ sky area. We employ both the pulsar backend and the spectrum backend to calibrate the spectral time-ordered-data (TOD) before projecting them onto HEALPix maps. We produce calibrated TOD with frequency resolution of 30kHz and time resolution of 1s and the map data-cube with frequency resolution of 30kHz and spatial resolution of $2.95\,\rm arcmin^2$. We examine the pointing errors, noise overflow, RFI contamination and their effect on the data quality. The resulting noise level is $\sim$ 5.7mJy for the calibrated TOD and 1.6mJy for the map, consistent with the theoretical predictions within 5\% at RFI-free channels. We also validate the data by Principal Components Analysis (PCA) and find the residual map looks thermal noise dominated after removing 30 modes. We identify 447 isolated bright continuum sources in our data matching the NRAO-VLA Sky Survey (NVSS) catalog, with relative flux error of 8.3\% for TOD and 6.6\% for the map-level. We also measure the \HI emission of 90 galaxies with redshift $z<0.07$ and compare with \HI-MaNGA spectra, yielding an overall relative \HI integral flux error of 16.7\%. These results provide an important first step in assessing the feasibility of conducting cosmological \HI detection with CRAFTS.

astro-ph.CO

FAST Drift Scan Survey for HI Intensity Mapping. II. Stacking-based Beam Construction of the 19-feed Array at $1.4$ GHz

Neutral hydrogen (HI) intensity mapping (IM) presents great promise for future cosmological large-scale structure surveys. However, a major challenge for HIIM cosmological studies is to accurately subtract the foreground contamination. An accurate beam model is crucial for improving the quality of foreground subtraction. In this work, we develop a stacking-based beam reconstruction method utilizing the radio continuum point sources within the drift-scan field. Based on the Five-hundred-meter Aperture Spherical radio Telescope (FAST), we employ two sets of drift-scan survey data and merge the measurements to construct the beam patterns of the 19 FAST L-band feeds. To model the beams, we utilize the Zernike polynomial (ZP), which effectively captures asymmetric features of the main beam and the different side lobes. Due to the symmetric location of the beams, the main features of the beams are closely related to the distance from the center of the feed array, e.g., as the distance increases, side lobes become more pronounced. This modeling pipeline leverages the stable drift-scan data to extract beam patterns while accounting for and excluding the reflector's changing effects. It provides a more accurate measurement beam and a more precise model beam for FAST HIIM cosmology surveys.

astro-ph.IM

LLM-Enhanced Data Management

Machine learning (ML) techniques for optimizing data management problems have been extensively studied and widely deployed in recent five years. However traditional ML methods have limitations on generalizability (adapting to different scenarios) and inference ability (understanding the context). Fortunately, large language models (LLMs) have shown high generalizability and human-competitive abilities in understanding context, which are promising for data management tasks (e.g., database diagnosis, database tuning). However, existing LLMs have several limitations: hallucination, high cost, and low accuracy for complicated tasks. To address these challenges, we design LLMDB, an LLM-enhanced data management paradigm which has generalizability and high inference ability while avoiding hallucination, reducing LLM cost, and achieving high accuracy. LLMDB embeds domain-specific knowledge to avoid hallucination by LLM fine-tuning and prompt engineering. LLMDB reduces the high cost of LLMs by vector databases which provide semantic search and caching abilities. LLMDB improves the task accuracy by LLM agent which provides multiple-round inference and pipeline executions. We showcase three real-world scenarios that LLMDB can well support, including query rewrite, database diagnosis and data analytics. We also summarize the open research challenges of LLMDB.

cs.DB

FAST drift scan survey for HI intensity mapping: I. preliminary data analysis

This work presents the initial results of the drift-scan observation for the neutral hydrogen (HI) intensity mapping survey with the Five-hundred-meter Aperture Spherical radio Telescope (FAST). The data analyzed in this work were collected in night observations from 2019 through 2021. The primary findings are based on 28 hours of drift-scan observation carried out over seven nights in 2021, which covers $60\,{\rm deg}^2$ sky area. Our main findings are: (i) Our calibration strategy can successfully correct both the temporal and bandpass gain variation over the $4$-hour drift-scan observation. (ii) The continuum maps of the surveyed region are made with frequency resolution of $28$ kHz and pixel area of $2.95\,{\rm arcmin}^2$. The pixel noise levels of the continuum maps are slightly higher than the forecast assuming $T_{\rm sys}=20\,{\rm K}$, which are $36.0$ mK (for 10.0 s integration time) at the $1050$--$1150$ MHz band, and $25.9$ mK (for 16.7 s integration time) at the $1323$--$1450$ MHz band, respectively. (iii) The flux-weighted differential number count is consistent with the NRAO-VLA Sky Survey (NVSS) catalog down to the confusion limit $\sim7\,{\rm mJy}/{\rm beam}^{-1}$. (iv) The continuum flux measurements of the sources are consistent with that found in the literature. The difference in the flux measurement of $81$ isolated NVSS sources is about $6.3\%$. Our research offers a systematic analysis for the FAST HI intensity mapping drift-scan survey and serves as a helpful resource for further cosmology and associated galaxies sciences with the FAST drift-scan survey.

astro-ph.CO