SearcharxivSearch

arXiv subjects

Zijie Yu

Publications and source records attributed to Zijie Yu.

11 recordsLinked to original sources

The Hyperspherical Geometry of CLIP Latent Space: A Semantic Mixture Model

Contrastive Language-Image Pretraining (CLIP) representations form a semantic embedding space governed by cosine similarity, reflecting an intrinsic hyperspherical geometry. However, existing probabilistic interpretations typically rely on Gaussian assumptions, which fail to capture this directional and multimodal structure. We propose a principled density model for the CLIP latent space based on Mixtures of von Mises-Fisher (MovMF) distributions defined on the unit hypersphere. Using the Expectation-Maximization (EM) algorithm, we efficiently learn a probabilistic model in which each mixture component corresponds to a coherent semantic concept. This formulation yields a closed-form likelihood naturally aligned with hyperspherical geometry, enabling accurate and interpretable density estimation. Empirically, our model significantly improves long-tailed and out-of-distribution detection and provides a natural semantic decomposition, representing each embedding as a sparse probabilistic combination of interpretable concepts. These results suggest that CLIP latent space is more faithfully characterized as a hyperspherical semantic mixture rather than an isotropic Gaussian, establishing a simple and geometrically consistent probabilistic framework for modeling and understanding multimodal representations. Project page is available at https://xiaoyuzhizi.github.io/movmf-clip/.

cs.LG

InternBootcamp: Boosting LLM Reasoning with Verifiable Task Scaling

Large language models (LLMs) have revolutionized artificial intelligence by enabling complex reasoning capabilities. While recent advancements in reinforcement learning (RL) have primarily focused on domain-specific reasoning tasks (e.g., mathematics or code generation), real-world reasoning scenarios often require models to handle diverse and complex environments that narrow-domain benchmarks cannot fully capture. To address this gap, we present InternBootcamp, an open-source framework comprising 1000+ domain-diverse task environments specifically designed for LLM reasoning research. With these bootcamps, we further establish Bootcamp-Eval, an automatically generated benchmark for comprehensive performance assessment. Evaluation reveals that frontier models still underperform in many reasoning tasks, while training with InternBootcamp provides an effective way to significantly improve performance, leading to our 32B model that achieves stateof-the-art results on Bootcamp-Eval and excels on other established benchmarks. In particular, we validate that consistent performance gain come from including more training tasks, namely task scaling, over two orders of magnitude, offering a promising route towards capable reasoning generalist. All data and code are publicly available.

cs.CL

FedMABench: Benchmarking Mobile Agents on Decentralized Heterogeneous User Data

Mobile agents have attracted tremendous research participation recently. Traditional approaches to mobile agent training rely on centralized data collection, leading to high cost and limited scalability. Distributed training utilizing federated learning offers an alternative by harnessing real-world user data, providing scalability and reducing costs. However, pivotal challenges, including the absence of standardized benchmarks, hinder progress in this field. To tackle the challenges, we introduce FedMABench, the first benchmark for federated training and evaluation of mobile agents, specifically designed for heterogeneous scenarios. FedMABench features 6 datasets with 30+ subsets, 8 federated algorithms, 10+ base models, and over 800 apps across 5 categories, providing a comprehensive framework for evaluating mobile agents across diverse environments. Through extensive experiments, we uncover several key insights: federated algorithms consistently outperform local training; the distribution of specific apps plays a crucial role in heterogeneity; and, even apps from distinct categories can exhibit correlations during training. FedMABench is publicly available at: https://github.com/wwh0411/FedMABench with the datasets at: https://huggingface.co/datasets/wwh0411/FedMABench.

cs.AI

MobileA3gent: Training Mobile GUI Agents Using Decentralized Self-Sourced Data from Diverse Users

The advancement of mobile GUI agents has opened new opportunities for automating tasks on mobile devices. Training these agents requires large-scale high-quality data, which is prohibitively expensive when relying on human labor. Given the vast population of global mobile phone users, if automated data collection from them becomes feasible, the resulting data volume and the subsequently trained mobile agents could reach unprecedented levels. Nevertheless, two major challenges arise: (1) extracting user instructions without human intervention and (2) utilizing distributed user data while preserving privacy. To tackle these challenges, we propose MobileA3gent, a collaborative framework that trains mobile GUI Agents using decentralized self-sourced data from diverse users. The framework comprises two components, each targeting a specific challenge: (1) Auto-Annotation, which enables the automatic collection of high-quality datasets during users' routine phone usage with minimal cost. (2) FedVLM-A, which enhances federated VLM training under non-IID distributions by incorporating adapted global aggregation based on both episode-level and step-level variability. Extensive experiments prove that MobileA3gent achieves superior performance over traditional approaches at only 1% of the cost, highlighting its potential for real-world applications

cs.AI

Self-Evolving Multi-Agent Collaboration Networks for Software Development

LLM-driven multi-agent collaboration (MAC) systems have demonstrated impressive capabilities in automatic software development at the function level. However, their heavy reliance on human design limits their adaptability to the diverse demands of real-world software development. To address this limitation, we introduce EvoMAC, a novel self-evolving paradigm for MAC networks. Inspired by traditional neural network training, EvoMAC obtains text-based environmental feedback by verifying the MAC network's output against a target proxy and leverages a novel textual backpropagation to update the network. To extend coding capabilities beyond function-level tasks to more challenging software-level development, we further propose rSDE-Bench, a requirement-oriented software development benchmark, which features complex and diverse software requirements along with automatic evaluation of requirement correctness. Our experiments show that: i) The automatic requirement-aware evaluation in rSDE-Bench closely aligns with human evaluations, validating its reliability as a software-level coding benchmark. ii) EvoMAC outperforms previous SOTA methods on both the software-level rSDE-Bench and the function-level HumanEval benchmarks, reflecting its superior coding capabilities. The benchmark can be downloaded at https://yuzhu-cai.github.io/rSDE-Bench/.

cs.SE

The FRB-searching pipeline of the Tianlai Cylinder Pathfinder Array

This paper presents the design, calibration, and survey strategy of the Fast Radio Burst (FRB) digital backend and its real-time data processing pipeline employed in the Tianlai Cylinder Pathfinder array. The array, consisting of three parallel cylindrical reflectors and equipped with 96 dual-polarization feeds, is a radio interferometer array designed for conducting drift scans of the northern celestial semi-sphere. The FRB digital backend enables the formation of 96 digital beams, effectively covering an area of approximately 40 square degrees with 3 dB beam. Our pipeline demonstrates the capability to make automatic search of FRBs, detecting at quasi-real-time and classify FRB candidates automatically. The current FRB searching pipeline has an overall recall rate of 88\%. During the commissioning phase, we successfully detected signals emitted by four well-known pulsars: PSR B0329+54, B2021+51, B0823+26, and B2020+28. We report the first discovery of an FRB by our array, designated as FRB 20220414A. We also investigate the optimal arrangement for the digitally formed beams to achieve maximum detection rate by numerical simulation.

astro-ph.IM

A Fast Transient Backend to Detect FRBs with the Tianlai Dish Pathfinder Array

The Tianlai Dish Pathfinder array is a radio interferometer array consisting of 16 six meter dish antennas. The original digital backend integration time is at the seconds level, designed for HI intensity mapping experiment. A new digital backend with millisecond response is added to enable it to search for fast radio burst (FRB) during its observations. The design and calibration of this backend, and the real time search pipeline for it are described in this paper. It is capable of forming 16 digital beams for each linear polarisation, covering an area of 19.6 square degrees. The search pipeline is capable of searching for, recording and classifying FRBs automatically in real time. In commissioning, we succeeded in capturing the signal pulses from the pulsars PSR B0329+54 and B2021+51.

astro-ph.IM

The Tianlai dish array low-z surveys forecasts

We present the science case for surveys with the Tianlai dish array interferometer tuned to the $\left[ 1300, 1400 \right] \mathrm{MHz}$ frequency range. Starting from a realistic generation of mock visibility data according to the survey strategy, we reconstruct a map of the sky and perform a foreground subtraction. We show that a survey of the North Celestial Polar cap during a year of observing time and covering an area of $150 \, \mathrm{deg^2}$ would reach a sensitivity of $ 1.5-2 \, \mathrm{mK} $ per $1 \, \mathrm{MHz} \times 0.25^2 \, \mathrm{deg^2 }$ voxel and be marginally impacted by mode-mixing. Tianlai would be able to detect a handful $(\sim 10)$ of nearby massive \HI clumps as well as a very strong cross-correlation signal of 21\,cm intensity maps with the North Celestial Cap Survey optical galaxies. We have also studied the performance of a mid-latitude survey, covering $\sim 1500 \, \mathrm{deg^2}$ centered on a declination of $\delta=55^\circ$, which overlaps the Sloan Digital Sky Survey footprint. Despite a higher noise level for the mid-latitude survey, as well as significant distortions due to mode mixing, Tianlai would be able to detect a highly significant cross-correlation between the 21\,cm signal and the Sloan spectroscopic galaxy sample. Using the extragalactic signals from either or both of these surveys, it will be possible to assess the impact of calibration uncertainties, antenna pattern uncertainties, sources of noise, and mode mixing for future surveys requiring higher sensitivity.

astro-ph.CO

The Tianlai Dish Pathfinder Array: design, operation and performance of a prototype transit radio interferometer

The Tianlai Dish Pathfinder Array is a radio interferometer designed to test techniques for 21~cm intensity mapping in the post-reionization universe as a means for measuring large-scale cosmic structure. It performs drift scans of the sky at constant declination. We describe the design, calibration, noise level, and stability of this instrument based on the analysis of about $\sim 5 \%$ of 6,200 hours of on-sky observations through October, 2019. Beam pattern determinations using drones and the transit of bright sources are in good agreement, and compatible with electromagnetic simulations. Combining all the baselines, we make maps around bright sources and show that the array behaves as expected. A few hundred hours of observations at different declinations have been used to study the array geometry and pointing imperfections, as well as the instrument noise behaviour. We show that the system temperature is below 80~K for most feed antennas, and that noise fluctuations decrease as expected with integration time, at least up to a few hundred seconds. Analysis of long integrations, from 10 nights of observations of the North Celestial Pole, yielded visibilities with amplitudes of 20-30~mK, consistent with the expected signal from the NCP radio sky with $<10\,$mK precision for $1 ~\mathrm{MHz} \times 1~ \mathrm{min}$ binning. Hi-pass filtering the spectra to remove smooth spectrum signal yields a residual consistent with zero signal at the $0.5\,$mK level.

astro-ph.IM

Reflections and Standing Waves on the Tianlai Cylinder Array

In 21~cm intensity mapping, the spectral smoothness of the foreground is exploited to separate it from the much weaker 21~cm signal. However, the non-smooth frequency response of the instrument complicates this process. Reflections and standing waves generate modulations on the frequency response. Here we report the analysis of the standing waves in the bandpass of the signal channels of the Tianlai Cylinder Array. By Fourier transforming the bandpass into the delay time domain, we find various standing waves generated on the telescope. A standing wave with time delay at about 142 ns is most clearly identified which is produced in the 15 meter feed cable. We also find a strong peak at a shorter delay of $\tau < 50 \ns$, which may be a mix of the standing wave between the reflector and feed, and the standing wave on the 4 m intermediate frequency (IF) cable. We also show that a smoother frequency response could be partially recovered by removing the reflection-inducted modulations. However, the standing wave on the antenna is direction-dependent, which poses a more difficult challenge for high precision calibration.

astro-ph.IM

The Tianlai Cylinder Pathfinder Array: System Functions and Basic Performance Analysis

The Tianlai Cylinder Pathfinder is a radio interferometer array designed to test techniques for 21 cm intensity mapping in the post-reionization Universe, with the ultimate aim of mapping the large scale structure and measuring cosmological parameters such as the dark energy equation of state. Each of its three parallel cylinder reflectors is oriented in the north-south direction, and the array has a large field of view. As the Earth rotates, the northern sky is observed by drift scanning. The array is located in Hongliuxia, a radio-quiet site in Xinjiang, and saw its first light in September 2016. In this first data analysis paper for the Tianlai cylinder array, we discuss the sub-system qualification tests, and present basic system performance obtained from preliminary analysis of the commissioning observations during 2016-2018. We show typical interferometric visibility data, from which we derive the actual beam profile in the east-west direction and the frequency band-pass response. We describe also the calibration process to determine the complex gains for the array elements, either using bright astronomical point sources, or an artificial on site calibrator source, and discuss the instrument response stability, crucial for transit interferometry. Based on this analysis, we find a system temperature of about 90 K, and we also estimate the sensitivity of the array.

astro-ph.IM