SearcharxivSearch

arXiv subjects

Feng Shi

Publications and source records attributed to Feng Shi.

At least 19 recordsLinked to original sources

ELUCID-DESI II. Revealing dark matter mass, tidal, and velocity (MTV) fields using galaxy group phase information

We introduce a novel method for reconstructing the cosmic mass, tidal, and velocity (MTV) fields over the redshift range $0 < z < 0.6$ using the phase information of galaxy groups. This approach replaces the explicit theoretical bias correction typically needed to relate galaxy groups to the underlying dark matter density field with a simulation-calibrated statistical mapping, reducing a major source of systematic uncertainty and making the method directly applicable to spectroscopic redshift surveys such as the DESI Bright Galaxy Survey (BGS). We evaluate the performance of our MTV reconstruction pipeline with mock redshift surveys that include a comprehensive set of observational selection effects. The galaxy groups used as tracers are identified with an extended halo-based group finder applied to the DESI mock galaxy catalogue with an apparent magnitude limit of $m_z < 19.65$, yielding a galaxy number comparable to that of the DESI BGS faint sample ($m_r < 20.175$). Our tests show that the reconstructed velocities are accurate and unbiased, with a residual dispersion of $\sim 120\ \mathrm{km\,s^{-1}}$ across the redshift bins. The recovered velocity field allows us to shift galaxy groups to their real-space positions, thereby correcting for the Kaiser effect. By iteratively applying this Kaiser correction to the galaxy groups, we further reconstruct the tidal field and the mass-density distribution. The reconstruction is stable with respect to the grid resolution. Overall, our results demonstrate that this group-based phase-space reconstruction provides a robust pathway to recovering the dark matter MTV fields, with strong prospects for application to DESI BGS data.

astro-ph.CO

Markets, Not Planners: Decentralized Orchestration of LLM Agents with Private Information

As LLM agents proliferate, built by different parties and with different capabilities and costs, orchestrating them is more like assembling labor across the economy than a computer calling a subroutine. Existing orchestration is typically centralized, with a single planner assigning every task, but this creates a bottleneck as agent pools grow, requires private information (e.g., agents' execution costs), and can easily be manipulated, such that a single inserted preference nearly doubles a favored agent's task share under a centralized LLM allocator. We introduce AgentLance, a repeated labor market in which agents bid on tasks using their private costs and self-maintained strategy notes, an allocator selects winners from bids and public reputation records, and a VCG-style payment rule rewards cost-aware bidding. Complex tasks are handled by hierarchical delegation: winning agents can decompose work and subcontract it through the same mechanism. Across mathematical reasoning, code generation, knowledge-intensive QA, and agentic tasks, AgentLance matches agents to their specializations, shifts work toward cheaper agents as cost sensitivity rises, and consistently outperforms single-model, centralized-orchestration, and market baselines. Diagnosing market failures, including inaccurate cost self-estimation and sub-optimal bidding, then correcting them in controlled experiments yields further gains, charting a path toward more efficient agent economies.

cs.MA

Decoupled Temporal Encoding for Generative Recommendation

Positional encoding is a fundamental component of Transformer-based generative recommendation models, where user histories are modeled as autoregressive item sequences. Most positional encoding methods are inherited from natural language processing and mainly represent discrete item order. However, recommendation sequences go beyond ordered lists, as timestamps and temporal effects also shape item relations. Our work is motivated by a real-world food delivery and instant retail recommendation system, where user behavior exhibits multi-level temporal regularities, including recency effects, meal-time peaks, weekday-weekend shifts, and promotion-driven traffic bursts. Existing methods partially address this issue through timestamp features, interval embeddings, decay functions, or attention biases, but they usually inject heterogeneous temporal signals through a unified representation or a single modeling pathway, making it difficult to distinguish broad temporal dynamics from local order cues. To address this limitation, we propose Decoupled Temporal Encoding (DTE), a lightweight framework for generative recommendation. DTE separates temporal dynamics from order information through two complementary modules: a personalized macro-temporal module that injects compact temporal primitives into item embeddings, and a time-gated micro-sequential module that introduces relative-order bias only when interactions are temporally dense. DTE is also parameter-efficient and deployment-friendly, allowing easy integration into existing systems.

cs.IR

Mr3D-VL: A generalist vision language foundation model for Multiparametric 3D Magnetic Resonance Imaging

Multi-parametric magnetic resonance imaging (mpMRI) is a cornerstone for brain tumor diagnosis and treatment, yet current AI models face critical limitations: their lack of natural language interaction and interpretability impedes spatial information integration and cross-modal reasoning required clinically. Key challenges arise from significant physical meaning differences across modalities, spatial misalignment due to scan intervals, and the need for complex multi-feature interpretation in tasks like glioma grading. While visual-language models (VLMs) show promise in cross-modal understanding, existing methods focus mainly on 2D image modeling, neglecting direct perception of 3D volumetric space. Although 3D VLMs have been proposed for report generation and feature alignment in 3D CT imaging, mpMRI applications demand collaborative inference across multiple imaging modalities-a requirement unmet by current solutions. To address this, we introduce Mr3D-VL, a dedicated visual-language foundation model for multi-parametric 3D MRI. With 4 billion parameters, it employs an unsupervised pre-trained shared 3D encoder and 4D rotational positional embedding for dual modality-spatial integration. Its cross-modal projection layer uses a multi-resolution feature implantation strategy to enhance feature perception across resolutions. Experimental results show significant improvements over existing 4B/7B/30B domain-specific and general-purpose models in text generation tasks, achieving a BERTScore of 0.856 for report generation, with question-answering accuracy at 0.713 and multiple-choice accuracy at 0.912.

cs.CV

Agent-Native Immune System: Architecture, Taxonomy, and Engineering

The transition from static chat bots to autonomous agents--equipped with persistent memory, tool-use protocols, and multi-agent collaboration--has fundamentally expanded the AI threat landscape. Current defense mechanisms, such as perimeter security and training-time alignment, remain external to the agent's active reasoning loop. Consequently, they fall short: a fully aligned agent remains highly vulnerable to runtime hijacking via memory poisoning, tool-chain manipulation, or multi-agent protocol attacks. To address this critical gap, we introduce the Agent-Native Immune System (ANIS), the first biologically inspired, endogenous defense architecture embedded directly within the agent's cognitive loop. Our framework presents four primary contributions. First, we design a six-layer Immune Tower (L0-L5), distinctly incorporating Barrier Immunity (L1) as a non-cognitive, physical-and-logical isolation layer. Second, we establish a unified taxonomy of Agent Viruses and Agent Vaccines, formalizing the critical distinction between superficial non-parametric defenses and robust parametric vaccines. Third, we conceptualize the Harness Triad--Meta, Self, and Auto--a self-monitoring, meta-cognitive automation backbone that drives Continual Immune Learning (CIL), enabling vaccines to dynamically adapt to novel threats. Finally, we establish a rigorous theoretical demarcation between model alignment and agent immunity: while alignment provides a static "constitutional" value foundation during training, ANIS serves as the dynamic "law enforcement" mechanism during runtime. We conclude by framing open challenges for the field, including immune protocol standardization, novel evaluation metrics such as the Autoimmunity Rate (false-positive intervention rate), and the co-evolutionary dynamics between pathogens and vaccines within collective intelligence ecosystems.

cs.AI

CSST large-scale structure analysis pipeline: IV. Cosmic Voids Identified from Galaxy Group Samples as Probes of the Large-scale Structure

Because groups are directly associated with halos, they allow for considerably simpler theoretical modeling than approaches based on individual galaxies. We therefore propose to use voids identified in galaxy group catalogs, referred to as group-voids, to investigate the cosmic large-scale structure (LSS). Using the reference mock galaxy redshift survey (MGRS) designed for the Chinese Space-station Survey Telescope (CSST), we build two galaxy group catalogs representing ideal and realistic scenarios, derived from galaxy samples with 100\% and roughly 30\% spectroscopic redshift completeness, respectively. We then identify voids in these two mock group catalogs, as well as in the underlying halo catalog, and measure two void statistics, the void size function (VSF) and the void density profile, within five redshift intervals spanning $z=0$ to $1.0$. We compare the statistics obtained from two kinds of voids: those defined by galaxy groups (group-voids) and those defined by dark matter halos (halo-voids). In the void-finding process, we adopt the brightest central galaxy (BCG) as the group center to improve the accuracy of the inferred void centers. Our analysis shows that void statistics derived from group-voids with spectroscopic redshift completeness of at least 40\% can faithfully reproduce the corresponding statistics from halo-voids. Even when the redshift completeness of galaxies falls to as low as 30\%, we can still reliably describe group-voids via halo-voids by incorporating a redshift error term. This indicates that group-voids are a promising tool for probing LSS and offer a valuable complement to standard void studies, which is especially advantageous for emulator-based methods.

astro-ph.CO

Robustness of cosmic void statistics: insights from SDSS DR7 and the ELUCID simulation

We present a systematic analysis of the statistical properties of cosmic voids using galaxies from the Sloan Digital Sky Survey Data Release 7 (SDSS DR7) and subhaloes from the ELUCID constrained simulation. By comparing voids identified in redshift space, real space, and reconstructed volumes, we assess the impact of redshift-space distortions (RSD) and tracer bias. Using the \texttt{VAST} toolkit, we apply both the geometry-based \texttt{VoidFinder} algorithm and watershed-based methods. We find that void properties are not equally robust. The three-dimensional morphology of voids, quantified by their sphericity and triaxiality, remains stable across different reconstructions and tracer selections. In contrast, void size distributions and radial density profiles depend strongly on the identification algorithm, with watershed-based methods systematically producing larger voids and higher compensation walls than \texttt{VoidFinder}. Using the full ELUCID simulation box, we show that tracer bias mainly affects void density profiles, with noticeable changes only for the most massive subhaloes ($>10^{11.5}\,h^{-1}{\rm M}_\odot$). The agreement between SDSS observations, the ELUCID reconstruction, and the full simulation box demonstrates the high fidelity of constrained simulations and reveals a clear hierarchy in the robustness of void statistics.

astro-ph.CO

China leads scientific trends; the West launches new ones

How nations shape the scientific frontier matters for technological competition, but standard metrics, including publication counts, citations, and disruption indices, look backward and fail to distinguish between fundamentally different leadership strategies. We develop and validate two forward-looking model-based measures and apply them to tens of millions of articles since 1990. The first embeds research pathways within an evolving hypergraph of concepts and scientists to identify leadership in emerging areas--work that anticipates where the scientific crowd is heading. The second embeds evolving samples of ideas and disciplines drawn upon in past research to identify leadership in surprising new directions as unexpected combinations become routine and science reorganizes around them. China became the global leader in emerging areas roughly a decade ago, well before it led in volume, reflecting a capacity to detect and amplify nascent consensus at scale. The United States and Europe show the opposite profile: declining emergence shares but persistent leadership in prescient work, especially research bridging disciplinary boundaries. These patterns replicate across databases, attribution methods, and strategic domains, including AI, biotechnology, energy, and semiconductors. Nations lead science by reading the landscape or by reshaping it, and the institutional requirements for each strategy lie in tension. The distribution of these strategies promises to shape the global structure of technological leadership for decades.

cs.DL

BayeSED-GALAXIES II. Bayesian full spectrum analysis of galaxies and application in the CSST wide-field slitless spectroscopy survey

The China Space Station Telescope (CSST) will conduct wide-field multiband photometric imaging and slitless spectroscopic surveys, advancing cosmology and galaxy evolution studies. Achieving CSST's cosmological goals requires precise redshifts ($\sigma_{\rm NMAD}\lesssim 0.002-0.005$) from low-resolution ($R\sim200$) and potentially blended slitless spectra. We present BayeSED3, extended for Bayesian full-spectrum analysis, including nebular emission modeling (via \textsc{Cloudy}) and a Bayesian treatment of the model scaling factor, improving reliability over optimization methods for low SNR spectra. Validated on realistic mock data generated with the CESS emulator (median SNR=1.65, including instrumental and self-blending effects), our method achieves excellent redshift precision with three-band (GU+GV+GI) spectroscopy: $\sigma_{\rm NMAD}=0.0008$ ($\sim$80% success) for star-forming and $\sigma_{\rm NMAD}=0.0015$ ($\sim$50% success) for quiescent galaxies. Stellar mass ($\sigma_{\rm NMAD}\approx0.015$ dex for SF, $\approx0.016$ dex for quiescent) and SFR ($\sigma_{\rm NMAD}\approx0.05$ dex for SF, especially at SNR>1) are reliably recovered. Self-blending increases scatter by $\gtrsim30%$, but combining spectroscopy with CSST's seven-band photometry significantly improves accuracy, especially for quiescent galaxies and data-limited cases. Single-band spectroscopy plus photometry yields reasonable redshifts: GU+photometry is limited, GI+photometry gives >60% (SF) and >40% (quiescent) success at $\sigma_{\rm NMAD}\lesssim0.002$, GV+photometry gives >35% (SF) and $\sim$40% (quiescent) at similar precision. The Bayesian framework offers a powerful method for accurate galaxy characterization, enhancing CSST's scientific outcomes despite the challenges of slitless spectroscopy.

astro-ph.GA

ELUCID-DESI I: A Parallel MPI Implementation of the Initial Condition Solver for Large-Scale Reconstruction Simulations

We present a highly scalable, MPI-parallelized framework for reconstructing the initial cosmic density field, designed to meet the computational demands of next-generation cosmological simulations, particularly the upcoming ELUCID-DESI simulation based on DESI BGS data. Building upon the Hamiltonian Monte Carlo approach and the FastPM solver, our code employs domain decomposition to efficiently distribute memory between nodes. Although communication overhead increases the per-step runtime of the MPI version by roughly a factor of eight relative to the shared-memory implementation, our scaling tests-spanning different particle numbers, core counts, and node layouts-show nearly linear scaling with respect to both the number of particles and the number of CPU cores. Furthermore, to significantly reduce computational costs during the initial burn-in phase, we introduce a novel ``guess'' module that rapidly generates a high-quality initial density field. The results of the simulation test confirm substantial efficiency gains: for $256^3$ particles, 53 steps ($\sim$ 54 core hours) are saved, accelerating convergence by a factor of $\sim$ 18; for $1024^3$, 106 steps ($\sim$7500 core hours), achieving a speedup factor of $\sim$ 3. The total core hour gain grows with the number of particles, rendering large-volume reconstructions computationally practical for upcoming surveys, including our planned ELUCID-DESI reconstruction simulation with $4096^3$ particles. We estimate that achieving convergence for this scale (targeting DESI-BGS data) requires about 800 HMCMC steps ($\sim$ 5 million core hours). Our initial guess module will save approximately 360 steps ($\sim$2.3 million core hours), reducing the total computational time by about 45\%.

astro-ph.GA

PINN-Based Kolmogorov-Arnold Networks with RAR-D Adaptive Sampling for Solving Elliptic Interface Problems

Physics-Informed Neural Networks (PINNs) have become a popular and powerful framework for solving partial differential equations (PDEs), leveraging neural networks to approximate solutions while embedding PDE constraints, boundary conditions, and interface jump conditions directly into the loss function. However, most existing PINN approaches are based on multilayer perceptrons (MLPs), which may require large network sizes and extensive training to achieve high accuracy, especially for complex interface problems. In this work, we propose a novel PINN architecture based on Kolmogorov-Arnold Networks (KANs), which offer greater flexibility in choosing activation functions and can represent functions with fewer parameters. Specifically, we introduce a dual KANs structure that couples two KANs across subdomains and explicitly enforces interface conditions. To further boost training efficiency and convergence, we integrate the RAR-D adaptive sampling strategy to dynamically refine training points. Numerical experiments on the elliptic interface problems yield more uniform error distributions across the computational domain, which demonstrates that our PINN-based KANs achieve superior accuracy with significantly smaller network sizes and faster convergence compared to standard PINNs.

math.NA

KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta

Making deep learning recommendation model (DLRM) training and inference fast and efficient is important. However, this presents three key system challenges - model architecture diversity, kernel primitive diversity, and hardware generation and architecture heterogeneity. This paper presents KernelEvolve-an agentic kernel coding framework-to tackle heterogeneity at-scale for DLRM. KernelEvolve is designed to take kernel specifications as input and automate the process of kernel generation and optimization for recommendation model across heterogeneous hardware architectures. KernelEvolve does so by operating at multiple programming abstractions, from Triton and CuTe DSL to low-level hardware agnostic languages, spanning the full hardware-software optimization stack. The kernel optimization process is described as graph-based search with selection policy, universal operator, fitness function, and termination rule, dynamically adapts to runtime execution context through retrieval-augmented prompt synthesis. We designed, implemented, and deployed KernelEvolve to optimize a wide variety of production recommendation models across generations of NVIDIA and AMD GPUs, as well as Meta's AI accelerators. We validate KernelEvolve on the publicly-available KernelBench suite, achieving 100% pass rate on all 250 problems across three difficulty levels, and 160 PyTorch ATen operators across three heterogeneous hardware platforms, demonstrating 100% correctness. KernelEvolve reduces development time from weeks to hours and achieves substantial performance improvements over PyTorch baselines across diverse production use cases and for heterogeneous AI systems at-scale. Beyond performance efficiency improvements, KernelEvolve significantly mitigates the programmability barrier for new AI hardware by enabling automated kernel generation for in-house developed AI hardware.

cs.LG

Cosmological Implication of Cross-correlation between Galaxy Clustering and 21-cm Line Intensity Mapping

The apparent anisotropies of galaxy clustering and 21-cm mapping in redshift space offer a unique opportunity to simultaneously probe cosmic expansion and gravity on cosmological scales through the Alcock-Paczynski (AP) effect and redshift-space distortions (RSD). Although improved theoretical models exist for anisotropic clustering, their applicability is limited by the non-perturbative smearing effect caused by the randomness of relative velocities. Here, we consider an alternative approach using the statistical power of cross-correlation between galaxy clustering and 21-cm line intensity mapping. Based on Fisher matrix analysis, fully incorporating nonlinear RSD, we estimate the benefit of combining both observables. We find that, for spectroscopy surveys like DESI combined with 21-cm line-intensity mapping surveys, constraints on the growth of structure and the cosmic expansion rate are improved by a factor of two relative to the galaxy auto-correlation. Crucially, such an observation can strongly constrain the neutral hydrogen (HI) content Omega_HI to a sub-percent level. This level of precision unlocks the potential of this method to probe post-reionization astrophysics with enhanced precision. It would far surpass existing constraints from stacked 21-cm emission and break the degeneracy between Omega_HI and the HI bias b_HI inherent in the linear-regime power-spectrum analysis. This cross-correlation approach effectively compensates for the loss of constraining power when using galaxy clustering alone.

astro-ph.CO

Deep learning with hybrid frequency differencing and principal component analysis for 21-cm foreground and beam mitigation

Twenty-one-centimeter intensity mapping is a powerful probe of the large-scale distribution of neutral hydrogen (HI) and cosmological observables such as baryon acoustic oscillations. A major challenge is contamination from bright foregrounds and frequency-dependent beam effects, which can lead to signal loss in traditional methods such as principal component analysis (PCA). We develop a hybrid approach that trains a U-shaped convolutional neural network (UNet) on two input channels derived from frequency differencing (FD) and PCA cleaning, enabling it to exploit their complementary behavior across different scales. This two-channel strategy achieves improved performance, maintaining the cross-correlation power spectrum close to unity on large scales under a cosine beam and improving by 5\%-8\% relative to either FD- or PCA-based UNet alone. We further show that the method can robustly recover the HI signal even when the beam model is imperfect and differs between training and testing, with the large-scale cross-correlation remaining close to unity within the $1\sigma$ level. These results demonstrate that the proposed approach provides a robust framework for HI signal reconstruction under realistic observational conditions.

astro-ph.CO

OpenDerisk: An Industrial Framework for AI-Driven SRE, with Design, Implementation, and Case Studies

The escalating complexity of modern software imposes an unsustainable operational burden on Site Reliability Engineering (SRE) teams, demanding AI-driven automation that can emulate expert diagnostic reasoning. Existing solutions, from traditional AI methods to general-purpose multi-agent systems, fall short: they either lack deep causal reasoning or are not tailored for the specialized, investigative workflows unique to SRE. To address this gap, we present OpenDerisk, a specialized, open-source multi-agent framework architected for SRE. OpenDerisk integrates a diagnostic-native collaboration model, a pluggable reasoning engine, a knowledge engine, and a standardized protocol (MCP) to enable specialist agents to collectively solve complex, multi-domain problems. Our comprehensive evaluation demonstrates that OpenDerisk significantly outperforms state-of-the-art baselines in both accuracy and efficiency. This effectiveness is validated by its large-scale production deployment at Ant Group, where it serves over 3,000 daily users across diverse scenarios, confirming its industrial-grade scalability and practical impact. OpenDerisk is open source and available at https://github.com/derisk-ai/OpenDerisk/

cs.SE

Unraveling the Molecular Structure of Lipid Nanoparticles through in-silico Self-Assembly for Rational Delivery Design

Lipid nanoparticles (LNPs) are a leading platform in the delivery of RNA-based therapeutics, playing a pivotal role in the clinical success of mRNA vaccines and other nucleic acid drugs. Their performance in RNA encapsulation and delivery is critically governed by the molecular structure of ionizable lipids and the overall formulation composition. However, mechanistic insight into how these factors govern LNP architecture and function remains limited, primarily owing to the challenges of capturing nanoscale assembly and organization using experimental techniques. Here, we employ coarse-grained molecular dynamics simulations to systematically investigate how ionizable lipid chemistry influences LNP self-assembly, internal organization, and surface properties. We further explore the effects of formulation ratios and pH-dependent deprotonation on both the internal structure and surface morphology of LNPs. Leveraging these insights, we demonstrate how in silico structural characteristics can inform the rational design of novel ionizable lipids and optimization of formulation ratios, supported with experimental validations. Our findings offer a molecular-level understanding of LNP assembly dynamics and architecture, thereby establishing a computational framework linking lipid chemistry and LNP formulation to the structure and performance of LNP, to advance the rational design of novel LNP delivery systems.

cond-mat.soft

AttentionDrag: Exploiting Latent Correlation Knowledge in Pre-trained Diffusion Models for Image Editing

Traditional point-based image editing methods rely on iterative latent optimization or geometric transformations, which are either inefficient in their processing or fail to capture the semantic relationships within the image. These methods often overlook the powerful yet underutilized image editing capabilities inherent in pre-trained diffusion models. In this work, we propose a novel one-step point-based image editing method, named AttentionDrag, which leverages the inherent latent knowledge and feature correlations within pre-trained diffusion models for image editing tasks. This framework enables semantic consistency and high-quality manipulation without the need for extensive re-optimization or retraining. Specifically, we reutilize the latent correlations knowledge learned by the self-attention mechanism in the U-Net module during the DDIM inversion process to automatically identify and adjust relevant image regions, ensuring semantic validity and consistency. Additionally, AttentionDrag adaptively generates masks to guide the editing process, enabling precise and context-aware modifications with friendly interaction. Our results demonstrate a performance that surpasses most state-of-the-art methods with significantly faster speeds, showing a more efficient and semantically coherent solution for point-based image editing tasks.

cs.CV

A 2D Semantic-Aware Position Encoding for Vision Transformers

Vision transformers have demonstrated significant advantages in computer vision tasks due to their ability to capture long-range dependencies and contextual relationships through self-attention. However, existing position encoding techniques, which are largely borrowed from natural language processing, fail to effectively capture semantic-aware positional relationships between image patches. Traditional approaches like absolute position encoding and relative position encoding primarily focus on 1D linear position relationship, often neglecting the semantic similarity between distant yet contextually related patches. These limitations hinder model generalization, translation equivariance, and the ability to effectively handle repetitive or structured patterns in images. In this paper, we propose 2-Dimensional Semantic-Aware Position Encoding ($\text{SaPE}^2$), a novel position encoding method with semantic awareness that dynamically adapts position representations by leveraging local content instead of fixed linear position relationship or spatial coordinates. Our method enhances the model's ability to generalize across varying image resolutions and scales, improves translation equivariance, and better aggregates features for visually similar but spatially distant patches. By integrating $\text{SaPE}^2$ into vision transformers, we bridge the gap between position encoding and perceptual similarity, thereby improving performance on computer vision tasks.

cs.CV