SearcharxivSearch

arXiv subjects

Jiacheng Ding

Publications and source records attributed to Jiacheng Ding.

At least 19 recordsLinked to original sources

HoopMind: A Real-Time Neural Game-Tree System for Opponent-Aware Possession Planning

School coaches prepare for opponents with game film and intuition. The analytics tools of professional teams stay out of reach. We ask how far public data can close this gap. Professional basketball is our case study, chosen for its data rather than the league. We fuse five public sources into one per-shot dataset of 4.23M shots over 21 seasons. The sources are shot locations, two play-by-play feeds, official matchup tracking, and player biometrics. Alignment across them is 99.5% to 100%. We also report two data pitfalls that are easy to miss. We then model a half-court possession as a sequential game. Shot values come from ShotNet, an embedding multilayer perceptron (MLP). On a held-out season it beats a zone-rate baseline and a logistic baseline, and its probabilities are well calibrated. A depth-limited expectimax search then solves the offensive decision tree, with branch-and-bound pruning to keep it real time. All training runs offline, so the online system stays light. A scouting planner and a playable simulator both run in a single browser page.

cs.LG

HAWKEYE: Seeing One Layer Deeper -- A Cohesion-Aware Structural Channel for Temporal Link Prediction

State-of-the-art temporal-link-prediction (TLP) models are, in essence, multi-channel information aggregators: they combine an interaction-history channel, a time-encoding channel, and a structure channel. The first two have been refined relentlessly; the structure channel remains a crude afterthought -- DyGFormer encodes it as a 1--2-bit neighbour-cooccurrence count. We begin with a measurement: on sparse temporal graphs the classical 1-hop common-neighbour signal is near-random (discriminative AUC $\approx 0.50$), because two nodes almost never share a direct neighbour; the genuinely discriminative signal lies one hop deeper -- the 2-hop cohesive bridge, whose discAUC reaches 0.73--0.98, on both bipartite and non-bipartite graphs. Motivated by this, we propose HAWKEYE, a cohesion-aware structural channel that incrementally maintains the classical k-family of cohesiveness indicators (degree $\to$ k-core $\to$ k-truss) and forms 2-hop cohesive-bridge features. HAWKEYE is a drop-in replacement for a temporal-graph model's native structure channel, with no change to the backbone. Swapping HAWKEYE into DyGFormer improves test AP/MRR over the cooccurrence channel by +0.6 to +10.8 points across six multi-seed-validated datasets (uci, enron, USLegis, CanParl, reddit, mooc). On the bipartite recommendation benchmark tgbl-subreddit, a 3-seed single-pass struct-only ablation shows HAWKEYE nearly doubling the baseline test MRR (0.103$\pm$0.003 $\to$ 0.204$\pm$0.005, +10.1 points across all three seeds); the streaming pipeline scales to the 67M-edge tgbl-flight in five minutes per pass. We further characterise when it helps: the gain tracks a graph's training-free 2-hop discAUC and vanishes on degenerate or saturated graphs -- a predictable boundary. All code, data, and figure-generation scripts are released.

cs.DB

Development of A Novel Compton Camera for MeV Gamma-Ray Measurement in Space

The astrophysical gamma rays in the MeV energy region have not yet been well-explored due to the limitation of detection technology in the past decades, and the famous gamma-ray "MeV gap" exists. Opening the window of MeV gamma-ray is not only critical for the gamma astronomy but also essential for rich frontier researches in astro-particle physics, such as detecting light dark matter, probing the primordial black hole and better understanding of nucleosynthesis. As a pilot experiment of the project for dark matter detection in space at Shanghai Jiao Tong University, a three-layer Compton camera with the energy resolution better than 4% and position resolution of ~2 mm is developed utilizing the novel scintillators. Here we show the design, detailed calibration and validation results of the novel Compton camera, and demonstrate its good ability of MeV gamma-ray source imaging for the upcoming in-orbit mission.

hep-ex

TriAgent: Divergence-Aware Multi-Agent Committees for Cost-Efficient Financial Sentiment Analysis

Production LLM-based financial sentiment analysis faces a structural cost trap: most queries are trivially classifiable, yet expensive cloud reasoners process them all, and the bill scales linearly with user count. We present TriAgent, a multi-agent committee stratified by contextual granularity -- a word-level lexicon (VADER), a sentence-level domain transformer (FinBERT), and a cross-sentence reasoner (Qwen2.5, 0.5B-14B-4bit, with Mistral-7B and Phi-3.5-mini cross-family checks). A three-way Semantic Divergence Index (SDI) measures pairwise disagreement across granularities and routes each query accordingly. Our central finding is the critic plateau: when the LLM is re-tasked as a critic over the smaller agents' outputs, F1 plateaus at ~0.87 across 1.5B-7B Qwen (bootstrap 95% CIs overlap), while a same-size 3-persona vote drops to F1=0.66, which is driven by granularity-stratified diversity. Three corollaries follow from the same SDI signal: (i) a Shared Consensus Dictionary on multilingual sentence-BERT answers 95% of Chinese queries from an English cache at F1=0.99 -- cross-border canonicalization at zero marginal cost; (ii) SDI doubles as a post-hoc LLM-hallucination detector at AUC=0.90; (iii) the SDI single-stage strategy attains the best risk-adjusted return (Sharpe=3.50) on a 20-ticker back-test, dominating both always-FinBERT (1.36) and always-LLM (0.11). At 10M-user scale, TriAgent saves $9.3M/year vs. a GPT-4o-mini baseline. Code, lexicons, and the SCD are released.

cs.CL

FIFA World Cup 2026 as a Contamination-Free Benchmark for LLM Forecasting Agents: Four Models, a Bookmaker, and 104 Matches

We introduce WC2026-Agents, a benchmark and dataset for evaluating large language models (LLMs) as autonomous forecasting agents on real, future events. For every one of the 104 matches of the 2026 FIFA World Cup, four frontier models -- Claude Opus 4.8, ChatGPT (GPT-5.5, high reasoning), Gemini 3.1 Pro, and Grok (Expert Mode) -- ran an identical search-act-reflect loop: gather evidence with a web tool, commit to a 1X2 (team-A win / draw / team-B win) distribution and a virtual 100-USD bet, and, after the match, reflect given only the final score. Because every match kicked off after the models' training cutoffs, the benchmark is contamination-free by construction. Crucially, we pair the four agents with a fifth competitor drawn from the same information environment -- the pre-match betting market -- collected as per-match 1X2 odds, giving an economically grounded baseline and letting us score not just what an agent predicts but what it does with money. The release contains 416 forecasts and 414 reflections with verbatim reasoning, ground truth (including penalty shootouts), odds, and a reproducible evaluation suite. A reference evaluation surfaces findings that raw accuracy hides: the four agents issue an identical top pick in 92% of matches and none beats the market's Brier score; indeed, a naive flat stake on the market favorite out-earns all four agents. Yet the agents diverge sharply as decision-makers: betting return-on-investment ranges from -18% to +10%, fading the market is unprofitable for all four, the share of forecasts that cite the market ranges from 12% to 100%, and self-reported error rates on wrong picks range from 36% to 86%. The benchmark thus measures calibration, decision quality, and self-knowledge -- axes on which frontier models differ even when their predictions do not. Data and code: https://github.com/graphuofm/FIFA2026LLM

cs.LG

SAGA: Synthetic Agentic Graph Architecture for Temporal Benchmark Generation

High quality temporal graph benchmarks with rich semantics and ground-truth anomaly labels are essential for training graph neural networks, yet remain scarce due to privacy constraints and annotation costs. We present SAGA (Synthetic Agentic Graph Architecture), a system for generating large-scale, semantically rich temporal graphs via a four-phase pipeline. Our Skeleton-First, Semantics-Second architecture decouples structure from semantics: (S) an O(1)-per-edge skeleton generator produces power-law graphs; (A) a dispatcher partitions causally ordered time blocks for parallel execution; (G) LLM agents inject domain semantics using RAG-based rule bases across four domains; and (A) a state alignment engine resolves conflicts via temporal replay, yielding anomaly labels as natural byproducts. Unlike structural generators (e.g., LDBC SNB, Kronecker/R-MAT) or purely LLM-based approaches, SAGA achieves structural realism, semantic richness, and automatic anomaly labeling in a unified framework. On a single H100 GPU with vLLM batching, SAGA generates 500,000 temporal edges with controlled anomalies in under 90 minutes, scaling to 100,000 nodes while maintaining clustering coefficients above 0.99. The system supports real-time pipeline visualization, interactive multi-domain tuning (Finance/AML, Network/IDS, Cyber/APT, Transportation), and a CLI for large-scale GPU-based experiments.

cs.DB

The Shear-to-Cosmology Paradigm I. Hybrid Field-Level and Simulation-Based Framework for Weak Lensing Surveys

Precise cosmological inference from next-generation weak lensing surveys requires extracting non-Gaussian information beyond standard two-point statistics. We present a hybrid machine-learning (ML) framework that integrates field-level inference (FLI) with simulation-based inference (SBI) to map observed shear fields directly to cosmological parameters, eliminating the need for convergence reconstruction. The FLI network extracts rich non-Gaussian information from the shear field to produce informative features, which are then used by SBI to model the resulting complex posteriors. To mitigate noise from intrinsic galaxy shapes, we develop a blind, training-free, PCA-based shear denoising method. Tests on CSST-like mock catalogs reveal significant performance gains. The shear-based inference achieves approximately twice the cosmological constraining power in Figure of Merit (FoM) compared to the conventional convergence-based approach. Moreover, the combination of PCA denoising and ML compression can deliver a 36.4% improvement in FoM over standard shear two-point statistics. This work establishes a scalable and robust pathway for cosmological inference, unlocking the full potential of Stage-IV weak-lensing surveys.

astro-ph.CO

Enhancing cosmological constraints with nonlinear tanh transformations of Hermite-Gaussian Derivative fields

A key goal in large-scale structure analysis is to extract multi-scale information to improve cosmological parameter constraints. In particular, higher-order derivative fields are especially valuable as they capture the geometric and topological information of the cosmic web that is highly sensitive to cosmological parameters. Traditional derivative-based methods, such as finite-difference or Fourier approaches, suffer from noise amplification at small scales and cannot stably capture multi-scale features. We present a robust two-step framework: first, stable multi-scale arbitrary-order derivatives are obtained via Hermite-Gaussian convolutional filters that suppress small-scale noise; second, a tanh nonlinear transformation compresses extreme density contrasts and enhances the visibility of cosmic web structures. Using the Quijote simulations, we show that combining multi-scale first-order spectra yields improvements of 1.2-3.0 times across all seven cosmological parameters, while multi-order spectra at a fixed scale provide 1.3-2.9 times gains. The most comprehensive combination achieves nominal gains of 2.0-5.3 times. Our method offers a robust approach to extracting additional cosmological information for future surveys.

astro-ph.CO

ChronoConnect: Tracking Pathways Along Highly Dynamic Vertices in Temporal Graphs

With the proliferation of temporal graph data, there is a growing demand for analyzing information propagation patterns during graph evolution. Existing graph analysis systems, mostly based on static snapshots, struggle to effectively capture information flows along the temporal dimension. To address this challenge, we introduce ChronoConnect, a novel system that enables tracking temporal pathways in temporal graph, especially beneficial to downstream mining tasks, e.g., understanding what are the critical pathways in propagating information towards a specific group of vertices. Built on ChronoConnect, users can conveniently configure and execute a variety of temporal traversal algorithms to efficiently analyze information diffusion processes under time constraints. Moreover, ChronoConnect utilizes parallel processing to tackle the explosive size-growth of evolving graphs. We showcase the effectiveness and enhanced performance of ChronoConnect through the implementation of algorithms that track pathways along highly dynamic vertices in temporal graphs. Furthermore, we offer an interactive user interface for graph visualization and query result exploration. We envision ChronoConnect to become a powerful tool for users to examine how information spreads over a temporal graph.

cs.DB

A New Wavelet Scattering Transform-Based Statistic for Cosmological Analysis of Large-Scale Structure

Large-scale structure (LSS) analysis in galaxy surveys is a powerful cosmological probe but is limited by tracer bias, which can obscure underlying information and weaken parameter constraints. Existing methods either model bias or restrict analyses to low-density regions, yet their sensitivity to bias remains poorly understood. We propose a novel method based on the wavelet scattering transform (WST) to distinguish LSS across cosmological models while mitigating tracer bias. Central to our approach are the WST $m$-mode ratios, $R^{\rm wst}$, a new statistical measure, and a high-density apodization preprocessing that smoothly rescales extreme values. We use a reduced chi-square to assess the cosmological parameter constraints and find that $R^{\rm wst}$, in the scale range $j \in [3,7]$, achieves $χ^2_{ν, \rm cos} \approx 6$ for cosmology while maintaining $χ^2_{ν, \rm bias} \sim 1$--a regime unattained by other statistics. $R^{\rm wst}$ thus provides robust cosmological sensitivity with effective bias mitigation for future surveys.

astro-ph.CO

Tripod in uniform spanning tree and three-sided radial SLE$_2$

Fix a bounded $3$-polygon $(\Omega; x_1, x_2, x_3)$ with three marked boundary points $x_1, x_2, x_3\in\partial\Omega$ and suppose $(\Omega^{\delta}; x_1^{\delta}, x_2^{\delta}, x_3^{\delta})$ is an approximation of $(\Omega; x_1, x_2, x_3)$ on $\delta$-scaled hexagonal lattice. We consider uniform spanning tree (UST) in $\Omega^{\delta}$ with wired boundary conditions. Conditional on the event that both branches from $x_1^{\delta}$ and $x_2^{\delta}$ hit the boundary through $x_3^{\delta}$, the two branches meet at a point $\trifurcation^{\delta}$ which we call trifurcation, and the union of the three branches from $x_j^{\delta}$ to $\trifurcation^{\delta}$ form a tripod in the UST. We compute the scaling limit of the tripod: the distribution of trifurcation is absolutely continuous with respect to Lebesgue measure with explicit density; given the trifurcation, the conditional law of the tripod is three-sided radial SLE$_2$. The proof relies on construction of a new observable for trifurcation in our key lemma--Lemma~3.1--where we use Fomin's formula and the geometry of the hexagonal lattice in an essential way. Interestingly, the scaling limit of the observable for trifurcation coincides with the partition function for three-sided radial $\SLE_2$. Our result gives a probabilistic interpretation of the correlation function in CFT which has conformal weights $1$ at the three boundary points and has a spinless field of weights $(1,1)$ at the bulk point.

math.PR

Testing the Equivalence Principle on Cosmological Scales Using Peculiar Acceleration Power Spectrum

While the (weak) Equivalence Principle (EP) has been rigorously tested within the solar system, its validity on cosmological scales, particularly in the context of dark matter and dark energy, remains uncertain. In this study, we propose a novel method to test EP on cosmological scales by measuring the peculiar acceleration power spectrum of galaxies using the redshift drift technique. We develop an EP estimator, $E_{\rm ep}$, to evaluate the consistency of the peculiar acceleration power spectrum across different tracers. By calculating the ratio of the peculiar acceleration power spectra of tracers, the ensemble average of $E_{\rm ep}$ is expected to be unity if EP holds on cosmological scales for these tracers. We validate this estimator using N-body simulations, focusing on four redshift bins with $z\leq 1.5$ and scales of $k$ in the range of $0.007$ and $0.2$ $h/\rm Mpc$. By fitting a single parameter $δ_{\rm ep}$ across redshifts, we find that DM particle mocks without EP violation yield $δ_{\rm ep}$ consistent with zero under the small redshift measurement uncertainty case, while the large redshift uncertainty case slightly induces biases at low redshifts. In addition, when using DM halo mocks with controlled EP violations, no-violation and mild-violation cases show no significant detection, while moderate and strong violations produce statistically significant $δ_{\rm ep}$ values and high $χ^2$, especially at low redshifts, confirming the estimator's sensitivity. Taking advantage of advanced observing capabilities, such as next-generation facilities that extend beyond the Square Kilometer Array, the proposed method offers a promising approach for future cosmological tests of EP.

astro-ph.CO

AI-Powered Reconstruction of Dark Matter Velocity Fields from Redshift-Space Halo Distribution

We propose a UNet-based deep learning model to reconstruct the real-space dark matter (DM) velocity field from the redshift-space distribution of sparse DM halos. Using various statistical measures, we show that the reconstructed velocity components--including velocity magnitude, momentum, and divergence--closely match the ground truth, achieving better than 10% relative error and a correlation coefficient of 0.88. In the power spectrum comparison over $k \in [0.05, 0.3] h/{\rm Mpc}$, the UNet reconstruction outperforms linear theory and agrees with the true field within $2σ$. The model also effectively corrects redshift-space distortions (RSD), yielding unbiased power spectrum multipoles of DM fields within $2σ$. Notably, the UNet remains robust even with incomplete halo mass information. These results highlight the model's broad applicability to cosmological analyses, including RSD, cosmic web studies, the kinetic Sunyaev-Zel'dovich effect, and BAO reconstruction.

astro-ph.CO

Enhancing Cosmological Constraints by Two-dimensional $β$-cosmic-web Weighted Angular Correlation Functions

In this study, we investigate the potential of mark-weighted angular correlation functions (MACFs), which integrate $β$-cosmic-web classification with angular correlation function analysis to improve cosmological constraints. Using SDSS DR12 CMASS-NGC galaxies and mock catalogs with $Ω_m$ varying from 0.25 to 0.40, we assess the discriminative power of different statistics via the average improvement in chi-squared, $Δ\overline{χ^2}$, across six redshift bins. This metric quantifies how effectively each statistic distinguishes between different cosmological models. Incorporating cosmic-web weights leads to substantial improvements. Using statistics weighted by the mean neighbor distance ($\bar{D}_{\rm nei}$) increases $Δ\overline{χ^2}$ by approximately 40%-130%, while applying inverse mean neighbor distance weighting ($1/\bar{D}_{\rm nei}$) yields even larger gains, boosting $Δ\overline{χ^2}$ by a factor of 2-3 compared to traditional unweighted angular statistics. These enhancements are consistent with previous 3D clustering results, demonstrating the superior sensitivity of the $β$-weighted approaches. Our method, based on thin redshift slices, is particularly suited for slitless surveys (e.g., Euclid, CSST) where redshift uncertainties limit 3D analyses. This study also offers a framework for applying marked statistics to 2D angular clustering.

astro-ph.CO

Restoring Missing Modes of 21cm Intensity Mapping with Deep Learning: Impact on BAO Reconstruction

In 21cm intensity mapping of the large-scale structure (LSS), regions in Fourier space could be compromised by foreground contamination. In interferometric observations, this contamination, known as the foreground wedge, is exacerbated by the chromatic response of antennas, leading to substantial data loss. Meanwhile, the baryonic acoustic oscillation (BAO) reconstruction, which operates in configuration space to "linearize" the BAO signature, offers improved constraints on the sound horizon scale. However, missing modes within these contaminated regions can negatively impact the BAO reconstruction algorithm. To address this challenge, we employ the deep learning model U-Net to recover the lost modes before applying the BAO reconstruction algorithm. Despite hardware limitations, such as GPU memory, our results demonstrate that the AI-restored 21cm temperature map achieves a high correlation with the original signal, with a correlation ratio of approximately $0.9$ at $k \sim 1 h/Mpc$. Furthermore, subsequent BAO reconstruction indicates that the AI restoration has minimal impact on the performance of the `linearized' BAO signal, proving the effectiveness of the machine learning approach to mitigate the impact of foreground contamination. Interestingly, we demonstrate that the AI model trained on coarser fields can be effectively applied to finer fields, achieving even higher correlation. This success is likely attributable to the scale-invariance properties of non-linear mode coupling in large-scale structure and the hierarchical structure of the U-Net architecture.

astro-ph.CO

Recovering Cosmic Structure with a Simple Physical Constraint

Radio observation of the large-scale structure (LSS) of our Universe faces major challenges from foreground contamination, which is many orders of magnitude stronger than the cosmic signal. While other foreground removal techniques struggle with complex systematics, methods like foreground avoidance emerge as effective alternatives. However, this approach inevitably results in the loss of Fourier modes and a reduction in cosmological constraints. We present a novel method that, by enforcing the non-negativity of the observed field in real space, allows us to recover some of the lost information, particularly phase angles. We demonstrate that the effectiveness of this straightforward yet powerful technique arises from the mode mixing from the non-linear evolution of LSS. Since the non-negativity is ensured by mass conservation, one of the key principles of the cosmic dynamics, we can restore the lost modes without explicitly expressing the exact form of the mode mixing. Unlike previous methods, our approach utilizes information from highly non-linear scales, and has the potential to revolutionize the analysis of radio observational data in cosmology. Crucially, we demonstrate that in long-baseline interferometric observations, such as those from the Square Kilometre Array (SKA), it is still possible to recover the baryonic acoustic oscillation (BAO) signature despite not directly covering the relevant scales. This opens up potential future survey designs for cosmological detection.

astro-ph.CO

Deep learning for cosmological parameter inference from a dark matter halo density field

We propose a lightweight deep convolutional neural network (lCNN) to estimate cosmological parameters from simulated three-dimensional dark matter (DM) halo distributions and associated statistics. The training dataset comprises 2000 realizations of a cubic box with a side length of 1000 $h^{-1}{\rm Mpc}$, and interpolated over a cubic grid of $300^3$ voxels, with each simulation produced using $512^3$ DM particles and $512^3$ neutrinos. Under the flat $Λ$CDM model, simulations vary standard six cosmological parameters including $Ω_m$, $Ω_b$, $h$, $n_s$, $σ_8$, $w$, along with the neutrino mass sum, $M_ν$. We find that: 1) within the framework of lCNN, extracting large-scale structure information is more efficient from the halo density field compared to relying on the statistical quantities including the power spectrum, the two-point correlation function, and the coefficients from wavelet scattering transform; 2) combining the halo density field with its Fourier transformed counterpart enhances predictions, while augmenting the training dataset with measured statistics further improves performance; 3) achieving high accuracy in inferring $Ω_m$, $h$, and $σ_8$ by the neural network model, while being inefficient in predicting $Ω_b$, { $n_s$}, $M_ν$ and $w$; 4) { compared to the simple fully connected network trained with three statistical quantities, our CNN yields statistically reduced errors, showing improvements of approximately 23\% for $Ω_m$, 11\% for $h$, 8\% for $n_s$, and 21\% for $σ_8$. Additionally, in comparison with the likelihood-based analysis on $P(k)$ data, our CNN provides much tighter constraints on parameters, especially on $Ω_m$ and $σ_8$.} Our study emphasizes this lCNN-based novel approach in extracting large-scale structure information and estimating cosmological parameters.

astro-ph.CO

Correlation-based Beam Calibration of 21cm Intensity Mapping

Foreground removal presents a significant obstacle in both current and forthcoming intensity mapping surveys. While numerous techniques have been developed that show promise in simulated datasets, their efficacy often diminishes when applied to real-world data. A primary issue is the frequency-dependent variations in the instrumental response. In this paper, we propose a novel approach utilizing the internal cross-correlation among different frequencies to calibrate the beam's frequency fluctuations. Using a simulated dataset that incorporates frequency-dependent random fluctuations into the beam model, we illustrate that our method can achieve considerable improvements over traditional techniques. Our results represent a step forward in enhancing the precision and reliability of foreground removal in intensity mapping surveys.

astro-ph.CO