SearcharxivSearch

arXiv subjects

Tong Pan

Publications and source records attributed to Tong Pan.

13 recordsLinked to original sources

Reduced latent leakage does not reliably predict lower likelihood bias in collider inference

Reusable collider representations can be evaluated through downstream discrimination and probes of retained information, but neither quantity directly tests the behaviour of score templates in a profiled likelihood. We test a specific prediction in a controlled two-channel routing protocol: if reduced physics-label readability in a nuisance branch indicates a more inference-robust representation, it should accompany a smaller profiled signal-strength bias under fixed unmodelled shifts. In a public Compact Muon Solenoid $H\rightarrow ZZ\rightarrow4\ell$ workflow, a downstream split of fixed EveNet embeddings preserves signal/background area under the receiver operating characteristic curve ($0.9894\pm0.0004$) while reducing nuisance-branch physics readability from $0.961\pm0.013$ to $0.593\pm0.030$. Probe-sensitivity and effective-rank controls exclude a failed readout and branch collapse. In a separate top quark jet-tagging workflow, the leakage reduction recurs with preserved task performance. Across two development event shards, however, its Spearman association with maximum absolute profiled bias is $0.036$, and three of six material leakage-improving transitions do not reduce that bias. A one-shot preregistered confirmation on an independently accessed shard produces material leakage reductions in all three paired seeds, while the maximum absolute bias increases in two. Thus, within the tested protocol, latent readability is a useful routing diagnostic but not a likelihood-robustness certificate. The result supports a practical validation rule: claims about inference robustness require a prespecified likelihood-facing stress test and held-out confirmation.

physics.data-an

Keyframe-Anchored Identity Preservation for Sequential-Action Video Generation

Identity-preserving text-to-video generation aims to synthesize a video that accurately follows a textual description while maintaining the recognizability of a user-specified subject throughout. The IPVG26 challenge extends this framework from a single holistic prompt to a temporally structured specification. The model additionally receives a sequence of timestamped action captions and must render the subject performing these actions in the specified order. This temporal structure presents a challenge not encountered in previous identity-preserving generation tasks, as the subject must continuously perform a scripted sequence of distinct actions while maintaining a consistent identity. However, end-to-end video generators are prone to appearance drift as motion accumulates and the depicted actions change. We address this challenge with a training-free, three-stage pipeline framework. An action-aware prompt polishment stage first rewrites the inputs into image-generation prompts that specify the terminal state of each action. An identity-preserving generation stage then produces the keyframe sequence by conditioning each frame jointly on the reference identity and its predecessor, thereby decoupling time-invariant appearance from time-varying pose. Finally, an identity-aware inference enhancement stage synthesizes the intermediate segments using multi-reference guidance and identity-driven noise searching, both of which reinforce identity fidelity during sampling. Our method ranked third on the official Track 2 leaderboard, demonstrating competitive performance and strong generality.

cs.CV

Probing the environments of FRI and FRII radio galaxies in LoTSS DR2 with galaxy clusters

The origin of the Fanaroff--Riley Class I/II (FRI/FRII) morphological dichotomy remains uncertain. We investigate whether cluster-scale environment contributes to this distinction using a morphologically classified LoTSS DR2 catalogue at \(z<0.4\). We construct a volume-limited sample with \(L_{144}>4\times10^{24}\,\mathrm{W\,Hz^{-1}}\) and a luminosity--redshift paired sample, and cross-match them with DESI Legacy Imaging Survey galaxy clusters. A radio galaxy is associated with a cluster if \(|\Delta z|<0.01\), projected separation \(<2R_{500}\). In the volume-limited sample, \(48.6\%\) of FRIs and \(30.6\%\) of FRIIs are cluster-associated; in the paired sample, the corresponding fractions are \(45.6\%\) and \(32.6\%\). The difference is stronger at \(L_{144}>10^{26}\,\mathrm{W\,Hz^{-1}}\), where the fractions are \(55.6\%\) versus \(19.0\%\) in the volume-limited sample and \(50.0\%\) versus \(6.7\%\) in the paired sample. However, cluster-associated FRIs and FRIIs occupy similar environments: their radio luminosities and stellar masses show similar trends with cluster richness and \(M_{500}\), and their radial distributions both peak near \(0.5R_{500}\) and decline beyond \(R_{500}\). Most cluster-associated sources are brightest cluster galaxies (BCGs), with fractions of \(74.8\%\) for FRIs and \(61.9\%\) for FRIIs in the volume-limited sample, and \(78.1\%\) and \(65.9\%\) in the paired sample. These results show that FRIIs are less frequently found in clusters, especially at high radio luminosity, consistent with dense intracluster gas disrupting or decelerating jets and suppressing stable FRII structures. Nevertheless, once inside clusters, FRIs and FRIIs inhabit similar large-scale environments, implying that cluster-scale properties alone are unlikely to be the primary driver of the FRI/FRII dichotomy.

astro-ph.GA

The Optical Properties of Host Galaxies of Radio Sources in the Coma Cluster

We present a comprehensive study of host galaxies of radio sources within the 1.35$R_{200}$ of the Coma cluster by combining deep 144MHz observations from the LOFAR Two-Metre Sky Survey (LoTSS-DR2) with optical spectroscopy and photometry from DESI and SDSS. We identify 79 spectroscopically confirmed cluster members with reliable radio emission and classify them into compact, extended, and tailed subsamples according to their radio morphologies. By combining their radio and optical properties, we find compact radio sources are predominantly associated with massive, quiescent galaxies driven by AGN activity, while tailed sources are largely hosted by star-forming galaxies, tracing ongoing ram pressure stripping (RPS). Using phase-space analysis and a projected infall time proxy ($d_R$), we find that extended sources are preferentially located in the cluster outskirts ($d_R > 1$), while tailed sources are concentrated in the intermediate infall region ($0.4 < d_R < 1.0$), highlighting the influence of the dense intracluster medium.

astro-ph.GA

The environments of radio galaxies and quasars in LoTSS data release 2

Aims. The orientation-based unification scheme of radio-loud active galactic nuclei (AGNs) asserts that radio galaxies and quasars are essentially the same type of object, but viewed from different angles. To test this unification model, we compared the environments of radio galaxies and quasars, which would reveal similar properties when an accurate model is utilized. Methods. Using the second data release of the LOFAR Two-metre Sky Survey (LoTSS DR2), we constructed a sample of 26,577 radio galaxies and 2028 quasars at 0.08 < z < 0.4. For radio galaxies with optical spectra, we further classified them as 3631 low-excitation radio galaxies (LERGs) and 1143 high-excitation radio galaxies (HERGs). We crossmatched these samples with two galaxy cluster catalogs from the Sloan Digital Sky Survey (SDSS). Results. We find that $17.1 \pm 0.2%$ of the radio galaxies and $4.1 \pm 0.4%$ of the quasars are associated with galaxy clusters. Luminous quasars are very rare in clusters, while $18.7 \pm 0.7%$ LERGs and $15.2 \pm 1.1%$ HERGs reside in clusters. We also note that in radio galaxies, both HERGs and LERGs tend to reside in the centers of clusters, while quasars do not show a strong preference for their positions in clusters. Conclusions. This study shows that local quasars and radio galaxies exist in different environments, challenging the orientation-based unification model. This means that factors other than orientation may play an important role in distinguishing radio galaxies from quasars. The future WEAVE-LOFAR survey will offer high-quality spectroscopic data for a large number of radio sources and allow for a more comprehensive exploration of the environments of radio galaxies and quasars.

astro-ph.GA

Evaluating the Generalization Ability of Spatiotemporal Model in Urban Scenario

Spatiotemporal neural networks have shown great promise in urban scenarios by effectively capturing temporal and spatial correlations. However, urban environments are constantly evolving, and current model evaluations are often limited to traffic scenarios and use data mainly collected only a few weeks after training period to evaluate model performance. The generalization ability of these models remains largely unexplored. To address this, we propose a Spatiotemporal Out-of-Distribution (ST-OOD) benchmark, which comprises six urban scenario: bike-sharing, 311 services, pedestrian counts, traffic speed, traffic flow, ride-hailing demand, and bike-sharing, each with in-distribution (same year) and out-of-distribution (next years) settings. We extensively evaluate state-of-the-art spatiotemporal models and find that their performance degrades significantly in out-of-distribution settings, with most models performing even worse than a simple Multi-Layer Perceptron (MLP). Our findings suggest that current leading methods tend to over-rely on parameters to overfit training data, which may lead to good performance on in-distribution data but often results in poor generalization. We also investigated whether dropout could mitigate the negative effects of overfitting. Our results showed that a slight dropout rate could significantly improve generalization performance on most datasets, with minimal impact on in-distribution performance. However, balancing in-distribution and out-of-distribution performance remains a challenging problem. We hope that the proposed benchmark will encourage further research on this critical issue.

cs.LG

Robust Traffic Forecasting against Spatial Shift over Years

Recent advancements in Spatiotemporal Graph Neural Networks (ST-GNNs) and Transformers have demonstrated promising potential for traffic forecasting by effectively capturing both temporal and spatial correlations. The generalization ability of spatiotemporal models has received considerable attention in recent scholarly discourse. However, no substantive datasets specifically addressing traffic out-of-distribution (OOD) scenarios have been proposed. Existing ST-OOD methods are either constrained to testing on extant data or necessitate manual modifications to the dataset. Consequently, the generalization capacity of current spatiotemporal models in OOD scenarios remains largely underexplored. In this paper, we investigate state-of-the-art models using newly proposed traffic OOD benchmarks and, surprisingly, find that these models experience a significant decline in performance. Through meticulous analysis, we attribute this decline to the models' inability to adapt to previously unobserved spatial relationships. To address this challenge, we propose a novel Mixture of Experts (MoE) framework, which learns a set of graph generators (i.e., graphons) during training and adaptively combines them to generate new graphs based on novel environmental conditions to handle spatial distribution shifts during testing. We further extend this concept to the Transformer architecture, achieving substantial improvements. Our method is both parsimonious and efficacious, and can be seamlessly integrated into any spatiotemporal model, outperforming current state-of-the-art approaches in addressing spatial dynamics.

cs.LG

STGformer: Efficient Spatiotemporal Graph Transformer for Traffic Forecasting

Traffic forecasting is a cornerstone of smart city management, enabling efficient resource allocation and transportation planning. Deep learning, with its ability to capture complex nonlinear patterns in spatiotemporal (ST) data, has emerged as a powerful tool for traffic forecasting. While graph neural networks (GCNs) and transformer-based models have shown promise, their computational demands often hinder their application to real-world road networks, particularly those with large-scale spatiotemporal interactions. To address these challenges, we propose a novel spatiotemporal graph transformer (STGformer) architecture. STGformer effectively balances the strengths of GCNs and Transformers, enabling efficient modeling of both global and local traffic patterns while maintaining a manageable computational footprint. Unlike traditional approaches that require multiple attention layers, STG attention block captures high-order spatiotemporal interactions in a single layer, significantly reducing computational cost. In particular, STGformer achieves a 100x speedup and a 99.8\% reduction in GPU memory usage compared to STAEformer during batch inference on a California road graph with 8,600 sensors. We evaluate STGformer on the LargeST benchmark and demonstrate its superiority over state-of-the-art Transformer-based methods such as PDFormer and STAEformer, which underline STGformer's potential to revolutionize traffic forecasting by overcoming the computational and memory limitations of existing approaches, making it a promising foundation for future spatiotemporal modeling tasks.

cs.LG

CoPRA: Bridging Cross-domain Pretrained Sequence Models with Complex Structures for Protein-RNA Binding Affinity Prediction

Accurately measuring protein-RNA binding affinity is crucial in many biological processes and drug design. Previous computational methods for protein-RNA binding affinity prediction rely on either sequence or structure features, unable to capture the binding mechanisms comprehensively. The recent emerging pre-trained language models trained on massive unsupervised sequences of protein and RNA have shown strong representation ability for various in-domain downstream tasks, including binding site prediction. However, applying different-domain language models collaboratively for complex-level tasks remains unexplored. In this paper, we propose CoPRA to bridge pre-trained language models from different biological domains via Complex structure for Protein-RNA binding Affinity prediction. We demonstrate for the first time that cross-biological modal language models can collaborate to improve binding affinity prediction. We propose a Co-Former to combine the cross-modal sequence and structure information and a bi-scope pre-training strategy for improving Co-Former's interaction understanding. Meanwhile, we build the largest protein-RNA binding affinity dataset PRA310 for performance evaluation. We also test our model on a public dataset for mutation effect prediction. CoPRA reaches state-of-the-art performance on all the datasets. We provide extensive analyses and verify that CoPRA can (1) accurately predict the protein-RNA binding affinity; (2) understand the binding affinity change caused by mutations; and (3) benefit from scaling data and model size.

q-bio.BM

Easy Begun is Half Done: Spatial-Temporal Graph Modeling with ST-Curriculum Dropout

Spatial-temporal (ST) graph modeling, such as traffic speed forecasting and taxi demand prediction, is an important task in deep learning area. However, for the nodes in graph, their ST patterns can vary greatly in difficulties for modeling, owning to the heterogeneous nature of ST data. We argue that unveiling the nodes to the model in a meaningful order, from easy to complex, can provide performance improvements over traditional training procedure. The idea has its root in Curriculum Learning which suggests in the early stage of training models can be sensitive to noise and difficult samples. In this paper, we propose ST-Curriculum Dropout, a novel and easy-to-implement strategy for spatial-temporal graph modeling. Specifically, we evaluate the learning difficulty of each node in high-level feature space and drop those difficult ones out to ensure the model only needs to handle fundamental ST relations at the beginning, before gradually moving to hard ones. Our strategy can be applied to any canonical deep learning architecture without extra trainable parameters, and extensive experiments on a wide range of datasets are conducted to illustrate that, by controlling the difficulty level of ST relations as the training progresses, the model is able to capture better representation of the data and thus yields better generalization.

cs.LG

Multiplex-detection Based Multiple Instance Learning Network for Whole Slide Image Classification

Multiple instance learning (MIL) is a powerful approach to classify whole slide images (WSIs) for diagnostic pathology. A fundamental challenge of MIL on WSI classification is to discover the \textit{critical instances} that trigger the bag label. However, previous methods are primarily designed under the independent and identical distribution hypothesis (\textit{i.i.d}), ignoring either the correlations between instances or heterogeneity of tumours. In this paper, we propose a novel multiplex-detection-based multiple instance learning (MDMIL) to tackle the issues above. Specifically, MDMIL is constructed by the internal query generation module (IQGM) and the multiplex detection module (MDM) and assisted by the memory-based contrastive loss during training. Firstly, IQGM gives the probability of instances and generates the internal query (IQ) for the subsequent MDM by aggregating highly reliable features after the distribution analysis. Secondly, the multiplex-detection cross-attention (MDCA) and multi-head self-attention (MHSA) in MDM cooperate to generate the final representations for the WSI. In this process, the IQ and trainable variational query (VQ) successfully build up the connections between instances and significantly improve the model's robustness toward heterogeneous tumours. At last, to further enforce constraints in the feature space and stabilize the training process, we adopt a memory-based contrastive loss, which is practicable for WSI classification even with a single sample as input in each iteration. We conduct experiments on three computational pathology datasets, e.g., CAMELYON16, TCGA-NSCLC, and TCGA-RCC datasets. The superior accuracy and AUC demonstrate the superiority of our proposed MDMIL over other state-of-the-art methods.

cs.CV

Catalog of One-side Head-Tail Galaxies in the FIRST Survey

One-side head-tail (OHT) galaxies are radio galaxies with a peculiar shape. They usually appear in galaxy clusters, but they have never been cataloged systematically. We design an automatic procedure to search for them in the Faint Images of the Radio Sky at Twenty-Centimeters source catalog and compile a sample with 115 HT candidates. After cross-checking with the Sloan Digital Sky Survey photometric data and catalogs of galaxy clusters, we find that 69 of them are possible OHT galaxies. Most of them are close to the center of galaxy clusters. The lengths of their tails do not correlate with the projection distance to the center of the nearest galaxy clusters, but show weak anticorrelation with the cluster richness, and are inversely proportional to the radial velocity differences between clusters and host galaxies. Our catalog provides a unique sample to study this special type of radio galaxies.

astro-ph.GA