SearcharxivSearch

arXiv subjects

Kunyang Li

Publications and source records attributed to Kunyang Li.

At least 19 recordsLinked to original sources

Set them free: extending RAMCOAL to model massive black hole triplets in hydrodynamical simulations of galaxies

Massive black hole binaries (MBHBs), and the higher-order multiples produced by repeated galaxy mergers, spend part of their lives in dynamical regimes that cosmological simulations cannot resolve, even though these regimes set their merger delays, spins, recoils, and host-galaxy context. We extend the RAMCOAL framework to follow such subgrid massive black hole triplets directly within hydrodynamical galaxy simulations. As in the original staged binary model, the black holes start as sink particles, pass through a dynamical-friction phase, and settle into bound binaries that harden through stellar scattering, gas torques, circumbinary-disc coupling, and gravitational-wave emission. When a hierarchical triplet becomes chaotic, RAMCOAL maps the encounter onto a library of three-body outcomes from direct N-body experiments and updates the surviving system, following the resulting mergers, exchanges, and ejections together with the accretion and spin evolution of each black hole. Using isolated-galaxy tests with contrasting geometries, we show that the encounter geometry alone can change which pair finally merges, and after how long. We demonstrate the first triplet MBH dynamical evolution all the way to coalescence inside a live hydrodynamical simulation. This establishes an end-to-end capability to predict triplet-driven MBH coalescences self-consistently coupled to the evolving host galaxy. Because each MBHB coalescence carries its environmental history through the subgrid phase, RAMCOAL offers a route toward merger catalogues that link the gravitational-wave signatures of coalescing black holes to the galaxies in which they form.

astro-ph.GA

MESA: Prioritizing Vulnerable Communication Channels for Securing Multi-Agent Systems

Multi-agent systems (MAS) are increasingly used to automate complex, distributed workflows. However, their inter-agent communication channels introduce new attack surfaces that remain poorly understood and are difficult to defend against. In this paper, we address how defenders should prioritize limited security effort to protect vulnerable communication channels before attacks are observed. This is motivated by our observation that the channel-level attack impact is highly non-uniform: a single compromised edge can account for up to 75% of total attack success. We introduce Mesa, a label-free framework for proactively ranking which MAS edges are most security-critical -- that is, most likely to affect the system's decision if compromised. Mesa combines six graph-theoretic metrics and two dynamic probes (ablation and masking) without requiring attack traces. We evaluate Mesa against a dynamic misinformation attack pipeline across three diverse MAS scenarios, eight network topologies, and five open-source LLMs from Qwen, Llama, and Gemma families. Mesa rankings correlate strongly with empirical per-edge attack success rate, achieving mean Spearman $\rho=+0.60$ (peaking at $+0.73$). In resource-constrained defense deployment, monitoring the top 10% of Mesa-ranked edges intercepts about 3x the successful attacks as random allocation. We further test Mesa under varying attacker and defender models and LangGraph workflows and characterize its limits under adaptive attacks and high-redundancy graphs. Overall, our results show that edge-level risk in MAS is often concentrated and predictable, allowing proactive hardening of multi-agent infrastructures.

cs.CR

Aero-World: Action-Conditioned Aerial Video Generation from Inertial Controls

Foundation video models produce visually impressive results, but their use in embodied AI remains limited because they are primarily trained on natural language rather than low-level control signals. This limitation is especially pronounced for aerial flight, where motion occurs in unconstrained 6-DoF space and small errors in ego-motion can produce large trajectory drift. Generating aerial videos that follow fine-grained inertial actions can support scalable training and evaluation of aerial agents by providing a controllable proxy for real-world or expensive simulation data. To address this problem, we propose \textbf{Aero-World}, a method for converting a pretrained image-to-video diffusion model into a controllable aerial video generator. Aero-World injects sequences of translational acceleration and angular velocity into a pretrained latent diffusion transformer through an action-token stream. A frozen latent-space Physics Probe, trained independently on real video--IMU pairs, provides differentiable inertial-consistency supervision during LoRA finetuning while avoiding computationally expensive video decoding. We further propose \textbf{AeroBench}, a benchmark for evaluating whether generated drone videos adhere to low-level action signals. AeroBench uses Action Alignment Score (AAS) to measure agreement with commanded inertial actions and Physical Consistency Rate (PCR) to measure temporal motion stability. On AeroBench, Aero-World improves mean AAS from 57.7 to 63.6 over action-only finetuning and gives a stronger quality-control trade-off than AirScape, with lower FVD (596.5 vs. 1058.6), higher SSIM (0.595 vs. 0.505), and higher Flow-IMU correlation (0.44 vs. 0.20). These results suggest that frozen Physics Probe supervision is a practical mechanism for adapting pretrained video generators toward more action-aligned aerial motion.

cs.CV

Attend Locally, Remember Linearly: Linear Attention as Cross-Frame Memory for Autoregressive Video Diffusion

Autoregressive (AR) video diffusion is a powerful paradigm for streaming and interactive video generation. However, its reliance on softmax self-attention leads to quadratic compute complexity in sequence length and memory usage due to key-value caching, which limits its scalability to long video horizons. Existing remedies (e.g., sparse attention and KV-cache compression) reduce per-step cost but still rely on a linearly growing cache or irreversibly discard past context, and thus fail to address linear memory growth and streaming context management. To address this scalability bottleneck, we propose ARL2 (Attend Locally, Remember Linearly), a hybrid attention module that replaces quadratic cross-frame attention with a fixed-size recurrent state. We decompose self-attention into two branches: an intra-frame softmax branch for spatial detail and local dependencies, and an inter-frame gated recurrent linear branch that maintains a fixed-size state for streaming context. Our key insight is that softmax attention captures fine-grained local interactions, while a recurrent state provides controllable long-range memory. This design achieves linear-time scaling with constant memory while improving temporal consistency over the full-softmax model. To prevent noisy intermediate states from corrupting memory, we update the recurrent state only after the denoised pass. To avoid within-frame information asymmetry, all tokens share the same pre-update state rather than sequential updates. To the best of our knowledge, this is the first work to convert a pretrained AR video diffusion model into a hybrid linear attention architecture, through an efficient two-stage training scheme for AR video. With 75% of layers replaced by hybrid linear attention, the model achieves up to 2.26 wall-clock speedup and 54% memory reduction, while maintaining comparable quality with improving temporal consistency.

cs.CV

The LISA Astrophysics MBHcatalogues Project: A comparison of predictions of simulated massive black hole binaries

In the hierarchical paradigm of galaxy formation, central massive black holes (MBHs) are expected to coalesce after the merger of their host galaxies. One of the main goals of the Laser Interferometer Space Antenna (LISA) is to constrain the origin and growth of MBHs through their merger rates and mass distribution. Predicting MBH merger rates requires not only tracing their statistical population from large to small physical scales (kpc to sub-pc) but also modelling their formation, accretion, dynamics, mergers, and their galactic physical processes across cosmic time. This project is the result of a large collaborative effort undertaken by the LISA Astrophysics Working Group, bringing together its collective expertise on MBH formation, evolution, and modelling, to build a comprehensive understanding of MBH merger rates across cosmic time. The project compares various theoretical predictions of MBH merger rates, quantifies the spread, and evaluates the global astrophysical uncertainties of the LISA event rates. To build a unique and complete view, our work is based on about 20 semi-analytical models and cosmological simulations from the literature, all employing distinct approaches to modelling MBH and galaxy physics. To compute the merger rates, we also incorporate delays arising from the dynamical phase of MBH hardening to coalescence. We present the expected LISA merger rates given current galaxy formation models and discuss how the merger rate depends on model assumptions, such as the seeding model and the resolution of cosmological simulations.

astro-ph.GA

A Non-compact Positivity-Preserving Numerical Scheme for Elliptic Differential Equations Based on Mathematical Expectation

We propose a novel non-compact, positivity-preserving scheme for linear non-divergence form elliptic equations. Based on the Feynman--Kac formula, the solution is represented as a conditional expectation associated with a diffusion process.Instead of using compact Markov chain approximations, we construct a wide-stencil scheme by approximating the expectation with carefully designed transition probabilities, ensuring both consistency and positivity preservation. The method is effective for anisotropic diffusion problems with mixed derivatives, where classical schemes typically fail unless the covariance matrix is diagonally dominant. A key feature of the proposed framework is its robust treatment of boundary conditions. For Dirichlet boundaries, we introduce a quadtree-based non-uniform stopping-time strategy, achieving $O(h)$ accuracy. For Neumann boundaries, a discrete specular reflection mechanism is employed, yielding $O(h^{1/2})$ convergence. Periodic boundaries are handled through modular wrapping, also achieving $O(h)$ accuracy. The resulting schemes are unconditionally stable and positivity-preserving due to their probabilistic structure. Numerical experiments confirm the theoretical convergence rates under all boundary conditions considered.

math.NA

PackCache: A Training-Free Acceleration Method for Unified Autoregressive Video Generation via Compact KV-Cache

A unified autoregressive model is a Transformer-based framework that addresses diverse multimodal tasks (e.g., text, image, video) as a single sequence modeling problem under a shared token space. Such models rely on the KV-cache mechanism to reduce attention computation from O(T^2) to O(T); however, KV-cache size grows linearly with the number of generated tokens, and it rapidly becomes the dominant bottleneck limiting inference efficiency and generative length. Unified autoregressive video generation inherits this limitation. Our analysis reveals that KV-cache tokens exhibit distinct spatiotemporal properties: (i) text and conditioning-image tokens act as persistent semantic anchors that consistently receive high attention, and (ii) attention to previous frames naturally decays with temporal distance. Leveraging these observations, we introduce PackCache, a training-free KV-cache management method that dynamically compacts the KV cache through three coordinated mechanisms: condition anchoring that preserves semantic references, cross-frame decay modeling that allocates cache budget according to temporal distance, and spatially preserving position embedding that maintains coherent 3D structure under cache removal. In terms of efficiency, PackCache accelerates end-to-end generation by 1.7-2.2x on 48-frame long sequences, showcasing its strong potential for enabling longer-sequence video generation. Notably, the final four frames - the portion most impacted by the progressively expanding KV-cache and thus the most expensive segment of the clip - PackCache delivers a 2.6x and 3.7x acceleration on A40 and H200, respectively, for 48-frame videos.

cs.CV

The multimessenger view of Pulsar Timing Array black holes with the Horizon-AGN simulation

We use the Horizon-AGN cosmological simulation to study the properties of supermassive black hole binaries (MBHBs) contributing most to the gravitational wave background (GWB) signal expected in the pulsar timing array (PTA) band. We develop a pipeline to generate realistic populations of MBHBs, allowing us to estimate both the characteristic strain and GWB time series observable by PTA experiments. We identify potential continuous wave (CW) candidates standing above the background noise, using toy PTA sensitivities representing the current EPTA and future SKA. We estimate the probability of detecting at least one CW with signal-to-noise ratio $>3$ to be $4\%$ ($20\%$) for EPTA (SKA)-like sensitivities, assuming a 10-year baseline. We find the GWB to be dominated by hundreds to thousands of binaries at redshifts in the range $0.05-1$, with chirp masses of $10^{8.5}-10^{9.5}\, M_\odot$, hosted mainly in quiescent massive galaxies residing in halos of mass $\sim 10^{13}\, M_\odot$. CW candidates have larger masses, lower redshifts and are found in even more massive halos, typical of galaxy groups and clusters. The majority of these systems would appear as AGN rather than quasars, because of their low Eddington ratios. Nevertheless, CW candidates with $f_{\rm Edd}>10^{-3}$ can still outshine their hosts, particularly in radio and X-ray bands, suggesting them as the most promising route for identification. Our findings imply that optical and near-infrared searches based on light curve variability are challenging and biased toward more luminous systems. Finally, we highlight important caveats in the common method used to compare PTA observations with theoretical models. We find that GWB spectral inferences used by PTAs could be biased toward shallower slopes and higher amplitudes at $f=1/\rm yr$, thereby reducing the apparent tension between astrophysical expectations and PTA observations.

astro-ph.GA

GVD: Guiding Video Diffusion Model for Scalable Video Distillation

To address the larger computation and storage requirements associated with large video datasets, video dataset distillation aims to capture spatial and temporal information in a significantly smaller dataset, such that training on the distilled data has comparable performance to training on all of the data. We propose GVD: Guiding Video Diffusion, the first diffusion-based video distillation method. GVD jointly distills spatial and temporal features, ensuring high-fidelity video generation across diverse actions while capturing essential motion information. Our method's diverse yet representative distillations significantly outperform previous state-of-the-art approaches on the MiniUCF and HMDB51 datasets across 5, 10, and 20 Instances Per Class (IPC). Specifically, our method achieves 78.29 percent of the original dataset's performance using only 1.98 percent of the total number of frames in MiniUCF. Additionally, it reaches 73.83 percent of the performance with just 3.30 percent of the frames in HMDB51. Experimental results across benchmark video datasets demonstrate that GVD not only achieves state-of-the-art performance but can also generate higher resolution videos and higher IPC without significantly increasing computational cost.

cs.CV

MP-GCAN: a highly accurate classifier for $\alpha$-helical membrane proteins and $\beta$-barrel proteins

Membrane protein classification is a fundamental task in structural bioinformatics, critical to understanding protein functions and accelerating drug discovery. In this study, we propose MP-GCAN, a novel graph-based classification model that leverages both spatial and sequential features of proteins. MP-GCAN combines GCN, GAT, and GIN layers to capture hierarchical structural representations from 3D protein graphs, constructed from high-resolution PDB files with $\alpha$-carbon coordinates and residue types. To evaluate performance, we curated a high-quality dataset of 500 membrane and 500 non-membrane proteins, and compared MP-GCAN with two baselines: a structure-confidence-based SGD classifier utilizing AlphaFold's pLDDT scores, and DeepTMHMM, a sequence-based deep learning model. Our experiments demonstrate that MP-GCAN significantly outperforms baselines, achieving an accuracy of 96% and strong F1-scores on both classes. The results highlight the importance of integrating pretrained GNN architectures with domain-specific structural data to enhance membrane protein classification.

q-bio.QM

Robust Lane Detection with Wavelet-Enhanced Context Modeling and Adaptive Sampling

Lane detection is critical for autonomous driving and ad-vanced driver assistance systems (ADAS). While recent methods like CLRNet achieve strong performance, they struggle under adverse con-ditions such as extreme weather, illumination changes, occlusions, and complex curves. We propose a Wavelet-Enhanced Feature Pyramid Net-work (WE-FPN) to address these challenges. A wavelet-based non-local block is integrated before the feature pyramid to improve global context modeling, especially for occluded and curved lanes. Additionally, we de-sign an adaptive preprocessing module to enhance lane visibility under poor lighting. An attention-guided sampling strategy further reffnes spa-tial features, boosting accuracy on distant and curved lanes. Experiments on CULane and TuSimple demonstrate that our approach signiffcantly outperforms baselines in challenging scenarios, achieving better robust-ness and accuracy in real-world driving conditions.

cs.CV

On the Robustness Tradeoff in Fine-Tuning

Fine-tuning has become the standard practice for adapting pre-trained models to downstream tasks. However, the impact on model robustness is not well understood. In this work, we characterize the robustness-accuracy trade-off in fine-tuning. We evaluate the robustness and accuracy of fine-tuned models over 6 benchmark datasets and 7 different fine-tuning strategies. We observe a consistent trade-off between adversarial robustness and accuracy. Peripheral updates such as BitFit are more effective for simple tasks -- over 75% above the average measured by the area under the Pareto frontiers on CIFAR-10 and CIFAR-100. In contrast, fine-tuning information-heavy layers, such as attention layers via Compacter, achieves a better Pareto frontier on more complex tasks -- 57.5% and 34.6% above the average on Caltech-256 and CUB-200, respectively. Lastly, we observe that the robustness of fine-tuning against out-of-distribution data closely tracks accuracy. These insights emphasize the need for robustness-aware fine-tuning to ensure reliable real-world deployments.

cs.LG

Alignment and Adversarial Robustness: Are More Human-Like Models More Secure?

A small but growing body of work has shown that machine learning models which better align with human vision have also exhibited higher robustness to adversarial examples, raising the question: can human-like perception make models more secure? If true generally, such mechanisms would offer new avenues toward robustness. In this work, we conduct a large-scale empirical analysis to systematically investigate the relationship between representational alignment and adversarial robustness. We evaluate 114 models spanning diverse architectures and training paradigms, measuring their neural and behavioral alignment and engineering task performance across 105 benchmarks as well as their adversarial robustness via AutoAttack. Our findings reveal that while average alignment and robustness exhibit a weak overall correlation, specific alignment benchmarks serve as strong predictors of adversarial robustness, particularly those that measure selectivity toward texture or shape. These results suggest that different forms of alignment play distinct roles in model robustness, motivating further investigation into how alignment-driven approaches can be leveraged to build more secure and perceptually-grounded vision models.

cs.CV

All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages

Existing Large Multimodal Models (LMMs) generally focus on only a few regions and languages. As LMMs continue to improve, it is increasingly important to ensure they understand cultural contexts, respect local sensitivities, and support low-resource languages, all while effectively integrating corresponding visual cues. In pursuit of culturally diverse global multimodal models, our proposed All Languages Matter Benchmark (ALM-bench) represents the largest and most comprehensive effort to date for evaluating LMMs across 100 languages. ALM-bench challenges existing models by testing their ability to understand and reason about culturally diverse images paired with text in various languages, including many low-resource languages traditionally underrepresented in LMM research. The benchmark offers a robust and nuanced evaluation framework featuring various question formats, including true/false, multiple choice, and open-ended questions, which are further divided into short and long-answer categories. ALM-bench design ensures a comprehensive assessment of a model's ability to handle varied levels of difficulty in visual and linguistic reasoning. To capture the rich tapestry of global cultures, ALM-bench carefully curates content from 13 distinct cultural aspects, ranging from traditions and rituals to famous personalities and celebrations. Through this, ALM-bench not only provides a rigorous testing ground for state-of-the-art open and closed-source LMMs but also highlights the importance of cultural and linguistic inclusivity, encouraging the development of models that can serve diverse global populations effectively. Our benchmark is publicly available.

cs.CV

Tracking on-the-fly massive black hole binary evolution and coalescence in galaxy simulations: RAMCOAL

The detection of gravitational waves (GWs) from massive black hole binary (MBHB) coalescence motivates the development of a sub-grid model. We present RAMCOAL, integrated into the RAMSES code, which simulates the orbital evolution of MBHBs, accounting for stellar and gaseous dynamical friction (DF), stellar scattering, circumbinary disk interactions, and GW emission at scales below the simulation resolution. Unlike post-processing approaches, RAMCOAL tracks the real-time evolution of MBHBs within hydrodynamical simulations of galaxies using local quantities to model dynamics and accretion. This enables more accurate predictions of both GW signals and the properties of merging black holes. We validate RAMCOAL across isolated and merging galaxy setups at resolutions of 10, 50, and 100 pc, with and without black hole accretion and feedback. In addition, we test the model in seven galaxy merger scenarios at 100 pc resolution. These tests demonstrate that RAMCOAL is largely resolution-independent and successfully captures the effects of DF from stars, dark matter, and gas, loss-cone scattering, viscous drag from circumbinary disks, and GW emission -- all within a realistic galactic environment, even at low resolutions. With RAMCOAL, we can better estimate MBHB coalescence rates and the GW background, while providing insights into the electromagnetic counterparts of GW sources. This approach bridges the gap between electromagnetic observations and GW detection, offering a more comprehensive understanding of MBHB evolution in cosmological simulations.

astro-ph.GA

ParTEETor: A System for Partial Deployments of TEEs within Tor

The Tor anonymity network allows users such as political activists and those under repressive governments to protect their privacy when communicating over the internet. At the same time, Tor has been demonstrated to be vulnerable to several classes of deanonymizing attacks that expose user behavior and identities. Prior work has shown that these threats can be mitigated by leveraging trusted execution environments (TEEs). However, previous proposals assume that all relays in the network will be TEE-based-which as a practical matter is unrealistic. In this work, we introduce ParTEETor, a Tor-variant system, which leverages partial deployments of TEEs to thwart known attacks. We study two modes of operation: non-policy and policy. Non-policy mode uses the existing Tor relay selection algorithm to provide users incident security. Policy mode extends the relay selection algorithm to address the classes of attacks by enforcing a specific TEE circuit configuration. We evaluate ParTEETor for security, performance, and privacy. Our evaluation demonstrates that at even a small TEE penetration (e.g., 10% of relays are TEE-based), users can reach performance of Tor today while enforcing a security policy to guarantee protection from at least two classes of attacks. Overall, we find that partial deployments of TEEs can substantially improve the security of Tor, without a significant impact on performance or privacy.

cs.CR

The Efficacy of Transformer-based Adversarial Attacks in Security Domains

Today, the security of many domains rely on the use of Machine Learning to detect threats, identify vulnerabilities, and safeguard systems from attacks. Recently, transformer architectures have improved the state-of-the-art performance on a wide range of tasks such as malware detection and network intrusion detection. But, before abandoning current approaches to transformers, it is crucial to understand their properties and implications on cybersecurity applications. In this paper, we evaluate the robustness of transformers to adversarial samples for system defenders (i.e., resiliency to adversarial perturbations generated on different types of architectures) and their adversarial strength for system attackers (i.e., transferability of adversarial samples generated by transformers to other target models). To that effect, we first fine-tune a set of pre-trained transformer, Convolutional Neural Network (CNN), and hybrid (an ensemble of transformer and CNN) models to solve different downstream image-based tasks. Then, we use an attack algorithm to craft 19,367 adversarial examples on each model for each task. The transferability of these adversarial examples is measured by evaluating each set on other models to determine which models offer more adversarial strength, and consequently, more robustness against these attacks. We find that the adversarial examples crafted on transformers offer the highest transferability rate (i.e., 25.7% higher than the average) onto other models. Similarly, adversarial examples crafted on other models have the lowest rate of transferability (i.e., 56.7% lower than the average) onto transformers. Our work emphasizes the importance of studying transformer architectures for attacking and defending models in security domains, and suggests using them as the primary architecture in transfer attack settings.

cs.CR

The Trade-off between Universality and Label Efficiency of Representations from Contrastive Learning

Pre-training representations (a.k.a. foundation models) has recently become a prevalent learning paradigm, where one first pre-trains a representation using large-scale unlabeled data, and then learns simple predictors on top of the representation using small labeled data from the downstream tasks. There are two key desiderata for the representation: label efficiency (the ability to learn an accurate classifier on top of the representation with a small amount of labeled data) and universality (usefulness across a wide range of downstream tasks). In this paper, we focus on one of the most popular instantiations of this paradigm: contrastive learning with linear probing, i.e., learning a linear predictor on the representation pre-trained by contrastive learning. We show that there exists a trade-off between the two desiderata so that one may not be able to achieve both simultaneously. Specifically, we provide analysis using a theoretical data model and show that, while more diverse pre-training data result in more diverse features for different tasks (improving universality), it puts less emphasis on task-specific features, giving rise to larger sample complexity for down-stream supervised tasks, and thus worse prediction performance. Guided by this analysis, we propose a contrastive regularization method to improve the trade-off. We validate our analysis and method empirically with systematic experiments using real-world datasets and foundation models.

cs.LG