SearcharxivSearch

arXiv subjects

Payel Das

Publications and source records attributed to Payel Das.

At least 19 recordsLinked to original sources

Can Vision-Language Models Reason about AI Edits in Images?

Detection and localization of AI-tampered images are critical for trustworthy AI, yet modern generative models have made such manipulations increasingly difficult to identify. While traditional binary classifiers can detect image tampering, they lack interpretability and generalization. Vision-Language Models (VLMs) offer a promising alternative due to their strong visual understanding and reasoning capabilities; however, existing approaches typically rely on supervised finetuning with curated explanations rather than exploiting their inherent reasoning capabilities. In this work, we investigate whether VLMs can be trained to reason about AI-generated image edits using reinforcement learning (RL) rather than explicit reasoning supervision. Motivated by the success in Group Relative Policy Optimization (GRPO), an RL technique that incentivizes the model to reason by generating thinking traces prior to giving the final answer, we propose a GRPO-based training framework that utilizes simple accuracy and format rewards. Given an input image, the model produces a structured reasoning trace and predicts whether the image has been tampered with. A lightweight segmentation model is then guided by the reasoning output to generate pixel-level localization masks. Experiments across multiple image manipulation datasets demonstrate that our approach achieves competitive detection and localization performance compared to state-of-the-art image forgery detectors, despite requiring substantially weaker supervision. We introduce effective intersection over union (eff-IoU), a unified metric to jointly evaluate detection and localization. These results suggest that reinforcement learning provides an effective and scalable mechanism for teaching VLMs to reason about AI-generated content.

cs.CV

Disentangling chemical evolution histories with phylogenetic trees

Chemical abundances encode the fossil record of galaxy evolution in a complex and diverse way that requires innovative approaches to reconstruct galactic histories. We investigate the power of using phylogenetic methods to disentangle different evolutionary pathways in analytical chemical evolution models. We ran 1024 one-zone chemical evolution models using flexCE. The resulting chemical abundances are combined with those of two fiducial models, mw-fid and dw-fid, and then used both to determine which combinations produce two-branched phylogenetic trees, as well as how purely these trees split the two input models. We used random forests and Shapley analysis to predict which model combinations return well-separated trees and explain which input parameters are most important for this. We also studied the abundance patterns, as well as star formation rates, mass accumulation, and branch lengths. We found that {\eta}, the mass-loading outflow parameter in flexCE, had the largest impact in separating models into separate branches, due to its importance in driving the chemical enrichment rates and total abundances. Star formation rates and mass accumulation had some impact on {\eta}, but no direct relation between these quantities and the abundances was found. We also found that branches connected through the most metal rich tips in our trees, which is opposite to how phylogenetic trees connect in biological systems. Phylogenetic trees help to reconstruct histories when there is information that is inherited between generations, which is the case of the chemical elements in galaxy evolution. Branch topologies can provide information about the rates of evolutionary change of the various populations, and the connection between branches also contains information about their shared history. This work brings us a step further understanding galaxy evolution through cross-disciplinary research.

astro-ph.GA

Reconstructing chemical enrichment pathways in disc galaxies: A phylogenetic approach

Phylogenetic methods, traditionally used in biology to trace the evolutionary relationships among species, are emerging as a powerful framework to reconstruct evolutionary processes in galaxies from chemical information. We apply galactic phylogenetics to study the chemical evolution of stellar populations in distinct regions of a simulated disc galaxy, assessing its capability to unveil assembly histories. We used a high-resolution simulation that follows the chemical enrichment of an isolated disc galaxy, by different stellar progenitors. We track gas particles as they turn into stars and inherit their parent gas chemical composition. Target particles are selected to store the chemical history of each chemical element considered in the simulation. Two regions were analysed: an inner ring, influenced by early bar-driven inflows, and an outer ring, shaped by spiral arms. We built phylogenetic trees for stellar populations in each region and quantified their structure using the Corrected Colless index, a standard metric of tree balance used in biology. The inner ring tree reveals a compact clade of old stars enriched by rapid SNII feedback, followed by a hierarchical sequence with increasing SNIa and AGB contributions. In contrast, the outer ring exhibits more symmetric, caterpillar-like trees with smoother abundance gradients, consistent with more prolonged star formation and efficient local mixing. Chemical enrichment rates corroborate these trends, showing fast early enrichment in the inner ring and gradual, spatially extended enrichment in the outer disc. The structural indices differ significantly between the two regions and converge robustly even for modest stellar samples (NSSP = 100). Galactic phylogenetics provides a novel and complementary tool to decode the fossil record of galaxies.

astro-ph.GA

A method for constructing the joint mass function of binary stars

The initial mass function (IMF) describes the distribution of stellar masses in a population of newly born stars and is amongst the most fundamental concepts in astrophysics. It is not only the direct result of the star formation process but it also explains the evolution of galaxies' luminosities, metal yields, star-formation efficiencies, and supernova production rates. Because most stars exist in binary systems, however, a full statistical account of stellar mass requires not the IMF but rather the joint distribution of a binary population's primary- and secondary-star masses. This joint distribution must respect the IMF of the stars from which the population has been assembled as well as the distribution of mass ratios that results from the assembly mechanism. Despite its importance, this joint distribution is known only in the case of random pairing. Here we present a method for constructing it in the general case. We also illustrate the use of our method by recovering the known result for random pairing and by finding the previously unknown result for uniform pairing.

astro-ph.SR

Continuous-Time Modelling of Black Hole Binary Evolution with Neural ODEs

Pulsar timing arrays (PTAs) can detect the low-frequency stochastic gravitational-wave background (GWB) generated by an ensemble of supermassive black hole binaries (BHBs). Accurate determination of BHB merger timescales is essential for interpreting GWBs and constraining key astrophysical quantities such as black hole (BH) occupation fractions and galaxy coalescence rates. High-accuracy $N$-body codes such as \texttt{Griffin} can resolve sub-pc BHB dynamics but are too costly to explore a wide range of initial conditions, motivating the need for surrogate models that emulate their long-term evolution at much lower computational cost. We investigate neural ordinary differential equations (NODEs) as surrogates for the secular orbital evolution of BHBs. Our primary contribution is a parameterised NODE (PNODE) trained on an ensemble of $N$-body simulations of galaxy mergers spanning a two-dimensional parameter space defined by the initial orbital eccentricity and particle resolution $(e_i, N)$, with the learned vector field explicitly conditioned on these parameters. A single PNODE thereby learns a simulation-parameter-conditioned dynamical model for the coupled evolution of the BH pair's orbital state across the ensemble, yielding smooth trajectories from which stable hardening and eccentricity growth rates can be extracted. The PNODE accurately reproduces the secular evolution of the specific orbital energy and angular momentum, and the corresponding Keplerian orbital elements, for held-out trajectories, with modest generalisation to a partially unseen high-resolution case. Combining PNODE predictions with semi-analytical prescriptions for stellar hardening and gravitational-wave emission yields BHB merger timescales consistent with those obtained from direct $N$-body inputs within current theoretical uncertainties.

astro-ph.GA

How much can we learn from resolved stellar kinematics of galactic haloes using action-based dynamical models?

Dynamical models are used to study dark matter (DM) in galaxies, how galaxies assemble through mergers, and to test galaxy formation models. Despite its widespread use, there has been no systematic study quantifying how much information can be obtained from just two on-sky positions and line-of-sight velocities, which are typically available for nearby external galaxies. In this work, we introduce axisymmetric, action-based dynamical models that use the positions and velocities of stellar halo stars to jointly constrain the total mass distribution of galaxies and the underlying DM component, as well as the stellar halo phase-space distribution. We rigorously test the method using both idealised equilibrium galaxy mocks and cosmological hydrodynamical simulations from the Auriga suite, systematically assessing how its performance degrades as the available phase-space information is progressively reduced. We further examine the impact of galaxy inclination, modelling assumptions, and methodological systematics on the recovered mass profiles. A crucial development in this work is the improved marginalisation of the model likelihood over missing phase-space dimensions. Our models successfully recover the total and DM mass distributions, as well as the kinematic properties of the stellar tracers, within the derived confidence intervals. However, we find that with limited (3D or 4D) phase-space information, the flattening of the DM halo cannot be constrained with any degree of certainty. Nevertheless, the recovered mass profile is insensitive to the flattening. This finding is independently validated by Schwarzschild modelling tests.

astro-ph.GA

Patching LLM Like Software: A Lightweight Method for Improving Safety Policy in Large Language Models

We propose patching for large language models (LLMs) like software versions, a lightweight and modular approach for addressing safety vulnerabilities. While vendors release improved LLM versions, major releases are costly, infrequent, and difficult to tailor to customer needs, leaving released models with known safety gaps. Unlike full-model fine-tuning or major version updates, our method enables rapid remediation by prepending a compact, learnable prefix to an existing model. This "patch" introduces only 0.003% additional parameters, yet reliably steers model behavior toward that of a safer reference model. Across three critical domains (toxicity mitigation, bias reduction, and harmfulness refusal) policy patches achieve safety improvements comparable to next-generation safety-aligned models while preserving fluency. Our results demonstrate that LLMs can be "patched" much like software, offering vendors and practitioners a practical mechanism for distributing scalable, efficient, and composable safety updates between major model releases.

cs.AI

Stellar velocity distributions in binary-rich ultrafaint dwarf galaxies

Ultrafaint dwarf (UFD) galaxies are dominated by dark matter, the distribution of which may be inferred from the kinematics of that galaxy's stellar population. Star-by-star observations are available for the satellite UFD galaxies of the Milky Way, making them uniquely good laboratories in which to test cosmological predictions at the smallest scales. However, the kinematics of these galaxies are complicated by the presence of binary stars, which alter the stellar velocity distribution. In particular these binary stars increase the galaxy's stellar velocity dispersion, which is related to the total galactic mass by the virial theorem. Without correctly eliminating or accounting for binary stars we may therefore overestimate the masses of UFD galaxies or even confuse globular clusters for UFD galaxies. Here we write down the probability density function for the observed line-of-sight (LOS) velocity of a stellar population containing both visual and spectroscopic binary stars, which we then use to determine the effect of those binary stars on the observed LOS velocity dispersion. For the coldest UFD galaxies the fractional increase in LOS velocity dispersion is of order one and for the coldest globular clusters is of order 100. However, if the stellar initial mass function is bottom light, as it may be for UFD galaxies and globular clusters, then both of these values increase by half a dex.

astro-ph.GA

A North-South Metallicity Asymmetry in the Outer Galactic disk -- Evidence for the Pericentric Passage of the Sagittarius Dwarf Galaxy

We present maps of the mean metallicity distributions on the Galactocentric $R$--$Z$ plane at different azimuthal angles using red clump stars selected from the LAMOST and APOGEE surveys. In the inner disk ($R < $ 11\,kpc), the metallicity distribution is symmetric between the upper and lower disk. However, we find a North-South metallicity asymmetry in the outer disk ($R > 11$\,kpc), especially towards the anti-Galactic center ($-5^\circ < \Phi < 15^\circ$) direction. By further dissecting the map in age space, we detect this asymmetry across all mono-age stellar populations. However, the asymmetry is less pronounced in older populations ($\tau > 8$ Gyr) compared to younger ones ($\tau < 6$\,Gyr). This reduced significance likely stems from three factors: larger age uncertainties, fewer stars in the outer disk, and the kinematically hotter nature of older populations. The observed metallicity asymmetry may be the consequence of the purturbation of the recent pericentric passage through the Galactic disk and tidal force of the well-known Sagittarius dwarf galaxy.

astro-ph.GA

GP-MoLFormer-Sim: Test Time Molecular Optimization through Contextual Similarity Guidance

The ability to design molecules while preserving similarity to a target molecule and/or property is crucial for various applications in drug discovery, chemical design, and biology. We introduce in this paper an efficient training-free method for navigating and sampling from the molecular space with a generative Chemical Language Model (CLM), while using the molecular similarity to the target as a guide. Our method leverages the contextual representations learned from the CLM itself to estimate the molecular similarity, which is then used to adjust the autoregressive sampling strategy of the CLM. At each step of the decoding process, the method tracks the distance of the current generations from the target and updates the logits to encourage the preservation of similarity in generations. We implement the method using a recently proposed $\sim$47M parameter SMILES-based CLM, GP-MoLFormer, and therefore refer to the method as GP-MoLFormer-Sim, which enables a test-time update of the deep generative policy to reflect the contextual similarity to a set of guide molecules. The method is further integrated into a genetic algorithm (GA) and tested on a set of standard molecular optimization benchmarks involving property optimization, molecular rediscovery, and structure-based drug design. Results show that, GP-MoLFormer-Sim, combined with GA (GP-MoLFormer-Sim+GA) outperforms existing training-free baseline methods, when the oracle remains black-box. The findings in this work are a step forward in understanding and guiding the generative mechanisms of CLMs.

cs.LG

Aligning Protein Conformation Ensemble Generation with Physical Feedback

Protein dynamics play a crucial role in protein biological functions and properties, and their traditional study typically relies on time-consuming molecular dynamics (MD) simulations conducted in silico. Recent advances in generative modeling, particularly denoising diffusion models, have enabled efficient accurate protein structure prediction and conformation sampling by learning distributions over crystallographic structures. However, effectively integrating physical supervision into these data-driven approaches remains challenging, as standard energy-based objectives often lead to intractable optimization. In this paper, we introduce Energy-based Alignment (EBA), a method that aligns generative models with feedback from physical models, efficiently calibrating them to appropriately balance conformational states based on their energy differences. Experimental results on the MD ensemble benchmark demonstrate that EBA achieves state-of-the-art performance in generating high-quality protein ensembles. By improving the physical plausibility of generated structures, our approach enhances model predictions and holds promise for applications in structural biology and drug discovery.

q-bio.BM

PEEL the Layers and Find Yourself: Revisiting Inference-time Data Leakage for Residual Neural Networks

This paper explores inference-time data leakage risks of deep neural networks (NNs), where a curious and honest model service provider is interested in retrieving users' private data inputs solely based on the model inference results. Particularly, we revisit residual NNs due to their popularity in computer vision and our hypothesis that residual blocks are a primary cause of data leakage owing to the use of skip connections. By formulating inference-time data leakage as a constrained optimization problem, we propose a novel backward feature inversion method, \textbf{PEEL}, which can effectively recover block-wise input features from the intermediate output of residual NNs. The surprising results in high-quality input data recovery can be explained by the intuition that the output from these residual blocks can be considered as a noisy version of the input and thus the output retains sufficient information for input recovery. We demonstrate the effectiveness of our layer-by-layer feature inversion method on facial image datasets and pre-trained classifiers. Our results show that PEEL outperforms the state-of-the-art recovery methods by an order of magnitude when evaluated by mean squared error (MSE). The code is available at \href{https://github.com/Huzaifa-Arif/PEEL}{https://github.com/Huzaifa-Arif/PEEL}

cs.LG

Fundamental Safety-Capability Trade-offs in Fine-tuning Large Language Models

Fine-tuning Large Language Models (LLMs) on some task-specific datasets has been a primary use of LLMs. However, it has been empirically observed that this approach to enhancing capability inevitably compromises safety, a phenomenon also known as the safety-capability trade-off in LLM fine-tuning. This paper presents a theoretical framework for understanding the interplay between safety and capability in two primary safety-aware LLM fine-tuning strategies, providing new insights into the effects of data similarity, context overlap, and alignment loss landscape. Our theoretical results characterize the fundamental limits of the safety-capability trade-off in LLM fine-tuning, which are also validated by numerical experiments.

stat.ML

Can Memory-Augmented Language Models Generalize on Reasoning-in-a-Haystack Tasks?

Large language models often expose their brittleness in reasoning tasks, especially while executing long chains of reasoning over context. We propose MemReasoner, a new and simple memory-augmented LLM architecture, in which the memory learns the relative order of facts in context, and enables hopping over them, while the decoder selectively attends to the memory. MemReasoner is trained end-to-end, with optional supporting fact supervision of varying degrees. We train MemReasoner, along with existing memory-augmented transformer models and a state-space model, on two distinct synthetic multi-hop reasoning tasks. Experiments performed under a variety of challenging scenarios, including the presence of long distractor text or target answer changes in test set, show strong generalization of MemReasoner on both single- and two-hop tasks. This generalization of MemReasoner is achieved using none-to-weak supporting fact supervision (using none and 1\% of supporting facts for one- and two-hop tasks, respectively). In contrast, baseline models overall struggle to generalize and benefit far less from using full supporting fact supervision. The results highlight the importance of explicit memory mechanisms, combined with additional weak supervision, for improving large language model's context processing ability toward reasoning tasks.

cs.CL

EDGE: The emergence of dwarf galaxy scaling relations from cosmological radiation-hydrodynamics simulations

We present a new suite of EDGE (`Engineering Dwarfs at Galaxy formation's Edge') cosmological zoom simulations. The suite includes 15 radiation-hydrodynamical dwarf galaxies covering the ultra-faint to the dwarf irregular regime ($10^4 \leq M_{\star}(z=0) \leq 10^8 \, M_{\odot}$) to enable comparisons with observed scaling relations. Each object in the suite is evolved at high resolution ($\approx 3 \, \text{pc}$) and includes stellar radiation, winds and supernova feedback channels. We compare with previous \textsc{edge} simulations without radiation, finding that radiative feedback results in significantly weaker galactic outflows. This generalizes our previous findings to a wide mass range, and reveals that the effect is most significant at low $M_{\star}$. Despite this difference, stellar masses stay within a factor of two of each other, and key scaling relations of dwarf galaxies (size-mass, neutral gas-stellar mass, gas-phase mass-metallicity) emerge correctly in both simulation suites. Only the stellar mass -- stellar metallicity relation is strongly sensitive to the change in feedback. This highlights how obtaining statistical samples of dwarf galaxy stellar abundances with next-generation spectrographs will be key to probing and constraining the baryon cycle of dwarf galaxies.

astro-ph.GA

EpMAN: Episodic Memory AttentioN for Generalizing to Longer Contexts

Recent advances in Large Language Models (LLMs) have yielded impressive successes on many language tasks. However, efficient processing of long contexts using LLMs remains a significant challenge. We introduce \textbf{EpMAN} -- a method for processing long contexts in an \textit{episodic memory} module while \textit{holistically attending to} semantically relevant context chunks. The output of \textit{episodic attention} is then used to reweigh the decoder's self-attention to the stored KV cache of the context during training and generation. When an LLM decoder is trained using \textbf{EpMAN}, its performance on multiple challenging single-hop long-context recall and question-answering benchmarks is found to be stronger and more robust across the range from 16k to 256k tokens than baseline decoders trained with self-attention, and popular retrieval-augmented generation frameworks.

cs.CL

Position: Theory of Mind Benchmarks are Broken for Large Language Models

Our paper argues that the majority of theory of mind benchmarks are broken because of their inability to directly test how large language models (LLMs) adapt to new partners. This problem stems from the fact that theory of mind benchmarks for LLMs are overwhelmingly inspired by the methods used to test theory of mind in humans and fall victim to a fallacy of attributing human-like qualities to AI agents. We expect that humans will engage in a consistent reasoning process across various questions about a situation, but this is known to not be the case for current LLMs. Most theory of mind benchmarks only measure what we call literal theory of mind: the ability to predict the behavior of others. However, this type of metric is only informative when agents exhibit self-consistent reasoning. Thus, we introduce the concept of functional theory of mind: the ability to adapt to agents in-context following a rational response to their behavior. We find that many open source LLMs are capable of displaying strong literal theory of mind capabilities, but seem to struggle with functional theory of mind -- even with exceedingly simple partner policies. Simply put, strong literal theory of mind performance does not necessarily imply strong functional theory of mind performance or vice versa. Achieving functional theory of mind, particularly over long interaction horizons with a partner, is a significant challenge deserving a prominent role in any meaningful LLM theory of mind evaluation.

cs.AI

Multi-Scale Representation Learning for Protein Fitness Prediction

Designing novel functional proteins crucially depends on accurately modeling their fitness landscape. Given the limited availability of functional annotations from wet-lab experiments, previous methods have primarily relied on self-supervised models trained on vast, unlabeled protein sequence or structure datasets. While initial protein representation learning studies solely focused on either sequence or structural features, recent hybrid architectures have sought to merge these modalities to harness their respective strengths. However, these sequence-structure models have so far achieved only incremental improvements when compared to the leading sequence-only approaches, highlighting unresolved challenges effectively leveraging these modalities together. Moreover, the function of certain proteins is highly dependent on the granular aspects of their surface topology, which have been overlooked by prior models. To address these limitations, we introduce the Sequence-Structure-Surface Fitness (S3F) model - a novel multimodal representation learning framework that integrates protein features across several scales. Our approach combines sequence representations from a protein language model with Geometric Vector Perceptron networks encoding protein backbone and detailed surface topology. The proposed method achieves state-of-the-art fitness prediction on the ProteinGym benchmark encompassing 217 substitution deep mutational scanning assays, and provides insights into the determinants of protein function. Our code is at https://github.com/DeepGraphLearning/S3F.

cs.LG