SearcharxivSearch

arXiv subjects

Shweta Jain

Publications and source records attributed to Shweta Jain.

At least 19 recordsLinked to original sources

The Mass Dependence of the Fundamental Metallicity Relation in Observations and Simulations

The metal content of galaxies provides direct insight into the underlying physical processes that drive galaxy evolution. An example of this is the three-parameter relationship between stellar mass, gas-phase metallicity, and star formation rate, commonly referred to as the Fundamental Metallicity Relation (FMR). Previous studies have suggested that the FMR is redshift-invariant (at $z \lesssim 4$) and fully accounts for the scatter in the mass-metallicity relation (MZR). In this work, we test this 'fundamental' relation in both cosmological simulations (EAGLE, SIMBA, Illustris, IllustrisTNG) and Sloan Digital Sky Survey (SDSS) observations. We find that the canonical anti-correlation between metallicity and specific star formation rate (sSFR) inverts in massive galaxies ($M_\star \gtrsim 10^{10.5} \mathrm{M}_\odot$) in EAGLE, IllustrisTNG, and SDSS. When including lower star forming galaxies, the positive correlation appears for all four simulations and SDSS. We speculate that this inversion may being driven by strong nuclear outflows (from, e.g., active galactic nuclei or stellar feedback), which quench star formation while simultaneously expelling preferentially enriched gas from the center of the galaxy. We also find that this 'inversion' appears in a number of metallicity diagnostics in observations (though the details depend on diagnostic) and persists out to $z \sim 1$ in the simulations. These results demonstrate that these strong nuclear outflows challenge simple gas regulator-type models and provide a new framework to test models of the baryon cycle in both future simulations and observations.

astro-ph.GA

PLACO: A Multi-Stage Framework for Cost-Effective Performance in Human-AI Teams

Human-AI teams play a pivotal role in improving overall system performance when neither the human nor the model can achieve such performance on their own. With the advent of powerful and accessible Generative AI models, several mundane tasks have morphed into Human-AI team tasks. From writing essays to developing advanced algorithms, humans have found that using AI assistance has led to an accelerated work pace like never before. In classification tasks, where the final output is a single hard label, it is crucial to address the combination of human and model output. Prior work elegantly solves this problem using Bayes rule, using the assumption that human and model output are conditionally independent given the ground truth. Specifically, it discusses a combination method to combine a single deterministic labeler (the human) and a probabilistic labeler (the classifier model) using the model's instance-level and the human's class-level calibrated probabilities.

cs.AI

Meritocratic Fairness via $K$-Shapley Values in Budgeted Combinatorial Bandits with Full-Bandit Feedback

We study meritocratic fairness in budgeted combinatorial multi-armed bandits with full-bandit feedback, where a learner selects at most $K$ arms per time step and observes only the noisy aggregate reward of the selected set. To define merit under budgeted coalition constraints, we introduce the $K$-Shapley value, an adaptation of the classical Shapley value that measures marginal contributions using only coalitions of size at most $K$. We show that the $K$-Shapley value is the unique solution concept satisfying symmetry, linearity, null player, and $K$-efficiency axioms. We then establish an $\Omega(T^{2/3})$ lower bound on fairness regret for monotone submodular valuation functions. We show that an explore-then-commit algorithm MURaS (Meritocratic Uniform Random Sampling) achieves $\tilde O(T^{2/3})$ fairness regret by exploring all arms uniformly in exploration phase. To improve empirical regret, we propose IW-KSVFair, a meritocratic full-bandit algorithm that learns a selection policy whose arm marginals are proportional to the unknown $K$-Shapley values. To correct the bias induced by adaptive sampling, IW-KSVFair uses importance-weighted estimation and mixes the adaptive set distribution with a uniform distribution to keep importance weights bounded. We prove that IW-KSVFair achieves $\tilde O(T^{2/3})$ fairness regret, matching the lower bound up to logarithmic factors. Experiments on synthetic and real-world datasets show that IW-KSVFair achieves low cumulative fairness regret and closely aligns empirical selection frequencies with $K$-Shapley value-based merit.

cs.LG

Lipschitz Dueling Bandits over Continuous Action Spaces

We study for the first time, stochastic dueling bandits over continuous action spaces with Lipschitz structure, where feedback is purely comparative. While dueling bandits and Lipschitz bandits have been studied separately, their combination has remained unexplored. We propose the first algorithm for Lipschitz dueling bandits, using round-based exploration and recursive region elimination guided by an adaptive reference arm. We develop new analytical tools for relative feedback and prove a regret bound of $\tilde O\left(T^{\frac{d_z+1}{d_z+2}}\right)$, where $d_z$ is the zooming dimension of the near-optimal region. Further, our algorithm takes only logarithmic space in terms of the total time horizon, best achievable by any bandit algorithm over a continuous action space.

cs.LG

Multi-Agent Combinatorial-Multi-Armed-Bandit framework for the Submodular Welfare Problem under Bandit Feedback

We study the \emph{Submodular Welfare Problem} (SWP), where items are partitioned among agents with monotone submodular utilities to maximize the total welfare under \emph{bandit feedback}. Classical SWP assumes full value-oracle access, achieving $(1-1/e)$ approximations via continuous-greedy algorithms. We extend this to a \emph{multi-agent combinatorial bandit} framework (\textsc{MA-CMAB}), where actions are partitions under full-bandit feedback with non-communicating agents. Unlike prior single-agent or separable multi-agent CMAB models, our setting couples agents through shared allocation constraints. We propose an explore-then-commit strategy with randomized assignments, achieving $\tilde{\mathcal{O}}(T^{2/3})$ regret against a $(1-1/e)$ benchmark, the first such guarantee for partition-based submodular welfare problem under bandit feedback.

cs.GT

Aggregating maximal cliques in real-world graphs

Maximal clique enumeration is a fundamental graph mining task, but its utility is often limited by computational intractability and highly redundant output. To address these challenges, we introduce \emph{$\rho$-dense aggregators}, a novel approach that succinctly captures maximal clique structure. Instead of listing all cliques, we identify a small collection of clusters with edge density at least $\rho$ that collectively contain every maximal clique. In contrast to maximal clique enumeration, we prove that for all $\rho < 1$, every graph admits a $\rho$-dense aggregator of \emph{sub-exponential} size, $n^{O(\log_{1/\rho}n)}$, and provide an algorithm achieving this bound. For graphs with bounded degeneracy, a typical characteristic of real-world networks, our algorithm runs in near-linear time and produces near-linear size aggregators. We also establish a matching lower bound on aggregator size, proving our results are essentially tight. In an empirical evaluation on real-world networks, we demonstrate significant practical benefits for the use of aggregators: our algorithm is consistently faster than the state-of-the-art clique enumeration algorithm, with median speedups over $6\times$ for $\rho=0.1$ (and over $300\times$ in an extreme case), while delivering a much more concise structural summary.

cs.DS

SDSS-C4 3028: The Nearest Blue Galaxy Cluster Devoid of an Intracluster Medium

SDSS-C4 3028 is a galaxy cluster at $z=0.061$, notable for its unusually high fraction of star-forming galaxies with 19 star-forming and 11 quiescent spectroscopically-confirmed member galaxies. From Subaru/HSC imaging, we derived a weak lensing mass of $M_{200} = (1.3 \pm 0.9) \times 10^{14} \rm M_\odot$, indicating a low-mass cluster. This is in excellent agreement with its dynamical mass of $M_{200} = (1.0\pm0.4)\times10^{14} \rm M_\odot$, derived from SDSS spectroscopic data. XMM-Newton observations reveal that its X-ray emission is uniform and fully consistent with the astrophysical X-ray background, with no evidence for an intracluster medium (ICM). The 3$\sigma$ upper limit of $L_{\rm X}(0.1-2.4\rm keV)=7.7\times10^{42}$ erg s$^{-1}$ on the cluster's X-ray luminosity falls below the value expected from the $L_{\rm X}-M_{\rm halo}$ scaling relation of nearby galaxy clusters. We derived star-formation histories for its member galaxies using the photometric spectral energy distribution from SDSS, 2MASS, and WISE data. Most of its quiescent galaxies reside within the central 300 kpc, while star-forming ones dominate the outer region (300 kpc - 1 Mpc). The core region has formed the bulk of its stellar mass approximately 1.5 Gyr earlier than the outskirts. We infer a long quenching time of $>3$ Gyr for its quiescent galaxies, consistent with slow quenching mechanisms such as galaxy-galaxy interaction or strangulation. These findings suggest that SDSS-C4 3028 may have undergone an "inside-out" formation and quenching process. Its ICM may have been expelled by intense AGN feedback after core formation but before full cluster assembly. The high fraction ($\sim$0.63) of star-forming members likely results from the absence of ram pressure stripping in this blue cluster, supporting the important role of ram pressure stripping in quenching galaxies in clusters.

astro-ph.GA

A Uniform Analysis of Gas-phase Metallicity Evolution with 1-3 Gyr Time Sampling over the Past 12 Billion Years

We present a systematic investigation of the evolution of the mass-metallicity relation (MZR) and fundamental metallicity relation (FMR) using uniform metallicity diagnostics across redshifts $z\sim0$ to $z\sim3.3$. We present new Keck/DEIMOS measurements of the [OII]$\lambda\lambda3726,3729$ emission line doublet for star-forming galaxies at $z\sim1.5$ with existing measurements of redder rest-optical lines from the MOSDEF survey. These new observations enable uniform estimation of the gas-phase oxygen abundance using ratios of the [OII], H$\beta$, and [OIII] lines for mass-binned samples of star-forming galaxies in 6 redshift bins, employing strong-line calibrations that account for the distinct interstellar medium ionization conditions at $z<1$ and $z>1$. We find that the low-mass power law slope of the MZR remains constant over this redshift range with a value of $\gamma=0.28\pm0.01$, implying the outflow metal loading factor ($\zeta_\text{out}=\frac{Z_{\text{out}}}{Z_{\text{ISM}}}\frac{\dot{M}_{\text{out}}}{\text{SFR}}$) scales approximately as $\rm \zeta_{out}\propto M_*^{-0.3}$ out to at least $z\sim3.3$. The normalization of the MZR at $10^{10}\ \text{M}_\odot$ decreases with increasing redshift at a rate of $d\log(\text{O/H})/dz =-0.11\pm0.01$ across the full redshift range. We find that any evolution of the FMR is smaller than 0.1 dex out to $z\sim3.3$. We compare to cosmological galaxy formation simulations, and find that IllustrisTNG matches our measured combination of a nearly-invariant MZR slope, rate of MZR normalization decrease, and constant or very weakly evolving FMR. This work provides the most detailed view of MZR and FMR evolution from the present day through Cosmic Noon with a fine time sampling of $1-3$ Gyr, setting a robust baseline for metallicity evolution studies at $z>4$ with JWST.

astro-ph.GA

Cooperative SGD with Dynamic Mixing Matrices

One of the most common methods to train machine learning algorithms today is the stochastic gradient descent (SGD). In a distributed setting, SGD-based algorithms have been shown to converge theoretically under specific circumstances. A substantial number of works in the distributed SGD setting assume a fixed topology for the edge devices. These papers also assume that the contribution of nodes to the global model is uniform. However, experiments have shown that such assumptions are suboptimal and a non uniform aggregation strategy coupled with a dynamically shifting topology and client selection can significantly improve the performance of such models. This paper details a unified framework that covers several Local-Update SGD-based distributed algorithms with dynamic topologies and provides improved or matching theoretical guarantees on convergence compared to existing work.

cs.LG

The Multi-Stage Assignment Problem: A Fairness Perspective

This paper explores the problem of fair assignment on Multi-Stage graphs. A multi-stage graph consists of nodes partitioned into $K$ disjoint sets (stages) structured as a sequence of weighted bipartite graphs formed across adjacent stages. The goal is to assign node-disjoint paths to $n$ agents starting from the first stage and ending in the last stage. We show that an efficient assignment that minimizes the overall sum of costs of all the agents' paths may be highly unfair and lead to significant cost disparities (envy) among the agents. We further show that finding an envy-minimizing assignment on a multi-stage graph is NP-hard. We propose the C-Balance algorithm, which guarantees envy that is bounded by $2M$ in the case of two agents, where $M$ is the maximum edge weight. We demonstrate the algorithm's tightness by presenting an instance where the envy is $2M$. We further show that the cost of fairness ($CoF$), defined as the ratio of the cost of the assignment given by the fair algorithm to that of the minimum cost assignment, is bounded by $2$ for C-Balance. We then extend this approach to $n$ agents by proposing the DC-Balance algorithm that makes iterative calls to C-Balance. We show the convergence of DC-Balance, resulting in envy that is arbitrarily close to $2M$. We derive $CoF$ bounds for DC-Balance and provide insights about its dependency on the instance-specific parameters and the desired degree of envy. We experimentally show that our algorithm runs several orders of magnitude faster than a suitably formulated ILP.

cs.MA

Analysis of Propaganda in Tweets From Politically Biased Sources

News outlets are well known to have political associations, and many national outlets cultivate political biases to cater to different audiences. Journalists working for these news outlets have a big impact on the stories they cover. In this work, we present a methodology to analyze the role of journalists, affiliated with popular news outlets, in propagating their bias using some form of propaganda-like language. We introduce JMBX(Journalist Media Bias on X), a systematically collected and annotated dataset of 1874 tweets from Twitter (now known as X). These tweets are authored by popular journalists from 10 news outlets whose political biases range from extreme left to extreme right. We extract several insights from the data and conclude that journalists who are affiliated with outlets with extreme biases are more likely to use propaganda-like language in their writings compared to those who are affiliated with outlets with mild political leans. We compare eight different Large Language Models (LLM) by OpenAI and Google. We find that LLMs generally performs better when detecting propaganda in social media and news article compared to BERT-based model which is fine-tuned for propaganda detection. While the performance improvements of using large language models (LLMs) are significant, they come at a notable monetary and environmental cost. This study provides an analysis of both the financial costs, based on token usage, and the environmental impact, utilizing tools that estimate carbon emissions associated with LLM operations.

cs.SI

A Resilience Framework for Bi-Criteria Combinatorial Optimization with Bandit Feedback

We study bi-criteria combinatorial optimization under noisy function evaluations. While resilience and black-box offline-to-online reductions have been studied in single-objective settings, extending these ideas to bi-criteria problems introduces new challenges due to the coupled degradation of approximation guarantees for objectives and constraints. We introduce a notion of $(\alpha,\beta,\delta,\texttt{N})$-resilience for bi-criteria approximation algorithms, capturing how joint approximation guarantees degrade under bounded (possibly worst-case) oracle noise, and develop a general black-box framework that converts any resilient offline algorithm into an online algorithm for bi-criteria combinatorial multi-armed bandits with bandit feedback. The resulting online guarantees achieve sublinear regret and cumulative constraint violation of order $\tilde{O}(\delta^{2/3}\texttt{N}^{1/3}T^{2/3})$ without requiring structural assumptions such as linearity, submodularity, or semi-bandit feedback on the noisy functions. We demonstrate the applicability of the framework by establishing resilience for several classical greedy algorithms in submodular optimization.

cs.LG

FLIGHT: Facility Location Integrating Generalized, Holistic Theory of Welfare

The Facility Location Problem (FLP) is a well-studied optimization problem with applications in many real-world scenarios. Past literature has explored the solutions from different perspectives to tackle FLPs. These include investigating FLPs under objective functions such as utilitarian, egalitarian, Nash welfare, etc. Also, there is no treatment for asymmetric welfare functions around the facility. We propose a unified framework, FLIGHT, to accommodate a broad class of welfare notions. The framework undergoes rigorous theoretical analysis, and we prove some structural properties of the solution to FLP. Additionally, we provide approximation bounds, which provide insight into an interesting fact: as the number of agents arbitrarily increases, the choice of welfare notion is irrelevant. Furthermore, the paper also includes results around concentration bounds under certain distributional assumptions over the preferred locations of agents.

cs.GT

Fairness Driven Slot Allocation Problem in Billboard Advertisement

In billboard advertisement, a number of digital billboards are owned by an influence provider, and several commercial houses (which we call advertisers) approach the influence provider for a specific number of views of their advertisement content on a payment basis. Though the billboard slot allocation problem has been studied in the literature, this problem still needs to be addressed from a fairness point of view. In this paper, we introduce the Fair Billboard Slot Allocation Problem, where the objective is to allocate a given set of billboard slots among a group of advertisers based on their demands fairly and efficiently. As fairness criteria, we consider the maximin fair share, which ensures that each advertiser will receive a subset of slots that maximizes the minimum share for all the advertisers. We have proposed a solution approach that generates an allocation and provides an approximate maximum fair share. The proposed methodology has been analyzed to understand its time and space requirements and a performance guarantee. It has been implemented with real-world trajectory and billboard datasets, and the results have been reported. The results show that the proposed approach leads to a balanced allocation by satisfying the maximin fairness criteria. At the same time, it maximizes the utility of advertisers.

cs.GT

Self-regulated growth of galaxy sizes along the star-forming main sequence

We present a systematic analysis of the spatially resolved star formation histories (SFHs) using Hubble Space Telescope imaging data of $\sim 997$, intermediate redshifts $0.5 \leq z \leq 2.0$ galaxies from the GOODS-S field, with stellar mass range $9.8 \leq \log \mathrm{M}_{\star}/\mathrm{M}_{\odot} \leq 11.5$. We estimate the SFHs in three spatial regions (central region within the half-mass radii $\mathrm{R}_{50s}$, outskirts between $1-3~\mathrm{R}_{50s}$, and the whole galaxy) using pixel-by-pixel spectral-energy distribution (SED) fitting, assuming exponentially declining tau model in individual pixels. The reconstructed SFHs are then used to derive and compare the physical properties such as specific star-formation rates (sSFRs), mass-weighted ages (t$_{\mathrm{50}}$), and the half-mass radii to get insights on the interplay between the structure and star-formation in galaxies. The correlation of sSFR ratio of the center and outskirts with the distance from the main sequence (MS) indicates that galaxies on the upper envelope of the MS tend to grow outside-in, building up their central regions, while those below the MS grow inside-out, with more active star formation in the outskirts. The findings suggest a self-regulating process in galaxy size growth when they evolve along the MS. Our observations are consistent with galaxies growing their inner bulge and outer disc regions, where they appear to oscillate about the average MS in cycles of central gas compaction, which leads to bulge growth, and subsequent central depletion possibly due to feedback from the starburst, resulting in more star formation towards the outskirts from newly accreted gas.

astro-ph.GA

Fair Federated Data Clustering through Personalization: Bridging the Gap between Diverse Data Distributions

The rapid growth of data from edge devices has catalyzed the performance of machine learning algorithms. However, the data generated resides at client devices thus there are majorly two challenge faced by traditional machine learning paradigms - centralization of data for training and secondly for most the generated data the class labels are missing and there is very poor incentives to clients to manually label their data owing to high cost and lack of expertise. To overcome these issues, there have been initial attempts to handle unlabelled data in a privacy preserving distributed manner using unsupervised federated data clustering. The goal is partition the data available on clients into $k$ partitions (called clusters) without actual exchange of data. Most of the existing algorithms are highly dependent on data distribution patterns across clients or are computationally expensive. Furthermore, due to presence of skewed nature of data across clients in most of practical scenarios existing models might result in clients suffering high clustering cost making them reluctant to participate in federated process. To this, we are first to introduce the idea of personalization in federated clustering. The goal is achieve balance between achieving lower clustering cost and at same time achieving uniform cost across clients. We propose p-FClus that addresses these goal in a single round of communication between server and clients. We validate the efficacy of p-FClus against variety of federated datasets showcasing it's data independence nature, applicability to any finite $\ell$-norm, while simultaneously achieving lower cost and variance.

cs.LG

Optimizing Probabilistic Propagation in Graphs by Adding Edges

Probabilistic graphs are an abstraction that allow us to study randomized propagation in graphs. In a probabilistic graph, each edge is "active" with a certain probability, independent of the other edges. For two vertices $u,v$, a classic quantity of interest, that we refer to as the proximity $\mathcal{P}_{G}(u, v)$, is the probability that there exists a path between $u$ and $v$ all of whose edges are active. For a given subset of vertices $V_s$, the reach of $V_s$ is defined as the minimum over pairs $u \in V_s$ and $v \in V$ of the proximity $\mathcal{P}_{G}(u,v)$. This quantity has been studied in the context of multicast in unreliable communication networks and in social network analysis. We study the problem of improving the reach in a probabilistic graph via edge augmentation. Formally, given a budget $k$ of edge additions and a set of source vertices $V_s$, the goal of Reach Improvement is to maximize the reach of $V_s$ by adding at most $k$ new edges to the graph. The problem was introduced in earlier empirical work in the algorithmic fairness community. We provide the first approximation guarantees and hardness results for Reach Improvement. We prove that the existence of a good augmentation implies a cluster structure for the graph. We use this structural result to analyze a novel algorithm that outputs a $k$-edge augmentation with an objective value that is poly($\beta^*$), where $\beta^*$ is the objective value for the optimal augmentation. We also give an algorithm that adds $O(k \log n)$ edges and yields a multiplicative approximation to $\beta^*$. Our arguments rely on new probabilistic tools for analyzing proximity, inspired by techniques in percolation theory; these tools may be of broader interest. Finally, we show that significantly better approximations are unlikely, under known hardness assumptions related to gap variants of the classic Set Cover problem.

cs.DS

Towards Fairness in Provably Communication-Efficient Federated Recommender Systems

To reduce the communication overhead caused by parallel training of multiple clients, various federated learning (FL) techniques use random client sampling. Nonetheless, ensuring the efficacy of random sampling and determining the optimal number of clients to sample in federated recommender systems (FRSs) remains challenging due to the isolated nature of each user as a separate client. This challenge is exacerbated in models where public and private features can be separated, and FL allows communication of only public features (item gradients). In this study, we establish sample complexity bounds that dictate the ideal number of clients required for improved communication efficiency and retained accuracy in such models. In line with our theoretical findings, we empirically demonstrate that RS-FairFRS reduces communication cost (~47%). Second, we demonstrate the presence of class imbalance among clients that raises a substantial equity concern for FRSs. Unlike centralized machine learning, clients in FRS can not share raw data, including sensitive attributes. For this, we introduce RS-FairFRS, first fairness under unawareness FRS built upon random sampling based FRS. While random sampling improves communication efficiency, we propose a novel two-phase dual-fair update technique to achieve fairness without revealing protected attributes of active clients participating in training. Our results on real-world datasets and different sensitive features illustrate a significant reduction in demographic bias (~approx40\%), offering a promising path to achieving fairness and communication efficiency in FRSs without compromising the overall accuracy of FRS.

cs.IR