SearcharxivSearch

arXiv subjects

Yifan Mai

Publications and source records attributed to Yifan Mai.

At least 19 recordsLinked to original sources

Hector Galaxy Survey: Falling in Between - Infalling Galaxies in the Midst of the Abell 3667 Merger

Whether cluster mergers enhance ram pressure stripping (RPS) and accelerate member galaxy evolution remains an open question. Here, we investigate galaxy populations in the nearby merging cluster Abell 3667 ($z\simeq0.0553$) using spatially resolved data from the Hector Galaxy Survey. We define an RPS sample combining Hector-selected galaxies with ionised gas disturbances (e.g., asymmetric tails or truncated disks) and supplementary, optically identified jellyfish galaxies lacking Hector data. Most of the RPS sample ($\sim 71^{+10}_{-7}\%$; 20/28) lies within $R_{200}$, where the merger impact is greater. Most asymmetric galaxies ($\sim 73^{+14}_{-8}\%$; 11/15), especially those with extreme RPS signatures, are concentrated in the inner cluster ($R \lesssim 0.6\, R_{200}$), along the merger axis between two shock-tracing radio relics. These central asymmetric galaxies show two spatial and kinematic groups: one at the North-West (NW) subcluster, downstream of its radio relic in a region of high-velocity intracluster medium (ICM) bulk motion, with blueshifted line-of-sight velocities; and a mostly redshifted population near the main cluster (MC), which also shows a turbulent ICM. Despite their projected association with the MC core and NW substructure, both samples' velocities indicate they are not bound to them. Tail orientations give insight into orbital histories: NW tails point away from the cluster centre and often align with the merger axis, suggesting merger-driven stripping, while MC tails show neither pattern clearly. Tails are broadly westward, with MC tails tracing due west and NW tails shifted northwest, pointing to two distinct filamentary accretion events for the NW and MC populations. Together, our results indicate enhanced RPS in the heart of A3667, driven mainly by infalling galaxies accreted along nearby filaments interacting with the merger-driven turbulent environment.

astro-ph.GA

DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots

Mental health professionals have raised concerns about risks of psychological harm from interaction with large language models (LLMs), including "delusional spirals" in which concerning human and LLM behaviors reinforce each other over time. With growing public use of LLM-powered chatbots, there is an urgent need to build evaluations grounded in real-world episodes of psychological harm experienced by users. We developed DelusionEval, an evaluation protocol that tests a model's tendencies to exhibit behaviors linked to promoting user delusions. We prompt each model with 589 unique conversation histories from 18 participants, comprising 12,591 messages from users who experienced delusions and psychological harm. We find that the tendency of an evaluated LLM to exhibit delusion-linked behavior does not reliably correlate with model size, release date, or the presence of test-time reasoning. However, extending the context of prior messages substantially increases rates of delusion-linked behaviors, providing evidence for the importance of context in LLM safety evaluation. For example, the rate of failing to discourage self-harm when the user expresses suicidal ideation increases from 30.0% to 41.1% when an additional 350 messages are prepended to the conversation history. All model families (e.g., GPT, Claude) exhibit substantial rates of delusion-linked behaviors. Within families, later, larger, or higher-reasoning models are not uniformly better across all behavior categories. Our results raise concerns regarding the potential psychological impact of LLMs and the need for more rigorous studies of real-world human-AI interaction.

cs.CL

Expert Evaluation of Clinical AI Tools on Real Point-of-Care Clinical Queries

Physicians now pose millions of clinical questions to AI tools each week, yet these tools are evaluated largely on hypothetical or exam-style questions, not those actually asked in practice. We report a blinded evaluation built on 620 Real-world Point-Of-Care Queries (Real-POCQi) submitted to the OpenEvidence (OE) platform by physicians spanning 30 specialties, as well as 187 questions from HealthBench. 149 practicing physicians across 36 states made head-to-head comparisons between answers from three frontier general-purpose models (Claude Opus 4.8, Gemini 3.1 Pro, and GPT-5.5) and a specialized clinical tool (OE), with graders matched to each question's specialty. When comparing answers along five dimensions relevant to clinical decision support -- accuracy, clinical utility, source quality, verifiability, & completeness -- physicians scored the specialized tool highest on all axes; in the primary analysis on Real-POCQi, win differences (margins between win and loss rates) ranged from 25 to 39 percentage points (p<0.001). Results remained consistent in sensitivity analyses stratifying by citation display, answer length, OE-user status, and Real-POCQi versus HealthBench. In parallel, LLM judges were found to systematically differ from expert judges, though both generally agreed on the best model. These findings underscore two conclusions: (i) AI tool evaluations should reflect real-world query distributions and use expert judges that mirror the specialization defining modern medicine and (ii) the consistent advantage of the specialized tool over general-purpose models does not necessarily mean that the latter cannot serve similar purposes, but that targeted engineering and customization can yield meaningful gains in performance for its users. We release Real-POCQi as a public benchmark, as well as the prespecified statistical analysis for reproducing results of this study.

cs.AI

Hector Galaxy Survey: Linking the low- and high-mass ends of the initial mass function in star-forming galaxies

The stellar initial mass function (IMF) is a fundamental ingredient in galaxy evolution, linking observed integrated light to galaxy properties. Constraining the full IMF shape beyond the Milky Way remains challenging, as most studies focus either on the low-mass end of quiescent galaxies or the high-mass end of star-forming galaxies. Here we present the first simultaneous analysis of both ends of the IMF in 214 star-forming galaxies from the Hector survey. We estimate the low-mass end slope using a stellar population approach that fits IMF-sensitive absorption features with extended star formation histories, while the high-mass end slope is derived via the Kennicutt diagnostic, which compares the observed H-alpha equivalent width and g-r colour with stellar population synthesis model predictions. We find substantial diversity in IMF shapes and a weak but statistically robust correlation between the low- and high-mass IMF slopes. Both IMF slopes show significant correlations with stellar mass, star formation activity, and stellar metallicity ([M/H]). In general, higher stellar mass, stronger star formation activity, and higher metallicity are associated with both bottom-heavy and top-heavy IMFs. Partial correlation analysis reveals that the low-mass end slope is primarily driven by [M/H], whereas the high-mass end is mainly linked to stellar mass and recent star formation. Because the low-mass end slope traces the IMF over long-term averages and the high-mass end slope captures only recent star formation, the processes shaping each end likely occur over different and possibly decoupled timescales. Our findings challenge the universality of the IMF and emphasise the need for galaxy evolution and stellar population models to incorporate a flexible IMF prescription. Accounting for these variations is essential to build an IMF-consistent picture of galaxy evolution across cosmic time.

astro-ph.GA

Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results

AI evaluations are widely used for testing and understanding progress. However, the diverse evaluators bring with them inconsistencies that challenge analysis and comparison. First, results are saved in incompatible formats, scattered across leaderboards, papers, blog posts, evaluation harness logs, and custom repositories. Second, results are created by different evaluation frameworks, which produce divergent scores for nominally identical evaluations and record metadata inconsistently, hindering comparison, cross-community evaluation science, cost reduction, and reuse. We introduce Every Eval Ever, the first shared schema and community-crowdsourced repository for AI evaluation results. The schema standardizes how evaluations are represented in a unified, single JSON document. It is source-agnostic by design, ingesting results from evaluation harnesses and papers alike, and optionally stores per-instance outputs for fine-grained analysis. We contribute: (i) a community-governed metadata schema with a companion instance-level schema, the first standardization effort of its kind; (ii) automatic converters from popular formats, evaluation harnesses, and leaderboards to the unified schema; and (iii) a crowdsourced community database hosted on Hugging Face, currently spanning to date 22,235 models, 2,273 unique benchmarks, and 31 evaluation formats.

cs.AI

Large-scale and local environmental drivers of quenching: tracing H$α$ concentration in X-ray and optical galaxy groups

To explore the environmental mechanisms causing quenching in nearby star-forming galaxies, we study the variation with local and large-scale environments of a star formation concentration index, C-index $\equiv\log{(r_{50,{\rm H}α}/r_{50,\rm cont}})$, that traces the spatially-resolved distribution of H$α$ emission. Our analysis combines (i) GAMA spectroscopic redshift survey data to optically select galaxy groups and reconstruct the cosmic web, (ii) eROSITA data to identify X-ray-emitting groups, and (iii) SAMI Galaxy Survey data to characterise spatially-resolved star formation. We find that galaxies in X-ray+optical groups exhibit the lowest median C-index and the highest fraction of centrally-concentrated star-forming galaxies relative to optical groups and the field (independently of group or stellar mass). Star-forming galaxies in more X-ray luminous groups at fixed dynamical mass show more concentrated star formation. At large scales, nodes show the lowest median C-index and the highest fraction of centrally-concentrated star-forming galaxies relative to filaments and voids, which have similar C-index distributions. C-index correlates most strongly with the distance to the closest node, leaving no significant role for other local or large-scale environment metrics. Finally, regular star-forming galaxies tend to have spins aligned parallel to filaments, consistent with smooth gas accretion, while centrally-concentrated galaxies tend have spins aligned perpendicular to filaments, likely driven by mergers and associated with bulge growth. These results suggest that multi-scale environmental processes, i.e. locally and at large-scale, act to concentrate star formation toward galaxy centres, via gas-related mechanisms in nodes and ram-pressure stripping in X-ray+optical groups.

astro-ph.GA

Adaptive auditing of AI systems with anytime-valid guarantees

A major bottleneck in characterizing the failure modes of generative AI systems is the cost and time of annotation and evaluation. Consequently, adaptive testing paradigms have gained popularity, where one opportunistically decides which cases and how many to annotate based on past results. While this framework is highly practical, its extreme flexibility makes it difficult to draw statistically rigorous conclusions, as it violates classical assumptions: the number of observations is typically limited (often 10 to 50 cases) and decisions regarding sampling and stopping are made in the midst of data collection rather than based a pre-specified rule. To characterize what statistical inferences can be drawn from highly adaptive audits, we introduce a hypothesis testing framework from two 'dueling' perspectives: (i) the model's null that asserts there is no failure mode with performance below a target threshold versus (ii) the auditor's null that asserts they have a sampling strategy that will uncover a failure mode. Leveraging Safe Anytime-Valid Inference (SAVI), we formalize the auditor as conducting 'testing by betting', which translates into simultaneous e-processes for testing the dueling null hypotheses. Furthermore, if the auditor is sufficiently powerful, we prove that these two hypotheses are asymptotically inverses of each other, in that passage of a stringent audit does in fact certify the AI system as being globally robust. Empirically, we demonstrate that our proposed testing procedures maintain anytime-valid type-I error control, outperform pre-specified testing methods, and can reach statistically rigorous conclusions sometimes with as few as 20 observations.

cs.AI

Optimization before Evaluation: Evaluation with Unoptimised Prompts Can be Misleading

Current Large Language Model (LLM) evaluation frameworks utilize the same static prompt template across all models under evaluation. This differs from the common industry practice of using prompt optimization (PO) techniques to optimize the prompt for each model to maximize application performance. In this paper, we investigate the effect of PO towards LLM evaluations. Our results on public academic and internal industry benchmarks show that PO greatly affects the final ranking of models. This highlights the importance of practitioners performing PO per model when conducting evaluations to choose the best model for a given task.

cs.AI

LLMs Judging LLMs: A Simplex Perspective

Given the challenge of automatically evaluating free-form outputs from large language models (LLMs), an increasingly common solution is to use LLMs themselves as the judging mechanism, without any gold-standard scores. Implicitly, this practice accounts for only sampling variability (aleatoric uncertainty) and ignores uncertainty about judge quality (epistemic uncertainty). While this is justified if judges are perfectly accurate, it is unclear when such an approach is theoretically valid and practically robust. We study these questions for the task of ranking LLM candidates from a novel geometric perspective: for $M$-level scoring systems, both LLM judges and candidates can be represented as points on an $(M-1)$-dimensional probability simplex, where geometric concepts (e.g., triangle areas) correspond to key ranking concepts. This perspective yields intuitive theoretical conditions and visual proofs for when rankings are identifiable; for instance, we provide a formal basis for the ``folk wisdom'' that LLM judges are more effective for two-level scoring ($M=2$) than multi-level scoring ($M>2$). Leveraging the simplex, we design geometric Bayesian priors that encode epistemic uncertainty about judge quality and vary the priors to conduct sensitivity analyses. Experiments on LLM benchmarks show that rankings based solely on LLM judges are robust in many but not all datasets, underscoring both their widespread success and the need for caution. Our Bayesian method achieves substantially higher coverage rates than existing procedures, highlighting the importance of modeling epistemic uncertainty.

cs.LG

Structured Prompts Improve Evaluation of Language Models

As language models (LMs) are increasingly adopted across domains, high-quality benchmarking frameworks are essential for guiding deployment decisions. In practice, however, frameworks such as Holistic Evaluation of Language Models (HELM) typically evaluate models under a single static prompt configuration, even though model behavior depends strongly on prompt choice. As a result, reported scores can reflect prompt choice as much as model capability. Declarative prompting frameworks such as DSPy offer a scalable way to evaluate models under a set of structured prompting strategies rather than a static prompt configuration. We present a reproducible DSPy+HELM framework for studying how prompt choice impacts reported benchmark outcomes. Using five prompting methods, we evaluate four frontier and two open-source LMs across seven benchmarks against existing HELM baseline scores. By evaluating LMs across a family of prompt configurations, we find that prompt choice can materially impact leaderboard outcomes. In particular, structured prompting improves performance (by 6% on average), alters comparisons (leaderboard rankings shift on 5/7 benchmarks), with most gains coming from introducing chain-of-thought, and little additional benefit from more advanced optimizers. To our knowledge, this is the first study to systematically integrate structured prompting into an established evaluation framework and quantify how prompt choice alone can impact benchmark conclusions. We open-source (i) DSPy+HELM Evaluation (https://github.com/stanford-crfm/helm/pull/3893) and (ii) Prompt Optimization Pipeline (https://github.com/StanfordMIMI/dspy-helm).

cs.CL

Characterizing Delusional Spirals through Human-LLM Chat Logs

As large language models (LLMs) have proliferated, disturbing anecdotal reports of negative psychological effects, such as delusions, self-harm, and ``AI psychosis,'' have emerged in global media and legal discourse. However, it remains unclear how users and chatbots interact over the course of lengthy delusional ``spirals,'' limiting our ability to understand and mitigate the harm. In our work, we analyze logs of conversations with LLM chatbots from 19 users who report having experienced psychological harms from chatbot use. Many of our participants come from a support group for such chatbot users. We also include chat logs from participants covered by media outlets in widely-distributed stories about chatbot-reinforced delusions. In contrast to prior work that speculates on potential AI harms to mental health, to our knowledge we present the first in-depth study of such high-profile and veridically harmful cases. We develop an inventory of 28 codes and apply it to the $391,562$ messages in the logs. Codes include whether a user demonstrates delusional thinking (15.5% of user messages), a user expresses suicidal thoughts (69 validated user messages), or a chatbot misrepresents itself as sentient (21.2% of chatbot messages). We analyze the co-occurrence of message codes. We find, for example, that messages that declare romantic interest and messages where the chatbot describes itself as sentient occur much more often in longer conversations, suggesting that these topics could promote or result from user over-engagement and that safeguards in these areas may degrade in multi-turn settings. We conclude with concrete recommendations for how policymakers, LLM chatbot developers, and users can use our inventory and conversation analysis tool to understand and mitigate harm from LLM chatbots. Warning: This paper discusses self-harm, trauma, and violence.

cs.CL

Talk, Evaluate, Diagnose: User-aware Agent Evaluation with Automated Error Analysis

Agent applications are increasingly adopted to automate workflows across diverse tasks. However, due to the heterogeneous domains they operate in, it is challenging to create a scalable evaluation framework. Prior works each employ their own methods to determine task success, such as database lookups, regex match, etc., adding complexity to the development of a unified agent evaluation approach. Moreover, they do not systematically account for the user's role nor expertise in the interaction, providing incomplete insights into the agent's performance. We argue that effective agent evaluation goes beyond correctness alone, incorporating conversation quality, efficiency and systematic diagnosis of agent errors. To address this, we introduce the TED framework (Talk, Evaluate, Diagnose). (1) Talk: We leverage reusable, generic expert and non-expert user persona templates for user-agent interaction. (2) Evaluate: We adapt existing datasets by representing subgoals-such as tool signatures, and responses-as natural language grading notes, evaluated automatically with LLM-as-a-judge. We propose new metrics that capture both turn efficiency and intermediate progress of the agent complementing the user-aware setup. (3) Diagnose: We introduce an automated error analysis tool that analyzes the inconsistencies of the judge and agents, uncovering common errors, and providing actionable feedback for agent improvement. We show that our TED framework reveals new insights regarding agent performance across models and user expertise levels. We also demonstrate potential gains in agent performance with peaks of 8-10% on our proposed metrics after incorporating the identified error remedies into the agent's design.

cs.AI

OCR or Not? Rethinking Document Information Extraction in the MLLMs Era with Real-World Large-Scale Datasets

Multimodal Large Language Models (MLLMs) enhance the potential of natural language processing. However, their actual impact on document information extraction remains unclear. In particular, it is unclear whether an MLLM-only pipeline--while simpler--can truly match the performance of traditional OCR+MLLM setups. In this paper, we conduct a large-scale benchmarking study that evaluates various out-of-the-box MLLMs on business-document information extraction. To examine and explore failure modes, we propose an automated hierarchical error analysis framework that leverages large language models (LLMs) to diagnose error patterns systematically. Our findings suggest that OCR may not be necessary for powerful MLLMs, as image-only input can achieve comparable performance to OCR-enhanced approaches. Moreover, we demonstrate that carefully designed schema, exemplars, and instructions can further enhance MLLMs performance. We hope this work can offer practical guidance and valuable insight for advancing document information extraction.

cs.CL

The Mighty ToRR: A Benchmark for Table Reasoning and Robustness

Despite its real-world significance, model performance on tabular data remains underexplored, leaving uncertainty about which model to rely on and which prompt configuration to adopt. To address this gap, we create ToRR, a benchmark for Table Reasoning and Robustness, measuring model performance and robustness on table-related tasks. The benchmark includes 10 datasets that cover different types of table reasoning capabilities across varied domains. ToRR goes beyond model performance rankings, and is designed to reflect whether models can handle tabular data consistently and robustly, across a variety of common table representation formats. We present a leaderboard as well as comprehensive analyses of the results of leading models over ToRR. Our results reveal a striking pattern of brittle model behavior, where even strong models are unable to perform robustly on tabular data tasks. Although no specific table format leads to consistently better performance, we show that testing over multiple formats is crucial for reliably estimating model capabilities. Moreover, we show that the reliability boost from testing multiple prompts can be equivalent to adding more test examples. Overall, our findings show that table understanding and reasoning tasks remain a significant challenge.

cs.CL

ErrorMap and ErrorAtlas: Charting the Failure Landscape of Large Language Models

Large Language Models (LLM) benchmarks tell us when models fail, but not why they fail. A wrong answer on a reasoning dataset may stem from formatting issues, calculation errors, or dataset noise rather than weak reasoning. Without disentangling such causes, benchmarks remain incomplete and cannot reliably guide model improvement. We introduce ErrorMap, the first method to chart the sources of LLM failure. It extracts a model's unique "failure signature", clarifies what benchmarks measure, and broadens error identification to reduce blind spots. This helps developers debug models, aligns benchmark goals with outcomes, and supports informed model selection. ErrorMap works on any model or dataset with the same logic. Applying our method to 35 datasets and 83 models we generate ErrorAtlas, a taxonomy of model errors, revealing recurring failure patterns. ErrorAtlas highlights error types that are currently underexplored in LLM research, such as omissions of required details in the output and question misinterpretation. By shifting focus from where models succeed to why they fail, ErrorMap and ErrorAtlas enable advanced evaluation - one that exposes hidden weaknesses and directs progress. Unlike success, typically measured by task-level metrics, our approach introduces a deeper evaluation layer that can be applied globally across models and tasks, offering richer insights into model behavior and limitations. We make the taxonomy and code publicly available with plans to periodically update ErrorAtlas as new benchmarks and models emerge.

cs.AI

The SAMI Galaxy Survey: Quenching of Star Formation in Clusters III. Ram-Pressure-Affected Galaxy Populations

Cluster environments influence galaxy evolution by curtailing star formation activity, notably through ram-pressure stripping (RPS). In this study, using spatially resolved spectroscopic data from the SAMI Galaxy Survey, we identify galaxies undergoing or recently affected by RPS in eight nearby clusters ($0.029 < z < 0.058$), through a visual classification scheme based on the ionised gas ($\rm Hα+ [NII]λ6584$) morphologies, split into unperturbed, asymmetric, and truncated. The projected phase-space analysis shows that asymmetric galaxies are found in a narrow region in cluster-centric distance ($\rm 0.1 < R/R_{200} < 0.6$) and have a larger dispersion in line-of-sight velocity ($σ(|v_{pec}|)_\mathrm{Asym} = 0.71^{+0.09}_{-0.07}\ σ_{200}$) compared to the truncated and unperturbed samples. In terms of star formation activity, RPS candidates yield a much steeper resolved star-forming main sequence (rSFMS; $Σ_\mathrm{SFR} - Σ_\ast$) relation compared to the unperturbed counterparts, primarily emerging from having lower $Σ_\mathrm{SFR}$ values for the low mass density regime, with the steepest gradient deriving from the truncated sample. Moreover, radial star formation profiles reveal that star formation in RPS candidates is suppressed in the outskirts relative to unperturbed galaxies and is more prominent for the truncated sample. In contrast, central ($\rm r/r_{eff}<0.5$) star formation activity in RPS candidates is comparable with that in their unperturbed and field counterparts, suggesting no elevated activity. Taken together, this suggests an evolutionary trend linked to the RPS stage, where unperturbed galaxies likely represent recently accreted systems (pre-RPS), while asymmetric and truncated galaxies may correspond to populations undergoing RPS and post-RPS phases, respectively, favouring outside-in quenching.

astro-ph.GA

The MAGPI Survey: co-evolution of baryons and dark matter in star-forming disk-like galaxies at $0.1 \lesssim z \lesssim 0.85$

We present a comprehensive analysis of the dark matter (DM) content and its structural dependence in star-forming disk-like galaxies at intermediate redshifts ($0.1 \lesssim z \lesssim 0.85$), utilizing spatially resolved kinematic data from the MAGPI survey. We report the following: (1) Low stellar mass galaxies ($M_{\rm star} < 10^{9.5}\, M_\odot$) are strongly DM dominated across all radii, with average $\langle f_{_{\rm DM}} \rangle \sim 0.85$, while high-mass ($M_{\rm star} > 10^{10.5}\, M_\odot$) systems exhibit relatively low DM fractions in their inner regions ($\langle f_{_{\rm DM}} \rangle \sim 0.47$) which is equivalent to local massive disk galaxies (e.g., Milky Way and Andromeda). This suggests a mass-dependent structural dichotomy, most-likely governed by a combination of internal galactic processes and environmental influences. (2) A tight inverse correlation between $f_{_{\rm DM}}$ and baryon mass surface density ($Σ_{\rm bar}$), with intrinsic scatter of $\sim 0.11$ dex. This is consistent with an inside-out baryon assembly scenario and suggests that the fundamental structural correlations of galaxies were already established by $z\sim 0.85$. (3) No significant evolution in $f_{_{\rm DM}}$ with redshift across the MAGPI window, and when combined with higher-redshift ($0.6 \leq z \leq 1.5$) data from Sharma et al. 2025, we quantitatively show that the reported decline in $f_{_{\rm DM}}(z)$ is most-likely due to observational biases against low-mass systems at $z > 1$. These results offer empirical evidence for a scenario in which disk-like galaxies evolve through a co-regulated build-up of baryonic and DM components, preserving internal structural regularities (such as the total mass distribution and rotation-curve shape) throughout cosmic time.

astro-ph.GA

The MAGPI Survey: forward modelled gas-phase metallicity gradients in galaxies at $z\sim 0.3$

We measure the seeing-deconvolved gas-phase metallicity gradients of 70 star-forming galaxies at $z\sim 0.3$ from the MAGPI survey and investigate their relationship with galaxy properties to understand the mechanisms that influence the distribution of metals and shape the evolution of the galaxies. We use a Bayesian modelling technique, Blobby3D, which accounts for seeing effects (beam smearing) and can model the substructures of the flux distribution. The median metallicity gradient of our sample is $\nabla \mathrm{[O/H]}=-0.013^{+0.059}_{-0.033}$ dex/kpc. Among the galaxies in our sample, 32.9% have negative metallicity gradients (2$σ$ significance), 10.0% have positive gradients and 57.1% have flat gradients. The $\nabla \mathrm{[O/H]}$-$M_*$ relation of the MAGPI galaxies generally agrees with theoretical predictions, where a combination of stellar feedback, gas transport, and accretion shapes the metallicity profile, with the dominant processes varying with galaxy mass. We find a positive correlation between $\nabla \mathrm{[O/H]}$ and gas velocity dispersion ($r=0.36$), indicating that stronger gas turbulence is associated with flatter or inverted metallicity gradients, likely due to enhanced gas mixing. Additionally, smaller galaxies tend to have flatter or positive gradients, suggesting that metal dilution by gas accretion or removal via feedback-driven winds may outweigh metal enrichment in small galaxies.

astro-ph.GA