SearcharxivSearch

arXiv subjects

Haochen Wang

Publications and source records attributed to Haochen Wang.

At least 19 recordsLinked to original sources

The CHIME/FRB Outriggers: Commissioning the Hat Creek Outrigger and an Updated Calibration Scheme for Mitigating RFI

This work presents commissioning of the Hat Creek Outrigger (HCO), a dual-polarization 256-element radio interferometer that is part of the Canadian Hydrogen Intensity Mapping Experiment Fast Radio Burst (CHIME/FRB) Outrigger project. Driven by a complex radio-frequency interference (RFI) environment that consistently contaminates $\sim40\%$ of HCO's usable bandwidth, we implement an improved calibration scheme using Gaussian Process Regression to recover complex gain solutions over RFI contaminated channels. To validate our method, we test the performance of the array using known transients and continuum sources over relevant timescales used for fast-transient research ($\lesssim$ seconds). We find that our updated calibration scheme results in a $\sim1.69\times~\mathrm{to}~1.87\times$ improvement in the array's point-source sensitivity, while simultaneously maintaining noise properties consistent with thermal statistics. As a result, we find that the array performs consistently within theoretical expectations across $\gtrsim80\%$ of HCO's bandpass. We further observe an improvement in the interferometric performance after applying recently developed spatial filtering techniques for RFI mitigation, which rely on accurate calibration solutions for effective removal of unwanted interference. We conclude that our approach provides a valid framework for improving calibration solutions over RFI contaminated channels for large-$N$ interferometric arrays more broadly. Our work motivates future development of more sophisticated techniques to recover astrophysical information in RFI-contaminated channels, departing from the historical practice of discarding them outright.

astro-ph.IM

Building the Harness Automatically: Self-Play in Code Distills a Text Harness for Black-Box Optimization

Can an agent learn a numerical search strategy through executable practice and then transfer that strategy as text? We study low-budget black-box optimization, where unaided language models remain well below strong classical optimizers. During development, an agent repeatedly writes and evaluates optimizer programs. It then distills the resulting program and practice record once into a 197-word primary Harness A, which is frozen before evaluation. Harness A reduces Gemini Flash regret by 48\% in an independent $N=30$ study ($p<.001$), enters the GP-BO performance range on the practice family, and lowers mean regret on all three held-out BBOB landscapes. The same text improves every tested Gemini executor and transfers to Claude Sonnet, reducing regret by 43\% and 49\% ($p\leq.005$). An independent end-to-end replication produces Harness B, a different program and text at the same performance tier. The same framework also attains the lowest regret on a sealed YouTube reward-tuning production benchmark. Executable practice is thus a viable way to discover a search policy, and language a portable medium for deploying it.

cs.LG

Halide donors in monoclinic- and corundum-phase Ga$_2$O$_3$ and Al$_2$O$_3$

We present a systematic first-principles investigation of halide impurities (F and Cl) in Ga$_2$O$_3$ and Al$_2$O$_3$, considering both monoclinic and corundum phases. Our study of the structural properties, formation energies, and charge-state transition levels establishes the relative stability of different atomic configurations and charge states. We find that F and Cl on oxygen sites act as shallow donors in Ga$_2$O$_3$ in both the monoclinic and corundum phases. However, their behavior differs substantially as the band gap increases with greater Al compositions. Fluorine is prone to $DX$-center formation with increased Al composition, leading to self-compensation at 38% Al concentration in monoclinic (Al$_x$Ga$_{1-x}$)$_2$O$_3$ and 70% Al concentration in corundum (Al$_x$Ga$_{1-x}$)$_2$O$_3$. Chlorine is more resistant to $DX$-center formation: in monoclinic (Al$_x$Ga$_{1-x}$)$_2$O$_3$, Cl$_\mathrm{O}$ on the lowest-energy oxygen site shows an onset of $DX$ behavior at 50% alloy composition, while in corundum (Al$_x$Ga$_{1-x}$)$_2$O$_3$ this onset for Cl$_\mathrm{O}$ occurs only at Al concentrations as high as 84%. We also study F and Cl interstitials, finding that they act as compensating centers but also exhibit migration barriers that are low enough for them to be removed by post-growth annealing. Surprisingly, in both monoclinic and corundum Al$_2$O$_3$, Cl$_\mathrm{O}$ exhibits a relatively shallow transition level located at $\sim$0.48 eV below the conduction-band minimum, much shallower than F$_\mathrm{O}$ and other donor candidates. These remarkable results identify Cl as an unusually promising donor candidate in (Al$_x$Ga$_{1-x}$)$_2$O$_3$ alloys and even pure Al$_2$O$_3$, although high formation energies and compensation will render observation of true $n$-type conductivity in Al$_2$O$_3$ difficult.

cond-mat.mtrl-sci

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation

Captions serve as a primary supervision signal for both multimodal understanding and text-to-image generation. However, previous evaluations treat the caption quality as a single scalar objective, which conflates two distinct properties: (1) how much visual information a caption covers and (2) how reliably the image supports its stated claims. To this end, we design a decoupled caption evaluation benchmark, CAPEval (Coverage And Precision Evaluation), with human-written ground-truth captions and human-verified atomic checklist items. Specifically, CAPEval decomposes caption quality into Coverage and Precision. The former quantifies how thoroughly a caption covers ground-truth factual content, while the latter reflects the factual correctness rate of all claims expressed in the caption. We select 10 captioners and further conduct controlled downstream end-to-end experiments with them from four model families, where the caption source is the only variable. Empirically, we find a consistent task-dependent dissociation: Coverage serves as the stronger correlate for understanding performance, whereas Precision acts as the dominant predictor for generation performance. This decoupled evaluation paradigm not only delivers a more fine-grained diagnosis of caption quality, but also offers actionable guidance for selecting and optimizing captioners tailored to different downstream tasks.

cs.CV

Penelope: Localized Latent Recurrence for Efficient Structured Reasoning

Complex structured reasoning tasks often require additional computation, yet current language models obtain it mainly by increasing parameter scale or by serializing intermediate steps as chain-of-thought (CoT) tokens. The former raises training and deployment costs, while the latter ties reasoning computation to autoregressive output length. We introduce Penelope, an efficient latent-reasoning framework for pretrained decoder-only Transformers that localizes recurrent computation to a selected decoder interval. The lower decoder prefix is evaluated once to construct a problem-conditioned boundary memory, which is then iteratively refined through time-modulated GRU dynamics and recurrent readout states before answer generation. A progressive CoT-to-latent curriculum transfers visible reasoning into this internal recurrent path, allowing additional computation to be allocated in latent space without repeatedly executing the complete decoder or generating a long intermediate trace. Experiments on open-source structured-reasoning benchmarks show that, at validation-selected latent budgets, Penelope attains competitive accuracy relative to established latent-reasoning models while reducing measured inference latency. These results show that latent refinement can be localized to a narrow decoder interval, reducing repeated full-decoder execution without generating a long visible reasoning trace and providing a practical accuracy-efficiency tradeoff for decoder-only Transformer models.

cs.AI

The FRB--Galaxy Overdensity Cross-Correlation Statistic in Dispersion Space

Cross-correlating the dispersion of fast radio bursts (FRBs) with galaxies provides a means to study the distribution of the baryons in the Universe, even in the absence of FRB redshifts. To this end, two variants of angular cross-power spectrum statistics have been proposed: one between DM and galaxy density binned by redshift $C^{Dg}_l(z_g)$ (abbreviated $D \times g$), and one between FRB counts binned by dispersion measure (DM) and galaxy density binned by redshift $C^{fg}_l(\textrm{DM}, z_g)$ (abbreviated $f \times g$). Here we show the $D \times g$ statistic can be recovered as a DM-moment of $f \times g$, implying the latter is strictly more informative. By slicing in both DM space and galaxy redshift space, the $f\times g$ statistic separates contributions from the clustering of free electrons and from the clustering of FRB sources. We perform Fisher forecasts for FRB samples consistent with CHIME (1,600 FRBs) and the upcoming CHORD (20,000 FRBs) survey cross-correlated against the DESI Legacy Survey BGS sample and Euclid galaxy surveys, respectively. We show that, compared to the $D \times g$ statistic, the $f \times g$ statistic results in $S/N\approx 12$ for CHIME$\times$DESI(LS) and SNR $\approx 54$ for CHORD$\times$Euclid. The $f \times g$ statistics is more sensitive to the redshift distribution of FRBs with forecasted errors on a simple parameterization of order $10 \%$. It measures the logarithmic cutoff scale for clustering of baryons due to feedback $k_{cut}$ to 26\% precision with CHIME and 14\% precision with CHORD. Since most FRBs currently lack host identifications, and scaling optical followup to large samples will remain challenging even with precise localizations, reliable redshifts will be unavailable for most FRBs for the foreseeable future. The $f \times g$ statistic provides a means to extract maximum cosmological information in their absence.

astro-ph.CO

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO

This paper introduces Actor as Its Own Critic, a unified reinforcement learning framework, Cycle Group Relative Policy Optimization (CycleGRPO), that jointly optimizes region understanding and localization for Multimodal Large Language Models (MLLMs). Unlike existing separate pipelines, we leverage the inherent duality between the two tasks to construct a self-evaluating reinforcement learning paradigm: "region $\to$ text $\to$ region''. Specifically, a single MLLM first acts as the actor to generate region captions, then immediately transitions to a critic to ground its generated text back in the spatial domain. Therefore, CycleGRPO requires only region inputs, e.g., masks or bounding boxes, entirely bypassing the need for textual ground truths. A quality-aware token-level cycle-consistency reward is employed to assess the semantic discriminability of text captions via their physical localization accuracy. Empirically, built upon SAMTok, our CycleGRPO framework successfully bootstraps both capabilities simultaneously. Without any task-specific fine-tuning, the framework yields consistent performance gains across a wide range of benchmarks, including region captioning, region VQA, grounded dialogue, and referring segmentation. Overall, CycleGRPO offers a straightforward and scalable way to advance pixel-level capabilities in MLLMs. Code and models are released at https://github.com/devinxzhang/CycleGRPO.

cs.CV

PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seeking

This paper explores multi-turn visual reasoning and observes that MLLMs repeatedly fail to localize the target, leading to long, redundant trajectories. We attribute this failure to the entanglement of reasoning and perception within a single model, the MLLM reasons and localizes simultaneously, and inaccurate localization triggers additional reasoning turns that bloat the trajectory. To solve this problem, we propose PixelEyes, a multi-turn visual reasoning agent that explicitly decouples reasoning from perception, i.e., the reasoner decides what to look for, while a specialized perception tool answers where it is. Specifically, PixelEyes introduces 1) Mask-guided Visual Search. A referring segmentation model is invoked to provide mask-precise localization, freeing the reasoner from the need to compensate for imprecise grounding. 2) Semantic-region Breadth-first Search (BFS). To eliminate redundant loops caused by repeatedly cropping incorrect sub-regions, we organize exploration as a breadth-first search over semantic regions. To internalize these capabilities, we construct the PixelEyes-6K dataset by resynthesizing expert trajectories from existing data. This explicitly embeds our mask-guided search and BFS logic into the model. We further introduce Pinpoint-Bench, a zero-hint visual search benchmark, i.e., no location cues are provided in the question, with instance-level masks and bounding boxes that separate localization failures from reasoning failures, enabling fine-grained analysis of failure modes such as inattentional blindness. Recent state-of-the-art MLLMs and visual reasoning agents leave large headroom on Pinpoint-Bench, demonstrating its quality and difficulty. Code and models are open-sourced.

cs.CV

MotionAtlas: Detailed Region Captioning for Motion-Centric Videos

We propose MotionAtlas, a system for detailed captioning of motion-centric videos, comprising (1) a dedicated human-annotated benchmark, (2) a scalable, high-quality pipeline to construct training samples, and (3) a family of powerful Video-MLLMs. Unlike conventional global motion captioning datasets, we focus on region-aware motion captioning: given a video and a spatiotemporal mask, the model generates precise descriptions of motion within the target region, thereby alleviating visual clutter and motion entanglement and enabling reliable, quantifiable evaluation. Concretely, we first build MotionAtlas-Bench, a comprehensive benchmark comprising 2,073 multiple-choice questions, meticulously annotated for a curated set of high-quality, motion-centric videos, to evaluate fine-grained motion understanding of the objects in question. Second, we design a rigorous and scalable data pipeline that leverages self-bootstrap refinement to suppress fine-grained hallucinations, yielding 159k high-quality motion captioning data. Third, we design a tailored training data composition strategy, which achieves consistent and substantial performance gains across diverse baseline Video-MLLMs, including Molmo2 and Qwen3-VL. For instance, MotionAtlas-4B surpasses Qwen3-VL-4B by an average of 5.2 percentage points across general motion benchmarks. The benchmark, dataset, and code have been released.

cs.CV

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models

Multimodal large language models (MLLMs) have achieved remarkable progress in visual understanding tasks. However, most existing MLLMs rely on autoregressive generation, which limits their efficiency for perception tasks that require captioning multiple regions. In this work, we propose PerceptionDLM, a multimodal diffusion language model optimized for efficient parallel region perception. Built upon PerceptionDLM-Base, a strong foundational baseline that achieves state-of-the-art performance among open-source diffusion MLLMs, our architecture fully leverages the parallel decoding nature of DLMs. Specifically, we introduce efficient prompting and structured attention masking to enable simultaneous perception of multiple masked regions, allowing the model to generate region descriptions in parallel at both the sequence and token levels. This design significantly improves inference efficiency compared with existing approaches that process regions sequentially. To systematically evaluate the parallelism property of visual perception capability for DLMs, we construct a new Parallel Detailed Localized Captioning Benchmark (ParaDLC-Bench) by scaling the DLC-Bench to include multiple region masks per image, enabling joint evaluation of both caption quality and inference efficiency. Experiments demonstrate that PerceptionDLM maintains competitive performance in region captioning while achieving substantial speed improvements for multi-region perception tasks. Our results highlight the potential of multimodal diffusion language models for efficient, parallel visual perception. To the best of our knowledge, we are the first to achieve parallel region caption and perception by leveraging the advantages of diffusion language models. Code, models, and datasets are released.

cs.CV

DRIVE: Distributional and Retrieval-Augmented Bidding with Value Evaluation

Auto-bidding is a core component of real-time advertising systems, where decisions must optimize long-term performance under budget and cost constraints, while online exploration is prohibitively risky. Offline reinforcement learning and, more recently, Transformer-based sequence modeling have shown promise for learning bidding policies from logged data, but their unimodal and purely parametric formulations often collapse multiple effective bidding strategies into suboptimal averaged actions and perform unreliably under sparse or long-tail traffic. To mitigate these limitations, we propose DRIVE (Distributional and Retrieval-Augmented Bidding with Value Evaluation), a unified Transformer-based framework that decouples candidate action generation from decision making for offline auto-bidding. DRIVE combines distributional action modeling, retrieval-augmented candidate generation from high-quality historical decisions, and value-based evaluation to select the most promising bid at inference time. Extensive experiments on AuctionNet and additional offline reinforcement learning benchmarks demonstrate that DRIVE consistently improves bidding performance and generalizes well across multiple Transformer-based methods.

cs.LG

Watch, Remember, Reason: Human-View Video Understanding with MLLMs

Video understanding is being rapidly transformed by multimodal large language models (MLLMs), as research moves from short clips to long, multimodal, and knowledge-intensive video scenarios. These scenarios require models to handle sparse evidence, long-range dependencies, multimodal alignment, and reliable inference under limited computational budgets. This work presents a human-view perspective on LLM-based video understanding, organized around three functional abilities: watching, remembering, and reasoning. Rather than treating video tasks as isolated benchmarks, this view provides a unified structure for analyzing how video MLLMs acquire evidence, preserve context, and produce grounded outputs. We introduce a formulation that characterizes video understanding systems by their perceptual representations, memory states, reasoning traces, and final predictions. Based on this formulation, we identify challenges in spatio-temporal perception, efficient long-video processing, memory modeling, streaming understanding, and faithful reasoning. Representative methods are organized by their roles in video MLLM systems. Watching covers fine-grained, comprehensive, audio-visual, and efficient perception. Remembering includes offline and streaming memory, while reasoning covers text-only reasoning and thinking with videos. We further examine application domains such as egocentric, sports, instructional, medical, and narrative videos, and cover training datasets and evaluation benchmarks across task types, supervision formats, modalities, and capability dimensions. Finally, we outline open problems and future directions for scalable, memory-aware, and evidence-grounded video intelligence. Related works will be continuously traced at https://github.com/marinero4972/Awesome-HumanView-VideoUnderstanding.

cs.CV

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison

Long-form image captioning exposes a reward granularity problem in RL: captions are judged as whole sequences, while the important errors occur at the level of individual visual claims. A good dense caption should be both faithful and informative, avoiding hallucination without omitting salient details. Yet pairwise preferences, reference-based metrics, and holistic scalar rewards compress these local errors into a single sequence-level signal, obscuring the tradeoff between factuality and coverage. We introduce ClaimDiff-RL, a framework that uses reference-conditioned atomic claim differences as the reward unit for caption RL. Given an image, an actor caption, and a reference caption, a multimodal judge enumerates visually grounded differences, verifies each difference against the image, assigns open-vocabulary error types and severity levels, and produces per-difference statistics for reward composition. This makes hallucinated claims and omitted salient facts separately measurable and tunable. Experiments show that holistic scalar rewards can reduce hallucination by increasing missing facts, while ClaimDiff-RL exposes this faithfulness and coverage tradeoff and enables more balanced operating points. On a 160-image human-labeled diagnostic benchmark, public captioning benchmarks, and VQA benchmarks, ClaimDiff-RL improves the hallucination--missing-fact balance, preserves general capability, and even surpasses Gemini-3-Pro-Preview on several fine-grained Capability dimensions such as object counting, spatial relations, and scene recognition. These results suggest that typed, verifiable claim differences are an effective reward unit for fine-grained and diagnosable caption RL.

cs.LG

Fast radio burst dispersion is an unbiased tracer of matter on large scales

The dispersion of fast radio bursts (FRBs) measures the column density of free electrons, tracing the diffuse ionized gas that contains more than $90\%$ of all baryons. On linear scales the FRB dispersion field is an approximately unbiased tracer of the matter distribution, an idea long assumed in the FRB large-scale structure literature and recently formalized by Zhou and Zhang [arXiv:2510.11022]. This follows from baryon-mass conservation, which forces the total baryon field to have unit linear bias, with dispersion inheriting this bias up to small corrections from the stellar and neutral-gas components. We show these corrections can be bounded at the percent level using existing galaxy and 21 cm surveys, and confirm with the FLAMINGO hydrodynamical simulations that the electron bias varies at the percent level across a wide range of feedback prescriptions. The dispersion-galaxy cross-power spectrum at linear scales directly constrains $B_8 \equiv \sigma_8(\Omega_b/0.05)^{1/2}$, a baryonic analog of $S_8$, independently of feedback physics. Because most of the per-object variance in dispersion is cosmological signal rather than noise, $\sim\!10^5$ localized FRBs can match the statistical power of $\sim\!10^8$ weak-lensing galaxy shape measurements. FRB dispersion thus joins weak lensing and redshift-space distortions as a new unbiased tracer of matter on large scales.

astro-ph.CO

VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification

Recent video multimodal large language models achieve impressive results across various benchmarks. However, current evaluations suffer from two critical limitations: (1) inflated scores can mask deficiencies in fine-grained visual understanding and reasoning, and (2) answer correctness is often measured without verifying whether models identify the precise spatio-temporal evidence supporting their predictions. To address this, we present VideoZeroBench, a hierarchical benchmark designed for challenging long-video question answering that rigorously verifies spatio-temporal evidence. It comprises 500 manually annotated questions across 13 domains, paired with temporal intervals and spatial bounding boxes as evidence. To disentangle answering generation, temporal grounding, and spatial grounding, we introduce a five-level evaluation protocol that progressively tightens evidence requirements. Experiments show that even Gemini-3-Pro correctly answers fewer than 17% of questions under the standard end-to-end QA setting (Level-3). When grounding constraints are imposed, performance drops sharply: No model exceeds 1% accuracy when both correct answering and accurate spatio-temporal localization are required (Level-5), with most failing to achieve any correct grounded predictions. These results expose a significant gap between surface-level answer correctness and genuine evidence-based reasoning, revealing that grounded video understanding remains a bottleneck for long-video QA. We further analyze performance across minimal evidence spans, atomic abilities, and inference paradigms, providing insights for future research in grounded video reasoning. The benchmark and code will be made publicly available.

cs.CV

Interpretation of 21 cm Auto Power Spectrum Measurement at $z\sim 1$ by the Canadian Hydrogen Intensity Mapping Experiment

Observations with the Canadian Hydrogen Intensity Mapping Experiment (CHIME) have been used to measure the 21 cm intensity mapping auto power spectrum, at $z\sim 1$, over a frequency range from 608.2 MHz to 707.8 MHz at wavenumbers $0.4~h~{\rm Mpc}^{-1} \lesssim k \lesssim 1.5~h~{\rm Mpc}^{-1}$. In this paper, we present the results of two different approaches to interpreting this measurement. In the first approach, we use a parametric power spectrum model to constrain an amplitude parameter, defined as $\mathcal{A}^2_{\rm HI} \equiv 10^6 \Omega_{\rm HI}^2(b^2_{\rm HI}+\langle f \mu^2\rangle)^2$, where $\Omega_{\rm HI}$ is the cosmological density parameter for atomic hydrogen ($\rm HI$), $b_{\rm HI}$ is the linear bias for $\rm HI$, and $\langle f \mu^2\rangle$ incorporates the dominant large-scale impact of redshift-space distortions on the angle-averaged power spectrum. Imposing an additional prior on either $\Omega_{\rm HI}$ or $b_{\rm HI}$, based on values in the literature, allows us to break the pairwise degeneracy between those two parameters. In the second approach, we compare CHIME's measurement with predictions for the power spectrum of $\rm HI$ from the IllustrisTNG simulations, finding that the measurement disagrees with the TNG100 run at $3.1\sigma$ and the TNG300 run at $4.0\sigma$. This disagreement is most likely attributable to the strength of nonlinear redshift-space clustering of $\rm HI$ in the simulations, rather than the total abundance of $\rm HI$, and invites further investigation of the physical processes in the simulations that determine the behavior of $\rm HI$ at nonlinear scales. These results exemplify the ability of 21 cm intensity mapping to provide astrophysical information using measurements at nonlinear scales.

astro-ph.CO

A spatial filter for mitigating radio interference and its application to CHIME/FRB Outriggers

The sensitivity of radio telescopes is becoming increasingly limited by the presence of radio frequency interference (RFI), which will worsen as the radio spectrum becomes more crowded. One context where this poses a challenge is the field of fast radio burst (FRB) science, where there is increasing scientific interest in capturing as large of a population of bursts as possible and accurately measuring their celestial coordinates using interferometry. With several modern radio facilities actively collecting data for large FRB surveys that will be transformative to the field, properly mitigating unwanted interference is essential for the science goals of these surveys to be met. In this work, we present variations of a spatial filter based on the Karhunen-Loeve (KL) Transform to enhance the sensitivity of radio interferometers and demonstrate its applicability to FRB detection and localization. We derive a particular variation of the filter for the case of point-like radio pulses, which we show reduces to the maximum-signal-to-noise beamformer. We apply this filter to CHIME/FRB baseband data and demonstrate its capability to enhance the sensitivity and overall localization rate of CHIME/FRB Outriggers. We compare the cross-correlation signal-to-noise obtained using the spatial filter with that obtained using a spectral-kurtosis RFI flagger for a sample of 100 FRBs recorded by CHIME and its Outriggers, and show that this filter will double the total number of FRBs successfully localized with the CHIME/FRB Outrigger telescopes. While demonstrated here in the context of CHIME/FRB Outriggers, the spatial filter presented in this work--which we have made publicly available--is broadly applicable to other interferometric radio facilities engaged in FRB science and transient detection, including next-generation telescopes such as CHORD, DSA-2000, BURSTT, and CHARTS.

astro-ph.IM

Strain effects on $n$-type doping in AlN

Controllable doping in AlN and its alloys is essential for deep-ultraviolet light sources. Ionization energies for donors in AlN ($\mathrm{Si_{Al}}$, $\mathrm{S_N}$, $\mathrm{Se_N}$) are high. We report first-principles calculations demonstrating that strain engineering can result in a reduction in ionization energies. The donor levels for $\mathrm{S_N}$ and $\mathrm{Se_N}$ shift closer to the conduction-band minimum (CBM) under in-plane tensile strains, driven by a downward shift of the CBM. The most widely used donor, $\mathrm{Si_{Al}}$, forms a $DX$ center in AlN. We find that a 2.5% in-plane tensile strain (which would be induced by pseudomorphic growth on GaN in experiment) shifts the ($+/-$) transition level from 271 meV to 98 meV below the CBM, which would enhance the electron concentration by three orders of magnitude. These results demonstrate that strain engineering offers an effective route to enhance doping levels in AlN.

cond-mat.mtrl-sci