SearcharxivSearch

arXiv subjects

Wei Du

Publications and source records attributed to Wei Du.

At least 19 recordsLinked to original sources

An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics

We study how model post-training and test-time inference design affect natural-language proof generation for hard olympiad mathematics. Starting from Nemotron 3 Ultra, we train two specialist checkpoints using supervised fine-tuning and reinforcement learning, and evaluate checkpoint choice, verification, and refinement. Based on these findings, we present an open-model test-time-compute pipeline. The system operates entirely in natural language, with no formal prover, external tools, or internet access. Three Nemotron 3 Ultra checkpoints - the general-availability model and two post-trained specialists - power an iterative search that generates, verifies, and refines candidate proofs; a separate high-compute stage then selects each final submission. The system scored 30 out of 42 points at IMO 2026, reaching the gold-medal threshold. We release the two post-trained checkpoints as well as the training data, the training and inference code, the submitted solutions, and Nemotron-IMO-Bench, a new benchmark of 200 novel olympiad-level problems.

cs.AI

Quantum Geometric Origin of Nonlinear Current Induced Orbital Magnetization

Electric generation of magnetization is a focus of condensed matter research, and has recently been advanced into the nonlinear regime. However, due to the nonlocal nature of orbital magnetism, how to properly formulate nonlinear current-induced orbital magnetization remains a fundamental challenge. Here, we develop the proper theory for this effect. This is based on the microscopic derivation of field-corrected orbital magnetic moment of a Bloch electron, a critical missing piece in the present theory. We show that the quantum geometric origin of this phenomenon lies in both the anomalous orbital polarizability and the Berry-connection polarizability, which often provide competing contributions. Combining our theory with first-principles calculations, we predict significant, experimentally accessible nonlinear orbital magnetization generated in strained bilayer graphene, monolayer 1T' $\mathrm{MoS_2}$ and $\mathrm{MoTe_2}$. Remarkably, nonlinear orbital magnetization can dominate over its spin counterpart in materials with topological band features, irrespective of the spin-orbit coupling strength.

cond-mat.mes-hall

Investigation of GeSn aspect ratio trapping growth up to 8% Sn

Aspect ratio trapping (ART) growth of germanium-tin (GeSn) is a promising approach to target important objectives on the quest towards commercialization of complementary metal-oxide-semiconductor (CMOS)-compatible GeSn optoelectronics devices. Its local growth on patterned substrate allows for versatile device integration into photonics integrated circuit or for stand-alone structure like focal plane array imager. Additionally, high aspect ratio from nano-sized window can terminate early threading dislocation propagation on the oxide sidewalls, leaving subsequent growth defect-free and potentially improving the device performance. Knowledge remains missing regarding GeSn ART growth kinetics, morphology and how they evolve from thin film growth, with successful growth itself yet to be demonstrated. In this work, we report GeSn ART growth up to 8% Sn. Two configurations -- self-induced Ge core/GeSn shell for Sn content between 6% and 8%, and bulk GeSn ART for Sn content below 1% -- are observed. We present a comprehensive study on GeSn ART growth kinetics through different growth rounds and designs, showing a link between pyramid shape of ART island and successful Sn incorporation, as well as the role of growth selectivity and local heating.

cond-mat.mtrl-sci

A RELHIC twin candidate near the galaxy M51

We report the discovery of a pair of H I clouds near M51 (NGC 5194) using the Five-hundred-meter Aperture Spherical radio Telescope (FAST). These clouds have no optical counterparts and are potential candidates for Reionization-Limited H I Clouds (RELHICs). We search for compact H I sources in deep FEASTS observations using SoFiA and remove objects with optical counterparts through cross-matching with the DESI Legacy Imaging Surveys. The remaining candidates are modelled as hydrostatic H I structures embedded in Navarro-Frenk-White dark matter haloes and compared with RELHIC predictions and TNG50 simulations. We identify two H I clouds, Cloud S and Cloud N, at projected distances of 70--90 kpc from M51. Each cloud has an H I mass of approximately 10^6.5 solar masses, a velocity dispersion of about 20 km/s, and no detectable optical counterpart down to a g-band surface brightness limit of approximately 27.5 mag arcsec^-2. Their stellar luminosities are constrained to be below 10^5 solar luminosities. Their H I properties are consistent with RELHIC predictions, corresponding to host halo masses of 3.7 +- 0.4 * 10^9 solar masses. Cloud S and Cloud N are promising but not definitive RELHIC candidates. A tidal origin remains possible in the interacting M51 system, especially because the clouds are unresolved by FAST and Cloud N may show a velocity gradient. Future high-resolution interferometric observations will be crucial for distinguishing between starless dark matter haloes and tidal debris.

astro-ph.GA

Visual Information Extraction from Documents via Classification-Guided Large Vision-Language Models

Visual information extraction (VIE) from visually rich documents remains challenging due to high layout variability and real-world impairments. Existing methods typically rely on sequential OCR pipelines or end-to-end models requiring extensive labeled data and layout-specific training, limiting their scalability.We propose a classification-guided large vision-language model (LVLM) framework for multi-type VIE that achieves high accuracy with minimal supervision. The approach decouples document-type classification from content extraction and employs in-context learning (ICL)-based dynamic prompt engineering to inject task-specific knowledge, enabling robust zero-shot inference across diverse layouts. From a theoretical perspective, the proposed method can be viewed as a form of conditional computation that reduces task uncertainty and improves information efficiency during prompt-based inference. Evaluated on a real-world bidding dataset with 16 certificate types, our zero-shot method (based on Qwen2.5-VL-7B) outperforms a strong supervised baseline by 18.35 percentage points in F1-score (86.43\% vs. 68.08\%) and 0.23 in normalized edit distance (0.90 vs. 0.67). Optional domain-specific fine-tuning further improves performance to 93.65\% F1 and 0.93 NED, demonstrating superior robustness against seals, watermarks, and low contrast. The framework offers an efficient, scalable solution for complex document understanding in office automation. Code is available at https://github.com/FairmeHIT/Multi-VIE, and fine-tuned models at https://huggingface.co/fairme/Qwen2.5-VL-7B-SFT.

cs.CV

Study of GeSn Selective Area Growth with Demonstration of SWIR Light Detection

As germanium-tin (GeSn) epitaxial growth quality continuously improves, the search for an efficient integration strategy of GeSn optoelectronics devices into complementary metal-oxide-semiconductor (CMOS) manufacturing line also accelerates. Selective area growth (SAG) on patterned substrate emerges as a promising approach for this quest, with locally controlled growth of GeSn laser/detector suitable for either co-integration with silicon-based waveguide structure or stand-alone module like focal plane array. In this work, we report successful GeSn SAG with Sn content ranging from 3.2% to 8.7% of good optical quality, with demonstration of tunable GeSn SAG photoluminescence and GeSn SAG photoconductor device, the latter with detection cutoff wavelength up to 2 um. In addition, we present a comprehensive study of GeSn SAG condition at different window sizes, from 2 um to 100 um, and shapes: circle, square, octagon, and rectangle. Presence of loading effect is revealed, where GeSn growth rate increases as pattern fill factor and window size shrink. It introduces a different growth condition compared to thin film growth, which can weaken or inhibit Sn incorporation at very small window size and induce Sn segregation in high Sn content SAG growth.

physics.app-ph

Joint constraints on gravity and stellar orbital anisotropy in massive galaxies

Strong gravitational lensing combined with stellar dynamics provides a complementary route for testing gravity on kiloparsec scales and probing the internal structure of massive galaxies. However, such studies remain limited by degeneracies among the mass-density profile, stellar orbital anisotropy and external convergence, and by modelling assumptions, especially when only single-aperture velocity dispersions are available. Here we develop a hierarchical Bayesian framework to disentangle gravity and stellar orbital anisotropy from other effects at the population level. By reconstructing the lens mass distribution with a flexible broken power-law model and propagating its posterior uncertainty into the predicted velocity dispersion, we obtain a likelihood for each lens in the plane of stellar orbital anisotropy and an effective mismatch parameter. This parameter encapsulates projection bias, external convergence, cosmological distance ratios and deviations from general relativity via the post-Newtonian parameter $\gamma_{\rm PPN}$. Applying this framework to 121 galaxy-scale lenses, we find $\gamma_{\rm PPN}=1.027^{+0.099}_{-0.095}$, consistent with general relativity, and obtain $2\sigma$ evidence that the stellar orbits of massive galaxies have become more radially biased over the past $\sim6$ Gyr. Forecasts show that future samples of order $10^5$ lenses could enable sub-percent tests of gravity, precise measurements of orbital-structure evolution and complementary constraints on the cosmological matter-density parameter.

astro-ph.GA

From Materials Database to Materials Bank: Assetizing Data for AI Driven Materials Innovation

Driven by high-throughput experimentation, computational modeling, and artificial intelligence (AI), materials data has expanded at an unprecedented rate. Conventional materials databases function only as passive repositories, archiving raw experimental records indiscriminately including both successful and failed data, without systematic value filtering or asset management. This creates a critical gap between massive data accumulation and actionable innovation, hindering the identification of high-potential materials and industrial translation. To address this bottleneck, we propose an industrialization-oriented Materials Bank, a dedicated valuefiltering and assetization layer that operates beyond traditional databases. It does not merely curate high-quality data but systematically elevates qualified candidates into standardized, upgradable materials assets via a multi-dimensional BankCard framework covering scientific validity, synthesis feasibility, application readiness, and industrial value. By unifying databases, AI models, automated experimentation, and multi-criteria assessment into a cohesive closed-loop ecosystem, the Materials Bank establishes a clear trajectory from data to knowledge, candidate, asset, and product. It serves not as an enhanced database or screening tool, but as a decision infrastructure bridging academic discovery and industrial demand, offering a scalable paradigm to accelerate AI-driven materials innovation and deliver tangible real-world impact.

cond-mat.mtrl-sci

Empowering Polymeric Materials Discovery by Artificial Intelligence

Polymeric materials underpin modern technologies spanning energy storage, microelectronics, healthcare and sustainable manufacturing. Yet their rational design remains exceptionally challenging because material performance emerges from complex interactions among molecular composition, chain architecture, processing history and hierarchical structural evolution across multiple length and time scales. Consequently, polymer research has long relied on labor-intensive experimentation and fragmented modeling approaches, limiting both mechanistic understanding and innovation efficiency. Recent advances in data infrastructure, machine learning, large artificial intelligence (AI) models and laboratory automation are beginning to reshape this landscape. Rather than functioning as isolated tools, polymer databases, predictive models, AI agents and automated laboratories are increasingly converging into interconnected discovery ecosystems. As a result, the central challenge is shifting from improving predictive accuracy alone to enabling reliable decision-making, adaptive learning and seamless integration across computation, experimentation and scientific reasoning. We argue that polymer science is entering an era of autonomous discovery, in which data, simulation, reasoning and experimentation operate within self-improving feedback loops that continuously generate hypotheses, design materials, execute experiments and refine predictive models. By unifying molecular design, process optimization, experimental validation and industrial translation, such autonomous ecosystems establish a more predictive, reproducible and scalable paradigm for polymer innovation, fundamentally transforming how polymer research is conducted.

physics.chem-ph

Surface Originated Cross-Field Anomalous Transport in Magnetoelectric Multilayers

In material systems with slab geometry, the surface contribution to physical responses is commonly expected to diminish rapidly with increasing thickness, giving way to the bulk response. Here, we show that this conventional wisdom is violated in a class of gate-induced responses, including gate-induced orbital and spin magnetization as well as cross-field anomalous thermoelectric transport. We develop a general framework for these effects, which naturally decomposes the total response into surface- and bulk-contributions treated on equal footing. Remarkably, the volume-averaged surface contribution remains finite in the thick-slab limit and exhibits the same thickness scaling as the bulk term. Furthermore, the surface response originates from band geometric quantities distinct from those in the bulk, being constrained solely by surface symmetries. As a result, it can dominate the overall response when the bulk contribution is symmetry-forbidden. Taking MnBi$_2$Te$_4$ multilayers as an example, we predict a strong surface-dominated cross-field anomalous Nernst effect arising from surface Berry curvature, which is readily accessible to experimental detection. These findings reveal a previously overlooked significance of surface response and open a new direction in the study of surface quantum geometry.

cond-mat.mes-hall

The FAST Hundred-Deg$^2$ HI Deep (HD$^2$) Survey: Early Results from the Pilot Survey

The Hundred-deg$^2$ HI Deep (HD$^2$) survey carried out with the Five-hundred-meter Aperture Spherical Telescope (FAST) is planned to map a contiguous region within the DESI DR1 footprint, achieving an effective integration time of 20 minutes for each pointing and a uniform detection sensitivity of 0.28 mJy beam$^{-1}$ at 4.8 km s$^{-1}$ resolution. We present early results from the pilot HD$^2$ survey: a 10 deg$^2$ field overlapping with HSC-SSP and the DESI EDR SV3, observed with an integration time of 7.3 minutes per beam and the rms of 0.45 mJy beam$^{-1}$ at 4.8 km s$^{-1}$ resolution. We identify 339 HI sources at $z<0.09$, corresponding to $\sim$34 detections per deg$^2$, nearly six times higher than the detection rate of the wide-field surveys. Optical counterparts are primarily identified using DESI redshifts, yielding a matching rate and correctness exceeding 90% for galaxies with $r<19.5$ mag, a substantial improvement over SDSS. Under the constraint of $r < 17.8$ mag and $0.01 < z < 0.05$, nearly 50% of galaxies in the DESI BGS samples have HI detections in this pilot survey. The optical properties of these HI-detected galaxies span nearly the entire parameter range of the DESI sample. The gas fraction scaling relations versus stellar mass, stellar mass surface density, NUV-r, and specific star formation rate are consistent with previous surveys, e.g., ALFALFA, DINGO, and xGASS. These results justify the feasibility of the full HD$^2$ survey, which will build a high-completeness HI census over a contiguous area to probe the cold gas scaling relations of galaxies over different scales.

astro-ph.GA

Comparative analysis of missing data imputation methods for CSST survey: Impact on photometric redshift estimation performance

Improving the accuracy of photometric redshifts (photo-$z$) is essential for reliable statistical studies of cosmology and galaxy evolution. However, missing photometric bands are a common observational challenge that can significantly degrade photo-$z$ estimation accuracy. In this work, we present a systematic evaluation of data imputation methods aimed at improving photo-$z$ performance. We benchmark a range of representative machine learning (ML) and deep learning (DL) architectures, identifying k-nearest neighbors (KNN) and the attention-based SAITS model as the leading performers. These models are then applied to China Space Station Survey Telescope (CSST) mock data to assess their performance under realistic observational conditions. Our results show that KNN yields the highest accuracy under idealized missing completely at random (MCAR) conditions with complete training sets, whereas robustness tests reveal that SAITS significantly outperforms KNN when training data is incomplete or when applied to realistic mixed-mechanism scenarios. We find that domain consistency between training and testing missingness patterns is a prerequisite for optimal performance, highlighting the risks of domain shift in supervised regression tasks. Furthermore, our analysis demonstrates that while general imputation models are highly effective for MCAR and missing at random (MAR) data, they are detrimental when applied to missing not at random (MNAR) data arising from flux limits, as statistical models fail to capture the physical information inherent in these non-detections. Consequently, we advocate for more sophisticated architectures capable of disentangling stochastic missingness from physical non-detections to address these distinct mechanisms individually.

astro-ph.GA

Germanium-tin (GeSn) avalanche photodiode with up to 2.7 micro cutoff wavelength for extended SWIR detection

Separate absorption charge multiplication germanium tin on silicon avalanche photodiode offers a viable solution to achieve CMOS compatible, high sensitivity detection technology in SWIR or extended SWIR range, leveraging the excellent k-factor of Si as multiplication layer and SWIR or e-SWIR band absorption of GeSn. However, unlike well-established growth of GeSn on Si with thick Ge buffer in-between to reduce threading dislocation density due to lattice mismatch, GeSn on Si APD design requires relatively thin Ge buffer to limit electric field drop through the background p-doped buffer and efficiently transporting photocarrier from GeSn absorber to Si multiplication layer, therefore making growth of high Sn content APD for e-SWIR coverage very challenging. In this work, we experimentally demonstrate GeSn on Si APD up to 12.7 percent Sn, monolithically grown on Si substrate with 122-nm-thick Ge buffer in between, which is considerably thinner than widely used 700-900 nm thick Ge buffer. Stronger relaxation of GeSn absorber via thin Ge buffer favors Sn incorporation, leading to higher Sn content than the nominal target of 8 percent Sn. Device detection range is significantly improved compared to previous work - with cutoff wavelength increased up to 2.7 micro at 300 K, in parallel with high avalanche gain at 77 K up to 21 at 1.55 micro and up to 52 at 2 micro, and good responsivity in SWIR or e-SWIR range, up to 1.45 AW-1 at 1.55 micro and 0.66 AW-1 at 2 micro.

physics.app-ph

EmbTracker: Traceable Black-box Watermarking for Federated Language Models

Federated Language Model (FedLM) allows a collaborative learning without sharing raw data, yet it introduces a critical vulnerability, as every untrustworthy client may leak the received functional model instance. Current watermarking schemes for FedLM often require white-box access and client-side cooperation, providing only group-level proof of ownership rather than individual traceability. We propose EmbTracker, a server-side, traceable black-box watermarking framework specifically designed for FedLMs. EmbTracker achieves black-box verifiability by embedding a backdoor-based watermark detectable through simple API queries. Client-level traceability is realized by injecting unique identity-specific watermarks into the model distributed to each client. In this way, a leaked model can be attributed to a specific culprit, ensuring robustness even against non-cooperative participants. Extensive experiments on various language and vision-language models demonstrate that EmbTracker achieves robust traceability with verification rates near 100\%, high resilience against removal attacks (fine-tuning, pruning, quantization), and negligible impact on primary task performance (typically within 1-2\%).

cs.CR

A Robust Geometric Distortion Solution for Main Survey Camera of CSST

The advancement in sensitivity and field of view of next-generation wide-field survey telescopes requires astrometric measurements with high precision, even in the presence of significant geometric distortions. To address this challenge, we develop a Weighted Polynomial Distortion Correction in 2-Phase (WPDC-2P) method. This approach enhances stellar cross-matching, incorporates distance-based weighting into the traditional polynomial fitting, and employs a look-up table to absorb the remaining distortion residuals. Validated on simulated data from the Main Survey Camera of the \emph{Chinese Space Station Survey Telescope} (CSST), incorporating geometric distortions up to approximately $200$ pixels, the method achieves astrometric standard deviation ranging from 0.013 to 0.107 pixels (0.03 pixels for the $g$-1 detector) across all 18 detectors. Under extreme crowding conditions (e.g., globular cluster NGC 2298), the astrometric precision for the $g$-1 detector reaches 0.05-pixel level within the central region ($r_d < 4000$), despite a centroiding precision of $\sim$0.04 pixels. When applied to the Beijing-Arizona Sky Survey data, for which the standard pipeline delivers an astrometric uncertainty of $\sim$20 mas, our method reduces the positional scatter to $ \sigma_{\Delta\alpha}=5.494$ mas (0.01 pixels) and $ \sigma_{\Delta\delta}=9.981$ mas (0.02 pixels) using only a weighted 3rd-order polynomial correction. The method has been integrated into the CSST data processing pipeline and is prepared for further refinement using on-orbit calibration data.

astro-ph.IM

ProtegoFed: Backdoor-Free Federated Instruction Tuning with Interspersed Poisoned Data

Federated Instruction Tuning (FIT) enables collaborative instruction tuning of large language models across multiple organizations (clients) in a cross-silo setting without requiring the sharing of private instructions. Recent findings on natural backdoors and the existing training data collection method suggest that poisoned samples may be pervasive and inadvertently embedded in real-world datasets, potentially distributed across all clients, even if the clients are benign. This work systematically examine this threat in FIT, demonstrating that existing defenses are ineffective when poisoned data is interspersed among all clients. Addressing this challenge entails two major difficulties: identifying the distinctive characteristics of poisoned samples at each client and enabling collaborative defense when some clients are heavily dominated by poisoned samples. To address these difficulties, we identify gradients in the frequency domain as a robust signal to distinguish poisoned data. We further propose a global secondary clustering mechanism that facilitates collaborative identification of poisoned samples across clients. In summary, this paper introduces ProtegoFed, the first backdoor-free FIT framework that accurately detects, removes, and even purifies interspersed poisoned data across clients during the training. Experimental results on four FL datasets show that ProtegoFed identifies $92.00\% \sim 100.00\%$ of poisoned samples, reduces the attack success rate to almost zero, and maintains utility on the main task. Code is available at https://github.com/dongdongzhaoUP/ProtegoFed.

cs.CR

Systematic study of high performance GeSn photodiodes with thick absorber for SWIR and extended SWIR detection

Germanium-tin (GeSn) photodiodes potentiate a viable solution to integrate SWIR and extended SWIR detection technology into CMOS processing line. However, challenges in the growth of thick, high quality GeSn limit the device absorber thickness, making it impossible to ascertain the performance limit of GeSn photodiodes. An in-depth understanding of their device physics and a clear optimization pathway towards commercial-grade devices remain elusive. This work presents a systematic empirical study of GeSn photodiodes with thick absorber (2 to 8% Sn content, up to 2630 nm thick), showing high responsivity up to 0.59 A.W-1 at 1.55 {\mu}m and 0.43 A.W-1 at 2 {\mu}m wavelengths, low dark current density down to 2 x 10-2 A.cm-2, and high detection cutoff wavelengths up to 2.1 and 2.5 {\mu}m at 5% and 8% Sn, respectively. Using specific doping design (P-i-N and N-i-P), an in-depth analysis is presented on the impact of junction position, p-type background carrier concentration, bulk/ surface defects and photocarrier diffusion length - on photodetection performance. Different optimization strategies for GeSn photodiodes, in particular at high Sn content, are proposed.

physics.ins-det

CFHT MegaCam Two Deep Fields Imaging Survey (2DFIS) II: Decoding the Lensing Profile of a "Rotating" Cluster with Deep CFHT Imaging

We present a multi-wavelength analysis of the galaxy cluster RXCJ0110.0+1358 ($z=0.058$), a rotating cluster candidate, combining deep CFHT imaging, SDSS photometry, spectroscopic redshifts, and XMM-Newton X-ray observations. We find a notable discrepancy between the optical and X-ray views: while optical data reveal a pronounced bimodal galaxy distribution with significant kinematic substructure signatures, the X-ray emission exhibits a single, smoothly extended component centered on the BCG. Our weak lensing analysis resolves this discrepancy by revealing that the mass is predominantly concentrated in the southeast ($\log M_{200}/M_\odot = 14.04_{-0.40}^{+0.24}$), while the northwestern substructure has a negligible mass ($\sim 10^{13} M_\odot$). This immense mass disparity rules out the dynamical possibility of a rotating system. We demonstrate that the apparent optical bimodality arises from the projection of a filament, which led optical group-finding algorithms to misclassify these galaxies as cluster members. This contamination creates a spurious substructure that mimics a rotation signal and leads to an overestimation of the luminosity-based halo mass, resolving the observed inconsistencies.

astro-ph.GA