Searcharxiv⌕ Search

arXiv subjects

Tao Han

Publications and source records attributed to Tao Han.

At least 37 records · Page 2Linked to original sources

Graph Attention-Based Virtual Metrology for Film Deposition Processes in Semiconductor Manufacturing

Artificial intelligence-driven semiconductor manufacturing increasingly operates at nanometer and angstrom scales, where precise process control depends on accurate and timely metrology. However, physical metrology is limited by measurement latency, cost, and sampling constraints, restricting its scalability in high-volume production. Virtual metrology (VM) has emerged as an effective alternative by predicting wafer-level characteristics from equipment sensor data. Despite recent advances, many existing VM models remain correlation-driven and lack the ability to capture structured dependencies among heterogeneous process variables, while providing limited interpretability. This study presents a graph attention-based VM framework for film deposition processes that integrates temporal feature learning with structured parameter-layer dependency modeling. The proposed approach represents each step-parameter pair as a node and extracts temporal embeddings from high-frequency equipment traces using convolutional feature encoders. A parameter-to-layer graph attention mechanism is employed to model directional dependencies, enabling each film layer to aggregate relevant process information. The framework is evaluated using industrial deposition data collected from production wafers, where the model predicts film thickness from multivariate sensor signals. Experimental results demonstrate improved predictive performance compared to baseline models. In addition, analysis of the learned attention weights reveals interpretable parameter-layer relationships consistent with physical process behavior, capturing dominant process factors and temporal dependencies across deposition stages. These results indicate that the proposed framework enhances prediction accuracy and provides meaningful insight into process dynamics, supporting effective monitoring and optimization in semiconductor manufacturing.

cs.CE↗

Train the Agent, Not the Expert: Learning to Harness Heterogeneous Experts for Multi-Turn Visual Reasoning

Recent progress in computer vision has produced a wide range of powerful specialized models for detection, segmentation, counting, and other visual tasks. However, these models are usually optimized for isolated task formulations, making it difficult to directly support general-purpose visual intelligence, especially when a task requires complex language understanding and dense small-object perception. In this paper, we propose VisHarness, a trainable visual agent that decouples high-level perception, reasoning, and decision-making from low-level task execution. Instead of training a model to solve a specific visual task, VisHarness learns to harness a set of carefully designed heterogeneous visual experts. This paradigm preserves the general intelligence of the agent while fully leveraging the precision advantages of specialized visual models in concrete visual tasks. With only lightweight training, VisHarness learns a generalizable visual expert-harnessing policy and can solve common fundamental vision tasks under various complex conditions through multi-turn interactions with visual expert models. To enable efficient on-policy reinforcement learning training in a live environment, we introduce dynamic visual memory archiving, which mitigates the rapidly accumulating visual-token overhead caused by multi-turn interactions with visual expert models. Experiments on four representative benchmarks covering reasoning segmentation, generalized referring segmentation, dense small-object detection, and referring counting demonstrate that VisHarness substantially outperforms existing general-purpose models and achieves competitive or superior performance compared with task-specific models.

cs.CV↗

RealBench: Benchmarking Data-Driven Numerical Weather Forecasting Under Operational Conditions and Extreme Event Challenges

Accurate evaluation of weather forecasting models is critical for their reliable deployment in real-world applications. However, existing benchmarks predominantly rely on reanalysis products such as ERA5, which are generated through delayed data assimilation and do not reflect the constraints of real-time operational forecasting, thereby resulting in a systematic mismatch between benchmark performance and real-world forecasting. In this work, we introduce RealBench, a next-generation benchmark for AI weather forecasting that emphasizes realistic evaluation under operational conditions. RealBench features a strictly out-of-distribution test set spanning 2025 to eliminate data leakage and capture recent atmospheric regimes. It integrates multiple data sources, including low-latency operational analysis and a large-scale global in-situ observation dataset comprising over 10,000 stations, enabling direct evaluation against real atmospheric measurements. Beyond standard global metrics, RealBench provides a comprehensive evaluation framework for high-impact extreme events, including heatwaves, cold surges, and tropical cyclones, using event-specific metrics that better reflect real-world forecasting priorities. The evaluation results reveal substantial discrepancies between reanalysis-based metrics and real-world performance, particularly concerning extreme events. By highlighting the limitations of existing benchmarks, this work establishes a more faithful and operationally relevant evaluation paradigm, providing a rigorous foundation for advancing next-generation AI weather forecasting systems. The benchmark implementation is available at: https://github.com/lixruize-del/NWP-Benchmark.

cs.LG↗

EMFormer: Efficient Multi-Scale Transformer for Accumulative Context Weather Forecasting

Long-term weather forecasting is critical for socioeconomic planning and disaster preparedness. While recent approaches employ finetuning to extend prediction horizons, they remain constrained by the issues of catastrophic forgetting, error accumulation, and high training overhead. To address these limitations, we present a novel pipeline across pretraining, finetuning and forecasting to enhance long-context modeling while reducing computational overhead. First, we introduce an Efficient Multi-scale Transformer (EMFormer) to extract multi-scale features through a single convolution in both training and inference. Based on the new architecture, we further employ an accumulative context finetuning to improve temporal consistency without degrading short-term accuracy. Additionally, we propose a composite loss that dynamically balances different terms via a sinusoidal weighting, thereby adaptively guiding the optimization trajectory throughout pretraining and finetuning. Experiments show that our approach achieves strong performance in weather forecasting and extreme event prediction, substantially improving long-term forecast accuracy. Moreover, EMFormer demonstrates strong generalization on vision benchmarks (ImageNet-1K and ADE20K) while delivering a 5.69x speedup over conventional multi-scale modules. Code: https://github.com/chenhao-zju/emformer

cs.CV↗

EdgeServing: Deadline-Aware Multi-DNN Serving at the Edge

As edge computing expands, serving multiple deep neural network (DNN) models on a single shared GPU has become a common yet challenging scenario, where each scheduling decision affects the tail latency of all concurrent queues. Existing schedulers rely on local heuristics and fail to capture this global impact, while GPU spatial-sharing approaches sacrifice latency predictability. In this paper, we propose EdgeServing, a deadline-aware multi-DNN serving system for edge devices. EdgeServing adopts time-division GPU sharing with early-exit inference for high inference predictability, and introduces a stability score to quantify how each candidate scheduling decision impacts the future queue status. At runtime, it cohesively selects the model, exit point, and batch size to minimize predicted system-wide SLO impact. Experimental results on multiple hardware platforms show that EdgeServing consistently outperforms representative baselines in both SLO violation ratio and P95 latency, enabled by early-exit mechanism, which expands the scheduling action space under tight latency constraints.

cs.DC↗

Earth-o1: A Grid-free Observation-native Atmospheric World Model

Despite the unprecedented volume of multimodal data provided by modern Earth observation systems, our ability to model atmospheric dynamics remains constrained. Traditional modeling frameworks force heterogeneous measurements into predefined spatial grids, inherently limiting the full exploitation of raw sensor data and creating severe computational bottlenecks. Here we present Earth-o1, an observation-native atmospheric world model that overcomes these structural limitations. Rather than relying on conventional atmospheric dynamical modeling systems or traditional data assimilation, Earth-o1 directly learns the continuous, three-dimensional physical evolution of the Earth system from ungridded observational data. By integrating diverse sensor inputs into a unified, grid-free dynamical field, the model autonomously advances the atmospheric state in space and time. We show that this fundamentally distinct paradigm enables direct, real-time forecasting and cross-sensor inference without the overhead of explicit numerical solvers. In hindcast evaluations, Earth-o1 achieves surface forecast skill comparable to the operational Integrated Forecasting System (IFS). These results establish that continuous, observation-driven world models -- a new class of fully observation-native geophysical simulators -- can match the fidelity of established physical frameworks, providing a scalable data-driven foundation for a digital twin of the Earth.

cs.CV↗

What You Think is What You See: Driving Exploration in VLM Agents via Visual-Linguistic Curiosity

To navigate partially observable visual environments, recent VLM agents increasingly internalize world modeling capabilities into their policies via explicit CoT reasoning, enabling them to mentally simulate futures before acting. However, relying solely on passive reasoning over visited states is insufficient for sparse-reward tasks, as it lacks the epistemic drive to actively uncover the ``known unknown'' required for robust generalization. We ask: Can VLM agents actively find signals that challenge and refine their internal world model through curiosity-driven exploration? In this work, we propose GLANCE, a unified framework that bridges reasoning and exploration by grounding the agent's linguistic world model into the stable visual representations of an evolving target network. Crucially, GLANCE leverages the discrepancy between linguistic prediction and visual reality as an intrinsic curiosity signal within reinforcement learning, steering the agent to actively explore areas where its internal model is uncertain. Extensive experiments across a series of agentic tasks show the effectiveness of GLANCE, and demonstrate that aligning ``what the agent thinks'' with ``what the agent sees'' is key to solving complex or sparse agentic tasks.

cs.AI↗

STCast: Adaptive Boundary Alignment for Global and Regional Weather Forecasting

To gain finer regional forecasts, many works have explored the regional integration from the global atmosphere, e.g., by solving boundary equations in physics-based methods or cropping regions from global forecasts in data-driven methods. However, the effectiveness of these methods is often constrained by static and imprecise regional boundaries, resulting in poor generalization ability. To address this issue, we propose Spatial-Temporal Weather Forecasting (STCast), a novel AI-driven framework for adaptive regional boundary optimization and dynamic monthly forecast allocation. Specifically, our approach employs a Spatial-Aligned Attention (SAA) mechanism, which aligns global and regional spatial distributions to initialize boundaries and adaptively refines them based on attention-derived alignment patterns. Furthermore, we design a Temporal Mixture-of-Experts (TMoE) module, where atmospheric variables from distinct months are dynamically routed to specialized experts using a discrete Gaussian distribution, enhancing the model's ability to capture temporal patterns. Beyond global and regional forecasting, we evaluate our STCast on extreme event prediction and ensemble forecasting. Experimental results demonstrate consistent superiority over state-of-the-art methods across all four tasks. Code: https://github.com/chenhao-zju/STCast

cs.LG↗

Spin Correlation and Quantum Entanglement of Fermion Pairs in Transversely Polarized $e^-e^+$ Collisions

We systematically study the spin correlations and quantum entanglement in transversely polarized electron-positron collisions. We find that the $s$-channel QED process $e^-e^+\to f\bar f$ produces a maximally entangled state in the entire phase space when the initial beams are transversely polarized, while the quantum magic varies in different phase space points for the maximally entangled Bell states. For electroweak processes, the spin configuration of final states depends on chiral couplings, and the entanglement is also greatly enhanced by transverse polarization as in the QED process. For Bhabha scattering with additional $t$-channel contributions, the transverse polarization still increases the final state entanglement, although with some dilution. The sensitive dependence of final spin states on the transverse polarization makes the beam polarization a powerful tool for generating and controlling quantum entanglement in collider experiments, opening up new opportunities for quantum information studies at high-energy colliders.

hep-ph↗

Generative 3D Gaussian Splatting for Arbitrary-ResolutionAtmospheric Downscaling and Forecasting

While AI-based numerical weather prediction (NWP) enables rapid forecasting, generating high-resolution outputs remains computationally demanding due to limited multi-scale adaptability and inefficient data representations. We propose the 3D Gaussian splatting-based scale-aware vision transformer (GSSA-ViT), a novel framework for arbitrary-resolution forecasting and flexible downscaling of high-dimensional atmospheric fields. Specifically, latitude-longitude grid points are treated as centers of 3D Gaussians. A generative 3D Gaussian prediction scheme is introduced to estimate key parameters, including covariance, attributes, and opacity, for unseen samples, improving generalization and mitigating overfitting. In addition, a scale-aware attention module is designed to capture cross-scale dependencies, enabling the model to effectively integrate information across varying downscaling ratios and support continuous resolution adaptation. To our knowledge, this is the first NWP approach that combines generative 3D Gaussian modeling with scale-aware attention for unified multi-scale prediction. Experiments on ERA5 show that the proposed method accurately forecasts 87 atmospheric variables at arbitrary resolutions, while evaluations on ERA5 and CMIP6 demonstrate its superior performance in downscaling tasks. The proposed framework provides an efficient and scalable solution for high-resolution, multi-scale atmospheric prediction and downscaling. Code is available at: https://github.com/binbin2xs/weather-GS.

cs.CV↗

Collider probes of baryogenesis with maximal CP asymmetry

We propose a novel collider probe of baryogenesis at TeV scale by measuring decay asymmetries into particle and anti-particle final states. Motivated by the idea of Dirac leptogenesis, we consider an extension of the standard model with new colored and $SU(2)_L$ singlet particles in such a way that the out-of-equilibrium decay of heavy colored fermions creates equal and opposite CP asymmetries in two sectors, prevented from equilibrating with each other. While the TeV scale viability of this mechanism requires a resonantly enhanced CP asymmetry, the latter also plays a crucial role leading to observable decay asymmetries in colliders. In addition to discussing conventional signatures of such heavy colored particles, namely, mono-jet plus missing transverse energy, displaced vertex, colored track at hadron colliders, we also show the unique possibility of measuring decay asymmetries via forward-backward and charge asymmetries at future muon colliders. In addition to being a verifiable TeV-scale baryogenesis scenario, the model also predicts a singlet scalar dark matter candidate consistent with the required thermal dark matter properties near the Higgs resonance.

hep-ph↗

UNICBench: UNIfied Counting Benchmark for MLLM

Counting is a core capability for multimodal large language models (MLLMs), yet there is no unified counting dataset to rigorously evaluate this ability across image, text, and audio. We present UNICBench, a unified multimodal, multi level counting benchmark and evaluation toolkit with accurate ground truth, deterministic numeric parsing, and stratified reporting. The corpus comprises 5,300 images (5,508 QA), 872 documents (5,888 QA), and 2,069 audio clips (2,905 QA), annotated with a three level capability taxonomy and difficulty tags. Under a standardized protocol with fixed splits/prompts/seeds and modality specific matching rules, we evaluate 45 state-of-the-art MLLMs across modalities. Results show strong performance on some basic counting tasks but significant gaps on reasoning and the hardest partitions, highlighting long-tail errors and substantial headroom for improving general counting. UNICBench offers a rigorous and comparable basis for measurement and a public toolkit to accelerate progress.

cs.CV↗

Benchmarking AI-based data assimilation to advance data-driven global weather forecasting

Research on Artificial Intelligence (AI)-based Data Assimilation (DA) is expanding rapidly. However, the absence of an objective, comprehensive, and real-world benchmark hinders the fair comparison of diverse methods. Here, we introduce DABench, a benchmark designed for contributing to the development and evaluation of AI-based DA methods. By integrating real-world observations, DABench provides an objective and fair platform for validating long-term closed-loop DA cycles, supporting both deterministic and ensemble configurations. Furthermore, we assess the efficacy of AI-based DA in generating initial conditions for the advanced AI-based weather forecasting model to produce accurate medium-range global weather forecasting. Our dual-validation, utilizing both reanalysis data and independent radiosonde observations, demonstrates that AI-based DA achieves performance competitive with state-of-the-art AI-driven four-dimensional variational frameworks across both global weather DA and medium-range forecasting metrics. We invite the research community to utilize DABench to accelerate the advancement of AI-based DA for global weather forecasting.

cs.LG↗

Quantum Tomography of Fermion Pairs in $e^+e^-$ Collisions: Longitudinal Beam Polarization Effects

We present a quantum tomography study of fermion pair production at future $e^+e^-$ colliders, emphasizing how longitudinal beam polarization controls the two-qubit spin density matrix. We study the processes $e^+ e^- \to t\bar{t},\ e^+e^-\to μ^+μ^-$ and Bhabha scattering $e^+e^-\to e^+e^-$, representing the mass threshold behavior, the $Z$ pole resonance and the $s/t$-channel interplay. We choose to focus on three key concepts: quantum entanglement via the concurrence $\mathcal{C}$, Bell nonlocality via the optimal Clauser Horne Shimony Holt (CHSH) parameter $\mathcal{B}$, and non-stabilizerness (``magic'') via the second stabilizer Rényi entropy $\mathcal{M}_2$. For the $s$-channel-dominated channels, longitudinal polarization mainly reshapes single-spin polarizations while leaving the spin-correlation matrix largely unchanged, rendering $\mathcal{C}$ and $\mathcal{B}$ comparatively robust, but inducing a pronounced variation of $\mathcal{M}_2$. In contrast, in Bhabha scattering, polarization modifies the relative contributions of the $s$-channel and $t$-channel and can strongly affect all three observables. The observability of entanglement, Bell nonlocality, and magic exceeds the $5σ$ level when both statistical and systematic uncertainties are included, establishing the fermion pair systems as ideal laboratories for quantum-information studies in high energy leptonic collisions. With optimized beam polarization, future $e^+e^-$ colliders will provide a unique opportunity to experimentally explore and influence quantum resources in particle interactions.

hep-ph↗

Drell-Yan Production of New Particles at Fixed-Target Experiments: Heavy Neutral Lepton as a Case Study

We demonstrate the sensitivity of Drell-Yan production processes from deep inelastic scattering in searches for beyond-the-Standard Model (BSM) physics at fixed-target or beam-bump experiments. We take heavy neutral leptons (HNLs) as a case study, produced from the decay of a light vector boson mediator with mass in the range of $2-20$ GeV, which itself is generated via the Drell-Yan process. The produced HNLs subsequently decay into Standard Model final states. We consider several current and future experiments, including SBND, DarkQuest, DUNE Near Detector (ND), and SHiP. Utilizing $νπ^0$ and $νe^+e^-$ final states from HNL decays, we find that the Drell-Yan mechanism provides important contributions and significantly enhances the HNL search sensitivity, owing to the production of energetic final-state particles that are more readily detectable over the expected backgrounds. We find that at $90\%$ C.L. sensitivity, for gauge couplings $g_{X} \sim 10^{-2}\ (10^{-3})$ and kinematically accessible mass range, SBND and DarkQuest can probe the HNL flavor mixing $|U_{\ell}| \sim 3\times 10^{-4}\ (10^{-3})$, whereas DUNE ND and SHiP may extend the sensitivity down to the Type-I Seesaw prediction of $|U_{\ell}| \sim 10^{-5}$. Finally, for our chosen benchmark $|U_{\ell}| = 10^{-3}$ outside of the current experimental constraints, with a fixed mass ratio $m_{Z'}/m_N = 2.1$, and working within the $U(1)_{B-L}$, $U(1)_{B-3L_τ}$, and $U(1)_{B}$ parameter spaces, we find that both SBND and DarkQuest can probe $g_{X} \sim 10^{-3}$, DUNE ND can reach $g_{X} \sim 10^{-4}$, and SHiP can probe down to $g_{X}\sim 5\times 10^{-6}$. Our approach provides a powerful new technique to study HNL production at future fixed-target experiments and can readily be extended to other light BSM particle production within a broader class of dark sector models.

hep-ph↗

Multi-messenger standard-siren cosmology for third-generation gravitational-wave detectors: forecasts considering observations of gamma-ray bursts and kilonovae

In the third-generation (3G) gravitational-wave (GW) detector era, GW multi-messenger observations for binary neutron star merger events can exert great impacts on exploring the cosmic expansion history. Extending the previous work, we explore the potential of 3G GW standard siren observations in cosmological parameter estimation by considering their associated electromagnetic (EM) counterparts, including $γ$-ray burst (GRB) coincidence observations by the Gravitational wave high-energy Electromagnetic Counterpart All-sky Monitor and GW-triggered target-of-opportunity observations of kilonovae by different optical survey projects. During an assumed 10-year observation, we predict that the number of detectable GW-kilonova events is $\sim 4900$ with redshifts below $\sim 0.4$ under GW network and Large Synoptic Survey Telescope in the $i$ band, which is three times more than that of GW-GRB detections. For the cosmological analysis, we find that with the inclusion of GW-kilonova detections, the constraints on cosmological parameters from GW-EM detections are significantly improved compared to those from GW-GRB detections. In particular, GW-EM detections can tightly constrain the Hubble constant with a precision ranging from $0.076\%$ to $0.034\%$. Moreover, GW multi-messenger observations could effectively break the cosmological parameter degeneracies generated by the mainstream EM observations, CMB+BAO+SN (CBS). The combination of CBS and GW-EM can tightly constrain the equation of state parameters of dark energy $w$ in the $w$CDM model and $w_0$ in the $w_0w_a$CDM model with precisions of $0.72\%$ and $0.99\%$, respectively, meeting the standard of precision cosmology. In conclusion, GW multi-messenger observations could play a crucial role in helping solve the Hubble tension and probing the fundamental nature of dark energy.

astro-ph.CO↗

R-Log: Incentivizing Log Analysis Capability in LLMs via Reasoning-based Reinforcement Learning

The growing complexity of log data in modern software systems has prompted the use of Large Language Models (LLMs) for automated log analysis. Current approaches typically rely on direct supervised fine-tuning (SFT) on log-label pairs. However, this exacerbates the domain discrepancy between general-purpose LLMs and specialized log data, causing overfitting. Furthermore, SFT's imbalanced loss computation often allows lengthy contexts to overwhelm critical, concise details in model answers, leading to hallucinations. To address these limitations, we propose R-Log, a novel reasoning-based paradigm that mirrors the structured, step-by-step analytical process of human engineers. This approach enhances generalizability by learning the underlying rules behind conclusions. We further employ Reinforcement Learning (RL) to optimize the model within a simulated O&M environment, thereby reducing hallucinations by directly rewarding correct outcomes. R-Log is first cold-started on a curated dataset of 2k+ reasoning trajectories, guided by 13 strategies from manual O&M practices, to establish an initial reasoning capability. This ability is then refined via RL using a joint reward function. Empirical evaluations on real-world logs show that R-Log outperforms existing methods across five log analysis tasks, particularly in unseen scenarios (by 228.05%). We also designed R-Log-fast with 5x speedup while keeping 93% of the efficacy.

cs.SE↗

Towards an end-to-end artificial intelligence driven global weather forecasting system

The weather forecasting system is important for science and society, and significant achievements have been made in applying artificial intelligence (AI) to medium-range weather forecasting. However, existing AI-based weather forecasting models rely on analysis or reanalysis products from traditional numerical weather prediction (NWP) systems as initial conditions for making predictions. The initial states are typically generated by traditional data assimilation components, which are computationally expensive and time-consuming. Here, by cyclic training to model the steady-state background error covariance and introducing the confidence matrix to characterize the quality of observations, we present an AI-based data assimilation model, i.e., Adas, for global weather variables. Further, we combine Adas with the advanced AI-based forecasting model (i.e., FengWu) to construct an end-to-end AI-based global weather forecasting system: FengWu-Adas. We demonstrate that Adas can assimilate global conventional observations to produce high-quality analysis, enabling the system to operate stably for long term. Moreover, the system can generate accurate end-to-end weather forecasts with comparable skill to those of the IFS, demonstrating the promising potential of data-driven approaches.

physics.ao-ph↗