SearcharxivSearch

arXiv subjects

Wei Luo

Publications and source records attributed to Wei Luo.

At least 19 recordsLinked to original sources

$g$-wave altermagnetic order parameter in hematite

Altermagnets combine the vanishing net magnetization of antiferromagnets with momentum-dependent spin splitting. Magnon band splitting provides a direct probe of altermagnetic order and may enable chirality-selective magnon transport, yet the momentum-space symmetry of this splitting has not been determined quantitatively. Here we use inelastic neutron scattering to map the momentum dependence of altermagnetic magnon splitting in hematite ($\alpha$-Fe$_2$O$_3$). The splitting vanishes along nodal directions and reaches maxima off the nodes, revealing the $g$-wave symmetry of the altermagnetic order parameter. These results agree with linear spin-wave theory calculations based on the altermagnetic model, which further identify the nondegenerate branches as magnons of opposite chirality and trace the splitting to symmetry-inequivalent long-range exchange interactions. Our results provide the first quantitative determination of the momentum-space symmetry of altermagnetic chiral magnons. These findings, together with hematite's high magnetic ordering temperature and low magnon damping, establish it as a promising platform for low-dissipation, symmetry-selective magnonic applications.

cond-mat.str-el

A Survey on Self-Improving Test-Time Intelligence: Feedback-Driven Adapting, Learning, and Scaling at Inference

The ability of AI systems to improve their behavior during deployment is becoming increasingly important. As inference moves beyond the static execution of a fixed trained model, a growing body of work studies how models can refine their behavior on the fly by exploiting test-time information and additional computation. These developments have largely evolved along two directions: methods that modify the model's state using test-time signals, and methods that improve predictions through extra inference-time resources such as more sampling and tool use. However, these directions are often studied in separate communities with different terminology, making their connections harder to see. In this survey, we present feedback-driven Test-Time Intelligence (TTI) as a unified perspective for understanding such deployment-time improvement. We use this view to relate test-time adaptation, test-time learning, and test-time scaling, highlighting both their distinctions and their growing overlap in hybrid systems. This unified framework helps connect previously fragmented ideas and provides a clearer conceptual foundation for studying inference-time self-improvement. We review major methodological paradigms, representative applications, and open challenges across vision, language, multimodal learning, generative models, robotics, and healthcare. Our goal is to provide a coherent foundation and research roadmap for the study of self-improving AI systems at test time.

cs.LG

Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills

Enterprise agent programs are moving from prototypes into production, where reusable skills, tools, and workflow packages must be reviewed with evidence rather than prose. Current gates often scan these artifacts for structure, style, and security, but they do not answer the deployment question: does the capability package help a live agent complete enterprise tasks under the same model, sandbox, and grading policy? We present ACES (Agentic Continuous Evaluation of Skills), a repository-native framework for evaluating skills and product capability packages as executable agent artifacts. ACES runs paired live trials with and without a target skill, normalizes trajectories into the Agent Trajectory Interchange Format (ATIF), grades six default runtime metrics, and reports Skill Lift: the target skill's added value for a fixed task, harness, workspace, and scorer. The same protocol supports product-owned task suites that compare baseline, skill, bundle, team-skill, and plugin targets. On 145 real skills from internal enterprise repositories and public catalogs, scan-only gates surface useful authoring issues but measure complementary facets (structural versus LLM-judge Spearman $\rho = 0.14$). Across 947 scored paired cases from 58 of 64 production skills and four primary harnesses, mean composite Skill Lift is 0.2134 (95\% paired-case CI [0.1967, 0.2301]); mean outcome-only lift, the average of accuracy and goal accuracy, is 0.1799. Composite lift is positive in 72.8\% of paired cases. The largest process-metric gains appear in skill execution, behavior check, and skill efficiency---signals about discovery, routing, workflow following, and tool use that document scans cannot observe. An open-source implementation of the methodology is available in NVIDIA SkillEvaluator.

cs.AI

IntHQ: Task-Interactive Hierarchical Query on Dual-Stream Representations for Generative Recommendation

Multi-task learning over heterogeneous data is fundamental to modern recommendation, while generative models are emerging as the backbone of next-generation recommenders. However, the integration of multi-task learning into the generative paradigm remains largely unexplored. Existing multi-task recommenders, in both discriminative and generative paradigms, extract task-relevant features from a single task-agnostic representation and wire tasks into a predefined conversion funnel. We show that this scheme is inherently prone to a threefold collapse. Source collapse, where task-specific signals are injected late and diluted in the shared latent space. Relational collapse, where task dependencies are either implicitly absorbed by the backbone or statically fixed by predefined funnels. Hierarchical collapse, where tasks depend on features at different scales and shift across training stages. We propose IntHQ, a multi-task generative recommender with three components, each alleviating one collapse. Dual-Stream Decoupling (DSD) injects task identity into computation stream early and separates the shared context stream from the task-specific stream, alleviating signal dilution. Task-Interactive Modeling (TIM) replaces the predefined funnel with explicit cross-task interaction, letting each task condition on the realized outcomes of its predecessors with learned, input-adaptive strength. Hierarchical Querying (HQ) lets each task gather multi-scale information across different layers at different training stages. In offline evaluations, IntHQ consistently outperforms competitive encoder backbones under four representative task-head configurations. Deployed in production on Amap, serving hundreds of millions of users for travel recommendation, IntHQ yields a 1.60\% relative UVCTR lift.

cs.IR

WearWow: Native 2K Multi-Garment Virtual Try-On via Adaptive Token Packing and Preference Alignment

Synthesizing native 2K multi-garment virtual try-on is a formidable frontier in digital fashion, critically bottlenecked by two fundamental limitations: the O(N^2) memory explosion induced by 2k conditions, and the spectral bias of diffusion models that over-smooths high-frequency fabric details. We present WearWow, an end-to-end, mask-free generative framework that pioneers ultra-high-resolution multi-garment synthesis. To mitigate the memory explosion , we propose Adaptive 2D Token Packing (ATP). ATP leverages inherent garment sparsity to algorithmically pack heterogeneous items onto a unified 2D canvas and prune uninformative background tokens, minimizing the effective sequence length and subsequent memory overhead while rigorously preserving 2D spatial priors. To rectify texture degradation, we introduce the Multi-dimensional Try-on Reward (MTR) system. MTR synergizes a Semantic Guidance Reward to explicitly drive tactile restoration with a Cloth Distribution Reward to implicitly anchor the physical distribution, a joint formulation that effectively mitigates the severe reward hacking. Furthermore, we curate WearWow-2K, an extreme-quality dataset comprising native 2K triplets, providing physically correct spatial interactions that naturally empower the model's mask-free generation. Extensive experiments demonstrate that WearWow establishes a new state-of-the-art, exceeding existing commercial baselines in native 2K multi-garment synthesis.

cs.CV

Tellurium sublattice instability driven amorphization in the chalcogenide AgSbTe2 under pressure

Pressure provides a powerful thermodynamic route to access hidden structural states in functional materials, yet the microscopic origin of pressure-induced amorphization remains elusive in many complex chalcogenides. Here we report a detailed high-pressure structural study of AgSbTe2,combining synchrotron X-ray diffraction with density functional theory and molecular dynamics calculations up to 60 GPa. We uncover a pressure-driven transformation from the ambient R-3m phase to a fully disordered cubic Im3m phase, through an extended intermediate amorphous state. Enthalpy calculations reveal a near-degeneracy between the R3m and Im3m structures over a broad pressure range, dictating amorphization. Contrary to previously speculated cation vacancies, the amorphization is governed by a pronounced displacement instability of the Te sublattice. Remarkably, the time dependent decompression pathway controls the final structural state, resulting in either amorphous (slow decompression) or fully crystalline (fast decompression) states, indicative of a strong counterintuitive kinetic effect.

cond-mat.mtrl-sci

On the Critical One Components Regularity for the $3-D$ Navier-Stokes System in $L^p_T(\dot{B}^{\frac 1 2+\frac 2 p}_{2,\infty})$ spaces

We consider the conditional regularity of the mild solution $v$ of the $3-D$ incompressible Navier-Stokes equations with initial data $v_0\in \dot{H}^{\frac 1 2}$ and vorticity $\Omega_0\in L^{r_0}$ for some $r_0\in (1,2)$. We prove that if the solution associated with initial data $v_0$ blows up at a finite time $T^\ast$, then for any $2<p<\infty$, and any unit vectors $e$ in $\mathbb{R}^3$, the integral $$\int_0^{T^\ast}\left\Vert (v(t)|e)_{\mathbb{R}^3}\right\Vert_{\dot{B}^{\frac 1 2+\frac 2 p}_{2,\infty}}^p{\rm d}t$$ blows up at $T^\ast$. The conclusion improves the recent results in Chemin et al. (Arch Ration Mech Anal 224(3):871-905, 2017) and Han et al. (Arch. Rational Mech. Anal. 231:939-970, 2019).

math.AP

Invariant Measure of the Camassa-Holm Equation with Linear Multiplicative Noise

In this paper, we prove that the solution map of Camassa-Holm equation with linear multiplicative noise $$ \left\{ \begin{array}{l} {\rm d}u+(u\partial_xu+\partial_xP[u])\,{\rm d}t=\beta u\,{\rm d}W, u(0,x)=u_0(x), P[u]=(1-\partial_x^2)^{-1}\left(u^2+\frac 1 2(\partial_x u)^2\right) \end{array} \right. $$ depends almost surely continuously on the deterministic initial data in $H^s$ for $s>3/2$. Furthermore, we prove the existence and non-uniqueness of an invariant measure for the Camassa-Holm equation with linear multiplicative noise.

math.AP

Global Existence of Weak Martingale Solutions to the Camassa-Holm Equation with Linear Multiplicative Noise

In this paper, we consider the global existence and properties of $H^1$ martingale solution to the Camassa-Holm equation with linear multiplicative noise under periodic boundary conditions. The solution is obtained as limit of regular viscous approximate solutions to parabolic SPDEs, which are constructed using the Galerkin approximations ans the stochastic compactness method. The proof of convergence to a solution argues via tightness of the laws of the viscous approximations and Skorokhod-Jakubowski a.s. representations of random variables in quasi-Polish spaces. In particular, by means of the Girsanov-type transform for regular viscous approximations and the convergence of Skorokhod-Jakubowski representations, we are able to establish the one-sided supernorm estimate and space-time higher regularity of the first-order spatial derivative, and large-time behavior of the weak martingale solution in the stochastic framework.

math.AP

Fast When, Careful Who: Dual-Process Multiparty Turn-Taking with Diffusion Augmentation

Reliable turn-taking is essential for spoken dialogue systems. However, most existing methods are designed for two-speaker interaction and struggle with realistic multiparty audio containing overlap and rapid speaker changes. We study multiparty turn-taking on the VoxConverse dataset and propose an audio-only two-stage pipeline that separates when to trigger a turn boundary from whether the floor is actually transferring. A fast trigger scans the audio and proposes candidate end-of-turn times, while a lightweight verifier runs only at those times to decide \textsc{Hold} or \textsc{Shift} and support next-speaker prediction. We report results in the full multiparty setting and a controlled dyadic top-2 projection for comparability. We also investigate diffusion-based, label-preserving background-audio mixing as a data augmentation strategy. Results show improved shift detection over a baseline, with further improvements from diffusion augmentation.

cs.CL

Implicit Identity Technologies for LLMs: Fingerprinting and Watermarking across Datasets, Models, and Generated Content

This paper presents a survey and taxonomy of LLM fingerprinting and watermarking for identity, ownership verification, provenance, and generated-content attribution. Large language models (LLMs) require substantial investments in data, computation, and expertise, and are increasingly deployed in high-stakes settings, making it critical to protect LLM-related assets and trace their origins. Existing work has rapidly expanded across dataset provenance, model ownership, and generated-content detection, but the field remains fragmented: fingerprinting and watermarking are often used inconsistently, and methods are typically studied within isolated asset-specific settings. To address this gap, we introduce implicit identity as a unifying abstraction for verifiable but not directly observable identity signals in LLM systems. We distinguish fingerprinting as non-intrusive identity derived from intrinsic characteristics, and watermarking as intrusive identity deliberately embedded into data, models, or generated content. We then propose a lifecycle-based taxonomy that organises techniques across datasets, models, and generated content, and further separates them by verification semantics: similarity-based attribution and keyed verification. Finally, we establish an evaluation framework centred on identifiability, robustness, and deployability, summarising representative metrics under realistic access and transformation regimes. By unifying terminology, lifecycle stages, and evaluation objectives, this survey provides a structured foundation for studying LLM identity technologies and for developing more reliable mechanisms for asset protection and provenance.

cs.CR

The Inclusion Depth of Pattern Languages: An Open Problem in Algorithmic Learning Theory

Pattern languages are a classical model in formal language theory and algorithmic learning theory. This note formulates the problem of computing the inclusion depth of a pattern language: the length of the longest strict inclusion chain from the universal pattern language to the language generated by a given pattern. Inclusion depth captures the mind-change complexity of pattern identification from positive data. The central open question is whether the inclusion depth ID_Sigma(p) is computable for every pattern p over every finite alphabet Sigma with at least two symbols, and whether it is computable in polynomial time. A simple conjectured formula, ID_Sigma(p) = 2|p| - #var(p) - 1, would imply a linear-time algorithm. The problem connects pattern language inclusion, combinatorics on words, language identification in the limit, and mind-change-bounded learning.

cs.FL

Correcting Visual Blur Induced by Attention Distraction to Reduce Hallucinations: Algorithm and Theory

Multimodal large language models (MLLMs) frequently suffer from object hallucinations, yet the visual perceptual mechanism underlying this failure remains poorly understood. In this work, we reveal that hallucinations are strongly associated with a human-like attention distraction phenomenon, where humans under divided focus experience degraded visual clarity and produce inaccurate descriptions, while in models the same mechanism manifests as spatial inconsistency in multi-head attention and temporal fading of attention to image tokens during decoding. We further provide theoretical insights that attention dispersion increases model complexity and degrades classification generalization. Motivated by these findings, we propose an Attention-Focused Approach for Improved Image Perception (AFIP), which corrects attention distraction via cross-head attention enrichment and reinforces visual grounding through dynamic historical attention enhancement. Extensive experiments on multiple benchmarks and models validate the effectiveness of AFIP without additional training. Code is available at: https://github.com/MIKUZ12/AFIP.

cs.CV

Meta-Soft: Leveraging Composable Meta-Tokens for Context-Preserving KV Cache Compression

The KV cache used in large language models has linearly growing time complexity, so LLMs face memory blow-up and reduced decoding efficiency when they process long contexts. Current KV Cache eviction has become an important research direction; however, existing methods based on fixed Soft Tokens (e.g., Judge Q) rely on a static parameter set as the query to evaluate the importance of KV pairs, so they cannot adapt dynamically to different input prompts, and they cannot precisely capture complex and changing task relevance. Also, evicted KV pairs are discarded permanently, so this causes irreversible information loss and context breaks. To address this problem, we propose Meta-Soft, a dynamic compression framework based on probe-driven context integration. Specifically, we build a meta-library with a learnable orthogonal basis matrix $\mathcal{L}$, and we use a selector network with Gumbel-Softmax to produce differentiable sparse combination weights, so we dynamically synthesize the most targeted $k$ Soft Tokens from the input prompt features. We append these Soft Tokens to the end of the input sequence to probe key information. We also introduce an attention-flow based integration mechanism, which redistributes the semantic information of removed tokens into retained tokens, and this keeps the dropped context information effectively. Experiments on multiple datasets show that our method outperforms existing state-of-the-art eviction methods and provides a new solution for KV Cache compression.

cs.AI

Unified definition of ferroelectricity

Recent theoretical and experimental advances in quantum ferroelectrics suggest that ferroelectricity can also emerge in non-polar space group, highlighting the limitations of conventional polar space group criteria in identifying ferroelectric materials. Here, we introduce a unified definition based on switchable polarization differences between energetically equivalent states, which naturally encompasses conventional and quantum ferroelectrics. Guided by this principle, we implement a high-throughput screening strategy that systematically identifies both conventional and quantum ferroelectrics among experimentally synthesized materials. In particular, we identify a new type of quantum ferroelectric in which the quantized polarization arises from arbitrary ionic displacements, in contrast to previous quantum ferroelectrics (including both fractional and integer quantum ferroelectrics) where quantized polarization results from fractional or integer ionic displacements. Notably, we find that materials such as Ba3I6 and Cs2PdC2 exhibit low switching barriers and robust insulating behavior, highlighting their experimental viability. Our results reconcile conventional and quantum ferroelectrics, expand the accessible materials landscape, and provide a practical roadmap for discovering next-generation ferroelectrics with advanced switchable functionalities.

cond-mat.mtrl-sci

Physical design of cold neutron direct geometry inelastic spectrometer at China Spallation Neutron Source

The Cold-Neutron Inelastic Spectrometer (CNIS) is a direct-geometry, time-of-flight instrument designed for China Spallation Neutron Source (CSNS) and optimized to probe low-energy lattice and magnetic excitations. The instrument integrates a long flight path with bent supermirror guides and an elliptical-focusing geometry to suppress high-energy background while improving cold-neutron delivery to the sample. A flexible multi-disk chopper suite provides pulse shaping, band selection and monochromatization, enabling multi-$E_\textrm{i}$ operation. Modular features, including an interchangeable high-focusing guide insert, radial collimation and a vacuum ``airbox'' for simplified sample-environment integration, enhance signal-to-noise and operational versatility. Through combined flight-path and chopper optimization, CNIS achieves excellent routine-mode energy resolution and can reach approximately $\sim 1\%$ in a dedicated high-resolution configuration. CNIS is planned to commence user operation in 2029, offering a highly flexible platform for cold-neutron inelastic scattering studies.

physics.app-ph

Photonic-Implemented Efficient Deep Quantum Neural Network via Virtual-Driven Hilbert Space Expansion

The growing computational demands of classical neural networks have intensified the search for energy-efficient and powerful computational alternatives. Quantum neural networks (QNNs) implemented on integrated photonic platforms offer a compelling avenue, offering exceptional computational power enhancements, with inherent programmability and scalability of integrated architectures. A critical challenge, however, is implementing the fundamental non-unitary and nonlinear activation function of QNNs within a linear quantum photonic system. Existing strategies, such as the adding ancillary qubits and measurement-based feedback or forward are constrained by high qubit resource costs, overhead devices, and poor cascadability. Here, we propose a novel deep photonic QNN with an expanded computational Hilbert space via input replication and mode expansion, which enables the realization of effective non-unitary and nonlinear activation on a linear programmable quantum photonic chip. This approach eliminates the need for physical ancillary qubits, measurement-induced qubit consumption and the measurement device burden, thereby significantly reduce resource costs. The fabricated chip integrates four high-quality entanglement sources and a programmable high-dimensional interferometric network, enabling a two-hidden-layer QNN that exhibits dimension-enhanced expressivity over the existing QNN architectures. We demonstrate its capabilities across diverse tasks, including nonlinear classification, image generation, and quantum Gibbs state preparation. This work establishes a scalable and efficient architecture toward practical quantum deep learning systems capable of tackling problems beyond the reach of classical computation.

quant-ph

LoViF 2026 The First Challenge on Holistic Quality Assessment for 4D World Model (PhyScore)

This paper reports on the LoViF 2026 PhyScore challenge, a competition on holistic quality assessment of world-model-generated videos across both 2D and 4D generation settings. The challenge is motivated by a central gap in current evaluation practice: perceptual quality alone is insufficient to judge whether generated dynamics are physically plausible, temporally coherent, and consistent with input conditions. Participants are required to build a metric that jointly predicts four dimensions, i.e., Video Quality, Physical Realism, Condition-Video Alignment, and Temporal Consistency. Depart from that, participants also need to localize physical anomaly timestamps for fine-grained diagnosis. The benchmark dataset contains 1,554 videos generated by seven representative world generative models, organized into three tracks (text-2D, image-to-4D, and video-to-4D) and spanning 26 categories. These categories explicitly cover physics-relevant scenarios, including dynamics, optics, and thermodynamics, together with diverse real-world and creative content. To ensure label reliability, scores and anomaly timestamps are produced through trained human annotation with an additional automated quality-control pass. Evaluation is based on both score prediction and anomaly localization, with a composite protocol that combines TimeStamp_IOU and SRCC/PLCC. This report summarizes the challenge design and provides method-level insights from submitted solutions.

cs.CV