SearcharxivSearch

arXiv subjects

Yang Xu

Publications and source records attributed to Yang Xu.

At least 19 recordsLinked to original sources

Agent-Based ML-LLM Fusion with Self-Optimizing Prompts for Plateau Weather Alerts

To address insufficient contextualization, weak generalization, and poor scenario adaptation in tourism meteorological services, we propose SmartWeatherAgent--a unified three-stage architecture integrating intent recognition, hazard prediction, and reasoning-enhanced generation. The system fuses rule-based methods with large language models to parse queries at multiple granularities and employs a LightGBM model enriched with highland-specific features (e.g., wind speed abruptness rate), achieving an F1-Macro score of 0.605 with 1.60 ms latency on high-wind, precipitation, and low-temperature events. A 12-round micro-step prompt self-optimization loop boosts the composite warning quality score S_final from 4.2 (B01) to 8.9 (B12, +112%). Key improvements include a sharp rise in B08 from data source citation (6.5 -> 8.5), sustained high performance in B10 via physical mechanism explanation, and a peak scientific rigor score of 9.2 in B12 through explicit uncertainty statements. The system autonomously generates structured warnings that integrate causal mechanisms, spatiotemporal evolution, quantitative evidence, regulatory references, and confidence statements--enhancing professional depth, logical rigor, and scientific soundness, and advancing meteorological services toward proactive perception, explainable decision-making, and intelligent agency.

cs.AI

CEDAR: Error-Bounded Residual Routing for Efficient Long-Context Attention

Post-hoc sparse attention accelerates long-context prefill by routing each query to a small set of token-level interactions. Hard selection, however, assigns zero probability to every omitted chunk: a routing miss cannot be recovered, and a fixed expansion budget spends the same work on easy and ambiguous queries. We introduce Coarse-to-fine Error-aware Dynamic Attention Routing (CEDAR), a coarse-to-fine method that keeps the language model frozen while preserving global coverage. Each semantic chunk contributes a cheap key--value summary to a residual attention path; chunks with high estimated approximation error are then expanded to exact token attention. Exact and summarized contributions are combined in a single softmax normalization, so refinement replaces, rather than duplicates, coarse evidence. We derive an output-error bound governed by within-chunk key/value dispersion and use it to allocate a variable refinement budget. A controlled clustered-attention study shows that residual summaries reduce reconstruction error by more than 98% relative to hard dropping at equal exact-chunk budgets. Experiments on long-context benchmarks demonstrate that CEDAR recovers most of the quality lost by hard sparse routing while maintaining approximately $3\times$ kernel speedup at 128K context.

cs.CL

On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability

We describe the architecture and ablations of Qwen3.8-Flash-Next, a sparse mixture-of-experts model with 125B parameters, 6B activated per token, and additional 51B parameters of n-gram embedding tables held off the accelerator. On fourteen pre-training benchmarks the model leads the 397B-A17B predecessor on eight and trails it on the rest by at most 2.6 points, at 1/3 the activated parameters, 1/3 the training tokens, and roughly 1/9 the training FLOPs. Token mixing uses a layer-wise hybrid of Gated DeltaNet (GDN) and global attention, with one full-attention layer in every four; at continued-pretraining time those full-attention layers are replaced by Qwen Sparse Attention (QSA), which scores context at micro-block granularity with a compressed lightweight indexer. The residual stream is widened to four branches and read through an elementwise gate, a design we call the Gated Residual (GR). Capacity is added outside the backbone by a single n-gram embedding layer whose tables are prefetched from host memory. We evaluate every candidate change along three axes: loss together with downstream benchmarks; the cost of the change in training, prefill and decode; and its effect on the optimal hyperparameters and training stability. Loss and downstream accuracy do not always move together: enlarging the n-gram vocabulary lowers loss monotonically while downstream accuracy saturates. The architecture and the Muon optimizer together shift the optimal learning rate and batch size upwards, render batch-size warmup unnecessary, and substantially improve stability under stress tests. Loss, benchmarks, efficiency and stability form one design problem. Solved jointly, they yield a recipe that is simultaneously more efficient, more capable and more stable.

cs.CL

Supermoir\'e Reconstruction and Topological Mosaics in Twisted Trilayer WSe$_2$ and MoTe$_2$

We investigate lattice relaxation and band structures of helical and alternating twisted trilayer WSe$_2$ and MoTe$_2$ using machine-learning force fields and large-scale ab initio calculations. Interference between the two bilayer moir\'e lattices generates a supermoir\'e lattice that, upon relaxation, reconstructs into a few dominant domain types with locally commensurate bilayer moir\'e lattices. Because the systems lack $C_{2z}$ symmetry, domains otherwise related by this symmetry become energetically and topologically distinct, unlike in twisted trilayer graphene. Band structure calculations show that the topmost valence bands originate from the $K$ ($K'$) valleys and carry domain-dependent valley Chern numbers. The resulting supermoir\'e lattice hosts a mosaic of topologically inequivalent domains, offering a platform for exploring correlated and topological physics.

cond-mat.mes-hall

Selection-Aware Stress Testing for Interactive Agents

Agent evaluations often use one benchmark to choose a workflow and then search for task types where its advantage weakens, so both conclusions are selected from the same data. We introduce Selection-Aware Semantic Stress Testing (\SASST{}), which learns a task reweighting from pre-execution features on discovery tasks and evaluates the same paired comparison on separate confirmation tasks. The protocol checks support and stability, uses joint bounds for all planned claims, and can return no claim. We prove conditional asymptotic validity under stated cluster assumptions. A forty-cluster audit finds Gaussian undercoverage and conservative Bonferroni $t$ bounds. In one 480-episode $\tau$-bench study, a $3.75$ point discovery gain vanished on confirmation. A second-model study likewise confirmed neither a workflow benefit nor a stable stress rule.

cs.LG

Evidence for Three-component Interlayer Coherent Exciton Condensation

Increasing the number of internal components in a quantum many-body system can host collective orders inaccessible to simpler settings. Quantum Hall bilayers provide a canonical realization of interlayer exciton condensation, yet extending such coherence across three independently addressable electronic fluids has remained elusive. Here we report evidence for three-component interlayer coherent exciton condensation in triple-layer graphene system. Using Rydberg excitons in an adjacent WSe2 monolayer as a layer-sensitive optical probe, we resolve interaction-induced incompressibility at zeroth-Landau-level crossings for all three pairwise layer combinations, establishing top-middle, middle-bottom and top-bottom exciton condensate channels within the same device. Independent control of displacement field and interlayer bias continuously tunes these pairwise states towards a regime where Landau levels from all three layers approach simultaneous degeneracy. At their convergence, incompressibility persists while the exciton energy and spectral weight evolve smoothly between the pairwise limits, suggesting coherent participation of all three layers in a single three-component state. More broadly, the ability to independently control layer potentials and engineer interlayer interactions establishes multilayer graphene as a programmable synthetic dimension for exploring higher-component quantum Hall order and simulating strongly correlated quantum matter.

cond-mat.mes-hall

Influence of twist direction and large deformation on soft material torsional contact

Shear-induced contact area reduction is widely observed in soft contacts, yet recent torsional experiments have revealed a more complex non-monotonic evolution in which the contact area first increases and then decreases with twist angle. The mechanism responsible for this initial area increase and the role of large deformation in the overall area evolution remain unclear. In this study, we experimentally investigate the torsional contact response of soft Polydimethylsiloxane (PDMS) spheres by combining forward-backward twist tests with a systematic variation of the curing-agent-to-base ratio to tune material softness and deformation level. The loading-unloading tests show that the torsional interface is strongly irreversible: during unloading, the contact area follows a decrease-increase-decrease path rather than retracing the loading branch, and repeatable petal-like edges appear, indicating a wrinkle-induced surface instability. By decreasing the mixing ratio, we find that larger deformation strengthens the area-reduction contribution and eventually suppresses the initial area increase, leading to a monotonic area decrease during loading for sufficiently soft PDMS. Softer PDMS also exhibits lower shear strength, weaker torque oscillations, and improved repeatability. The results provide experimental evidence that large deformation can drive shear-induced contact area reduction, while the origin of the initial area increase remains unresolved. These findings narrow the possible mechanisms (e.g., triboelectrification) responsible for the initial area increase and provide a stringent benchmark for frictional contact models of soft interfaces.

cond-mat.soft

Breaking the mutual exclusivity between metallicity and ferroelectricity in a non-polar covalent semiconductor via orbital selective doping

The mutual exclusion of ferroelectricity and metallic conductivity is a long-standing tenet because itinerant electrons screen long-range Coulomb forces that stabilize the bulk polar order. Here, we break this paradigm by heavily doping a non-polar covalent semiconductor of cubic silicon carbide (3C-SiC) with nitrogen. This introduces heavy electron doping, inducing metallicity and driving a structural transition from the non-polar F-43m to the polar R3m symmetry via the pseudo-Jahn-Teller effect. Remarkably, we provide direct, atomic-scale visualization of about 180{\deg} polarization reversal under an external voltage bias in a ferroelectric metal. The strongly directional character of antibonding orbitals occupied by conduction electrons prevents them from screening the local Si-C polarization, resulting in the coexistence of metallicity and ferroelectricity. Ferroelectric tunnel junctions demonstrate nonvolatile memory properties with a well-defined high-resistance state (HRS) and low-resistance state (LRS), an ultrahigh response speed (~50 ns), an ultralow operating voltage (1 V), an endurance exceeding 85927 cycles, and a projected retention time of 100 years. Our results provide a novel strategy for pioneering ferroelectricity in a metal, a new ferroelectric metal platform for exploring exotic properties, and a ferroelectric device with high performance that meets the requirements for low consumption and high-speed non-volatile devices.

cond-mat.mtrl-sci

LiveMem: Maintaining Memory State Continuity in Long-Running LLM Inference

Long-running assistants and agents consume interaction streams that eventually outgrow the context. Existing context retention, summarization, and retrieval preserve access to selected history, but do not provide a persistent state over the full lifecycle when working context changes. We formulate this missing inference capability as \emph{state continuity under context turnover}: carrying computation forward through a fixed-capacity memory state whose lifetime is independent of the active context. We introduce an intrinsic memory method, \textbf{LiveMem}, which augments a pretrained full-attention LLM with a memory state that preserves the historical information over the whole lifecycle while the main attention path retains a bounded KV window. Context turnover and memory state maintaining, memory-oriented post-training, and state-aware serving jointly make this memory state load bearing after its originating tokens are released. Our experiments show that LiveMem achieves leading overall performance among evaluated systems and other intrinsic memory methods. Experiments on LongMemEval show that LiveMem is able to answer the question based on the memory state, even when the supporting evidence has been removed from the current context, and evidence-distance analysis shows that useful information persists beyond the active window. LiveMem thus establishes state continuity as a distinct and complementary abstraction for continual LLM inference.

cs.CL

SeededGrasp: Language-Guided Grasping in Complex Scenes with Multiple Embodiments

Practical robotic grasping in complex scenes requires both 3D spatial reasoning and alignment with task-specific requirements. Vision-language models (VLMs) offer a natural way to specify these requirements using language, but existing approaches either use a VLM to predict the grasp directly with limited spatial awareness, or train the VLM together with the grasping model, which requires significantly more data and compute. These limitations impede performance and have prevented scaling to multiple embodiments in complex scenes. We address this by proposing SeededGrasp, a novel data-efficient framework that enables a VLM to predict a seed point to be used as conditioning for a subsequent lightweight grasp-generation model. Our architecture decouples high-level semantic reasoning from low-level geometric execution, enabling multi-embodiment support while bypassing the need for expensive end-to-end training. To enable training such models, we release the first multi-embodiment tabletop grasping dataset comprising over 2.5M grasps in cluttered scenes. Experimental results demonstrate that our approach outperforms existing baselines, achieving 72% success in simulation and 78% in real-world grasping experiments. See our project site for data and code: https://uoft-isl.github.io/seeded-grasp/

cs.RO

MagicSelector: Joint Optimization for Agent Tool Selection via Counterfactual Decomposition and Progressive Reranking

We present MagicSelector, a joint optimization framework integrating Counterfactual task decomposition, Progressive reranking, and Dynamic Top-K, designed to address the fundamental challenges of tool retrieval in agents. MagicSelector is a specialized framework capable of translating ambiguous user instructions into executable atomic subtasks and guiding high-precision tool retrieval, effectively mitigating redundant noise and severe context distraction in out-of-domain (OOD) scenarios. We empower MagicSelector with these capabilities through three key contributions: (1) a preference-guided counterfactual task decomposition mechanism that utilizes a counterfactual reward to quantify the marginal causal gain of decomposition on retrieval ranking, effectively imposing fine-grained structural supervision on logical coherence; (2) a progressive tool reranking method driven by self-distillation hard negative mining, which optimizes both point-wise and list-wise relevance to enhance fine-grained discrimination among highly similar tools; and (3) a dual semantic boundary-aware dynamic Top-K strategy that adaptively monitors reranking score cliffs and inter-tool semantic shifts to dynamically truncate the candidate list, maximizing relevant tool recall while filtering long-tail noise. Evaluated on MTDTool, the first task decomposition benchmark we constructed tailored for mobile multi-turn interactions with process-level annotations, MagicSelector yields promising performance. Extensive experiments demonstrate that MagicSelector significantly outperforms state-of-the-art methods in terms of tool retrieval accuracy, OOD generalization capability, and overall token efficiency, thereby demonstrating the effectiveness of our proposed framework.

cs.IR

Unconventional superconductivity in ScIr$_2$ chiral crystal with a kagome lattice

Materials with a kagome lattice host exotic quantum phenomena driven by the interplay between band topology, spin-orbit coupling, magnetism, and electronic correlations. While magnetism of kagome materials has been widely investigated, their unconventional superconductivity (SC) remains largely unexplored due to the limited availability of suitable materials. Here, we report evidence of unconventional SC in the ScIr$_{2-x}$Si$_{x}$ family by combining muon-spin spectroscopy measurements with band-structure calculations. The parent ScIr$_2$ undergoes a structural phase transition from a high-$T$ cubic- to a low-$T$ rhombohedral phase, while the Ir kagome layer remains, albeit slightly distorted. Although the structural transition is suppressed by Si substitution, the superconducting pairing of ScIr$_{2-x}$Si$_{x}$ remains well described by a two-gap model. Since at least one of the gaps has nodes, this indicates an unconventional SC. Its unconventional nature can be explained by the distinct flat bands occurring near the Fermi level, leading to strong electronic correlations in the ScIr$_{2-x}$Si$_{x}$ family. Moreover, the low-$T$ phase of ScIr$_2$ exhibits an Ir chiral chain; therefore, it can be classified as a topological chiral crystal. Overall, the unusual properties of the ScIr$_{2-x}$Si$_{x}$ family make it an interesting, albeit rare, system for studying the interplay between unconventional SC, flat bands, and chirality.

cond-mat.supr-con

Stochastic process model of rough surface contact

The stochastic process model of rough surface contact, widely known as Persson's theory of contact, serves as a representative multi-scale model that has been extensively applied across various fields of tribology. In this chapter, we briefly introduce the background of the development of Persson's theory of contact. We thoroughly discuss Persson's theory for purely normal elastic contact, with a special focus on solving the probability density of the contact pressure and the interfacial gap using partial differential equations. Subsequent applications of these fundamental results in addressing more complex interfacial properties in other fields of tribology are also examined. Finally, several recommendations regarding future studies of Persson's theory are proposed. This review article is expected to assist researchers in quickly familiarizing themselves with the current state of the art of Persson's theory and to attract more attention from tribologists and solid mechanicians, thereby contributing to the development and application of Persson's theory of contact.

cond-mat.soft

$r$-Minimal Poset Codes

In this paper, we propose and study $r$-minimal codes with respect to $\mathbf{P}$-support, where $\mathbf{P}=(\Omega,\preccurlyeq_{\mathbf{P}})$ is a poset defined on the coordinate set of the ambient space $\mathbf{H}$. $r$-Minimal $\mathbf{P}$-codes are natural extensions of Hamming metric minimal codes that have been extensively studied in the literature. We characterize $r$-minimal $\mathbf{P}$-codes in terms of the notion so called cutting $r$-blocking maps, which generalizes the well-known equivalence between minimal Hamming metric codes and cutting blocking sets. We also give a necessary and sufficient condition for $r$-minimality in terms of $(\mathbf{P},\omega)$-weight defined on $\mathbf{H}$, where $\omega:\Omega\longrightarrow\mathbb{R}^{+}$ is an arbitrary weight function. This leads to a generalization of the well-known Ashikhmin-Barg criterion for Hamming metric minimal codes. We then prove two existence results for $r$-minimal $\mathbf{P}$-codes, both for general $\mathbf{P}$ and for the special case that $\mathbf{P}$ is a disjoint union of chains. When $\mathbf{P}$ is hierarchical, we characterize $r$-minimal $\mathbf{P}$-codes in terms of $r$-minimal Hamming metric codes. Finally, we characterize cutting $r$-blocking sets induced by hierarchical posets with two levels, which further enables us to answer a question raised in Hyun, Kim, Wu and Yue \cite{28}.

cs.IT

List-Decoding Counterexamples Yield Lower Bounds on Mutual Correlated Agreement Error

Mutual correlated agreement captures whether a random linear combination of received words can create a new large agreement with a code, a property relevant to the soundness of batched proximity testing. We show constructively that list-decoding counterexamples yield lower bounds on the mutual correlated agreement error. Given an explicit counterexample to the $(p,L)$-list-decodability of a linear code over $\mathbb{F}_q$, we construct a related code $C'$ of the same length and dimension such that $\operatorname{err}_{\mathrm{MCA}}(C',p)\ge\frac{1}{q}\left\lceil\frac{(L+1)q}{q+L}\right\rceil$, while decreasing its minimum distance by at most one. The construction also produces an explicit pair of words witnessing this error. We further give a structure-preserving version for code families whose coordinates are indexed by a finite set $\Omega$, with each index determining a generator-matrix column through a map $v:\Omega\to\mathbb{F}_q^k$. The construction changes at most one coordinate index and ensures that the output code remains in the same indexed family. As applications, we instantiate this principle for algebraic-geometry (AG) evaluation codes and Reed--Solomon codes. For AG codes, if $G$ is the divisor defining the underlying Riemann--Roch space and $N$ is the number of rational places outside $\operatorname{supp}(G)$ available for evaluation, the resulting code remains over the same function field and Riemann--Roch space, with a modified set of evaluation places. Its mutual correlated agreement error is at least $\frac{1}{q}\left\lceil\frac{(L+1)N}{N+L\mathrm{deg} G}\right\rceil$. The Reed--Solomon conclusion follows as the Vandermonde-column specialization.

cs.IT

Pelican-VLA 0.5: Attending Before Acting Benefits Generalization

In this report, we present Pelican-VLA 0.5, a unified VLA model that integrates vision-language understanding, future-frame generation, and action prediction within a single architecture. Pelican-VLA 0.5 achieves attention-level generalization: without object annotations, segmentation masks, attention supervision, or task-specific fine-tuning, its action pathway already focuses on the manipulation-relevant object and contact region. This behavior persists across unseen scenes and unseen robot embodiments, and is substantially stronger than in other open-source VLA baselines. We verify that this ability originates from the learnable Bottleneck Token inserted between perception and action: by routing task-relevant visual information through a compact bottleneck, the tokens interface induces manipulation-centric attention during pre-training and remains effective across different policy structures, including a MoT-style architecture.

cs.RO

Mass weighting algorithm optimizes Fourier-based physics-informed neural network in adhesive contact mechanics

Physics-informed neural networks (PINNs) for elastic contact mechanics suffer from a spectral stiffness imbalance,that is, the elastic kernel grows linearly with wave number, causing short-wavelength modes to dominate gradient updates and stall convergence of the macroscopic deformation. We introduce a spectral preconditioning strategy that reweights displacement gradients in Fourier space before back-propagation, amplifying low wavenumber components through a mass weighting (MW) function while suppressing sub-grid noise via a built-in low-pass filter. Applied to adhesive line contact problems, the mass weighted PINN reaches machine-zero residual loss within 400 Adam iterations for specified benchmark, whereas the reference benchmark stalls at three orders of magnitude higher loss. The converged displacement and contact stress fields agree quantitatively with Green's function molecular dynamics (GFMD) solutions for both smooth Hertz contact at pressures spanning tension to compression and rough surfaces with roughness covering several decades of wavelength. The method operates directly on a uniform real-space grid, requires no explicit Green's function integration or quadrature rules, and is formulated entirely in terms of minimising a scalar energy function. Extension to two-dimensional rough surfaces is direct, as both the Fourier elastic energy and the spectral preconditioner depend only on the wave-number magnitude.

cond-mat.soft

Generalized Rank Weight and Extended Generalized Poset Weight Defined For Codes Over Rings: A Galois Connection Approach

In this paper, we study generalized rank weights (GRWs) and extended generalized poset weight (EGPWs) of codes over rings via a Galois connection approach. First, we show that various coding-theoretic properties related to generalized weights, including security drops of a code employed in wire-tap channel of type II, connections between generalized weights of a Gabidulin code and its associated Delsarte code, (generalized) Singleton bound, MDS discrepancy of a code, characterizations of MDS, near MDS, $i$-MDS, MRD, near MRD, $i$-MRD, (dually) quasi-MRD codes as well as evasive property of subspaces, can be reformulated in terms of Galois connections. Next, we study GRWs and rank profiles defined for modules over principal ideal rings, especially those over chain rings. Generalizing GRWs defined for vector spaces over fields, we establish a singleton bound and a Wei-type duality theorem, characterize MRD, near MRD and dually quasi-MRD codes and determine their GRWs; moreover, we characterize $i$-MRD codes and establish a scattered bound for $(h,h)$-evasive codes over chain rings, generalizing counterpart result established for vector space over finite fields. Finally, we propose and study EGPWs and extended poset profiles defined for modules with a composition series, which in fact form a Galois connection. Generalizing EGPWs defined for modules over finite Galois rings, we establish a Wei-type duality theorem for modules over arbitrary quasi-Frobenius rings, which unifies the two Wei-type duality theorems derived in both \cite{32} and \cite{33}.

cs.IT