SearcharxivSearch

arXiv subjects

Shihao Zhang

Publications and source records attributed to Shihao Zhang.

At least 19 recordsLinked to original sources

Global well-posedness of strong solutions to the initial-boundary value problem for a two-dimensional stress-diffusive Oldroyd-B model in the creeping flow regime

This paper investigates the global well-posedness of strong solutions to a stress-diffusive Oldroyd-B system in two-dimensional smooth bounded domains in the zero Reynolds number (creeping flow) regime. The stress-diffusion term is kept explicit throughout the paper and is understood as the usual center-of-mass diffusion regularization of the Oldroyd-B constitutive equation. In the corresponding non-diffusive creeping-flow setting, the strongest available result is a Beale-Kato-Majda type breakdown criterion for the three-dimensional Cauchy problem due to Kupferman, Mangoubi and Titi [Commun. Math. Sci. 6 (2008)], and global well-posedness remains open even in two dimensions. We prove the global existence and uniqueness of strong solutions for arbitrarily large H1 initial polymeric stresses satisfying the natural non-negativity condition on the conformation tensor. The result covers the initial-boundary value problem on general smooth bounded domains and shows how the Stokes elliptic structure, the preservation of the non-negativity of the conformation tensor, and the stress diffusion combine to close large-data estimates at the H1 level. We also point out a regularity feature specific to the zero Reynolds number regime: the velocity field gains higher spatial regularity from the elliptic Stokes equation than the polymeric stress tensor.

math.AP

Visualizing flat-band spatial renormalization in rhombohedral graphene superlattices

Rhombohedral graphene/hBN moir\'e superlattices exhibit flat-band-driven emergent phases, including superconductivity and the fractional quantum anomalous Hall effect (FQAHE), yet the microscopic role of the moir\'e potential remains unclear. Here, using scanning tunneling microscopy, we visualize moir\'e-modulated spatial renormalization of flat bands in rhombohedral pentalayer and tetralayer graphene/hBN superlattices. We observe spatially hierarchical filling, manifested as periodic energy shifts of the flat bands at the moir\'e scale, leading to spatial reshaping of correlated states in the interacting regime. Remarkably, this modulation vanishes below a ~10 nm moir\'e period--the same threshold below which the FQAHE is absent. Theoretical modeling attributes this mechanism to atomic-corrugation-induced charge redistribution. Our work provides real-space visualization of moir\'e-engineered flat-band reconstruction, resolving a key link between moir\'e periodic potential and emergent topological order.

cond-mat.mes-hall

Orbital-Selective Spin Splitting in the Altermagnetic Ti$_2$XX' Monolayers

Altermagnets have attracted intensive attention recently, owing to their unique combination of zero net magnetic moment of antiferromagnets and momentum space spin-splitting of ferromagnets. Herein, we systematically explore monolayer altermagnetic Ti$_2$XX' (X/X' = F, Cl, Br, I) materials for novel electronic properties. Especifically, spin-splitting exists in the valence bands of all studied materials, whereas several compositions exhibit nearly spin-degenerate conduction bands near the Fermi level, which is markedly different from conventional altermagnets. This phenomenon can be attributed to the bond-angle-sensitive super-exchange occurring in the $d_{xz}$ or $d_{yz}$ orbitals, while the super-exchange of the $d_{x^2-y^2}$ orbital is not sensitive to the Ti-X-Ti bond angle. Motivated by this intriguing phenomenon, Fermi level modulation via doping or heterostructure construction enables the transition between two electronic states, endowing these materials great application potential in information storage devices.

cond-mat.mtrl-sci

New Mid-Band (FR3, 6-24 GHz) XL-MIMO for 6G: Channel Modeling, Algorithm Evaluation, and Field Trials

The new mid-band (FR3, 6-24 GHz) spectrum is expected to play an important role in future 6G networks by providing a favorable balance among coverage, capacity, and deployment feasibility. Meanwhile, extremely large-scale multiple-input multiple-output (XL-MIMO) has emerged as a key enabling technology to exploit the propagation and spatial multiplexing potential of these frequency bands. Firstly, this paper provides a systematic review of spectrum allocation and standardization activities for new mid-band spectrum, together with the 6G spectrum planning strategies of countries and regions. Secondly, the wideband massive MIMO channel sounder is also introduced, which is specially developed for channel measurements of new mid-band with over a thousand elements. Thirdly, propagation characteristics and channel modeling approaches of four representative XL-MIMO architectures, including co-located, cell-free, and intelligent XL-MIMO, are comprehensively reviewed and analyzed, with particular emphasis on near-field propagation, spatial non-stationarity, and capacity performance. Then, recent advances in channel estimation, beamforming, and artificial-intelligence-assisted signal processing are summarized. In addition, the performance of new mid-band XL-MIMO systems equipped with 1536 and 768 antenna elements is comparatively evaluated. Finally, real communication environment prototype system field trials conducted in the Upper 6 GHz (U6GHz) band are used to investigate practical system performance under realistic deployment conditions. The results indicate that the target signal-to-noise ratio is a critical factor affecting XL-MIMO performance in the U6GHz band.

eess.SP

SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-scale LLM training systems are built around GPU-based clusters, this report presents an end-to-end optimization practice on the Ascend NPU SuperPOD. Using the DeepSeek-V4 model family as the target workload, we develop a hierarchical optimization framework spanning model-level parallelism, computation-communication orchestration, and low-level kernel execution. The resulting system achieves 34.22% Model FLOPs Utilization (MFU) with a 2.93x improvement over the open-source baseline recipe while maintaining training stability. Building on this optimized infrastructure, we further establish a CPT and SFT workflow for complex Operations Research (OR) tasks. We refer to the integrated framework as SLAI T-Rex. Using DeepSeek-V4-Flash, we develop OR-oriented CPT and SFT data pipelines that combine collected domain resources with solver-verified synthetic optimization documents. The resulting dataset contains 10K high-quality SFT samples spanning four task categories and three problem representations. The specialized model achieves the highest average zero-shot Pass@1 score among the evaluated models, reaching 71.81% and outperforming GPT-5.4-Mini and the base DeepSeek-V4-Flash model by 3.98 and 11.27 percentage points, respectively. Overall, this work demonstrates a full-stack pathway from efficient trillion-parameter model post-training on Ascend infra to domain-specialized Flash models for solver-grounded mathematical modeling, advancing frontier-model systems for complex reasoning.

cs.CL

Quantum-metric-driven light-induced ferrovalley state in d-wave altermagnets

Isolating the quantum metric from the Berry curvature remains a central challenge in quantum materials, as the two geometric quantities nearly always coexist and their contributions are difficult to disentangle. We show that d-wave altermagnets, whose real Hamiltonian possesses strictly vanishing Berry curvature, constitute an ideal platform for overcoming this obstacle. Using the Magnus expansion and exact Floquet diagonalization, we demonstrate that linearly polarized off-resonant light drives an orbital-selective ferrovalley phase through a purely quantum-metric--mediated band-gap renormalization, with no Berry curvature contribution at any order. The orbital selectivity originates from the hopping anisotropy, which generates a pronounced metric anisotropy between the $d_{xz}$ and $d_{yz}$ orbitals, and the gap reduction is expressed analytically in terms of the quantum metric. The resulting valley gap difference provides a direct, quantitative measure of the quantum metric, accessible to spin-resolved ARPES and optical pump-probe spectroscopy. This establishes d-wave altermagnets as a pristine, tunable platform in which quantum metric effects can be isolated, controlled by light polarization, and read out through valley polarization.

cond-mat.mes-hall

Large Language Model Enhanced Differentiable Trajectory Planning for IoT-Enabled Autonomous Driving

Autonomous driving planning is a key component of IoT-enabled intelligent transportation systems, requiring vehicles to generate safe, efficient, and executable trajectories in complex urban environments from multi-source contextual information. While imitation learning (IL) has shown promise on large-scale datasets, IL-based planners still suffer from limited coverage of complex long-tail interactions, weak consistency with downstream constrained refinement, and insufficient use of high level scene semantics under real time constraints. To address these issues, this paper proposes a large language model (LLM) enhanced differentiable trajectory planning framework for IoT-enabled autonomous driving. Specifically, we introduce a surrounding agent centric data augmentation strategy to reorganize sur rounding agent trajectories as additional planning supervision, thereby improving the training distribution without collecting additional raw data. We further design a complexity-aware asyn chronous LLM-based semantic enhancement module to extract scene-related high-level semantic features with controlled online overhead. In addition, a differentiable optimization module is incorporated to refine generated trajectories with explicit residual penalties while backpropagating optimization gradients to the upstream planner. Experiments show that the proposed method achieves the best overall scores of 83.63 and 78.29 on the nuPlan closed-loop nonreactive and reactive Hard20 benchmarks, respectively, and CARLA-ROS tests further verify its online deployment and real time closed-loop execution capability.

cs.RO

Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation

Recent text-to-video (T2V) diffusion models rely heavily on auxiliary reward signals (e.g., via reward models or DPO) to align generated content with human aesthetics and improve realism. These signals, however, incur substantial computational overhead, require costly human annotations, and often yield limited improvement in fine-grained local details. In this paper, we argue that your data manifold is secretly a reward model. By explicitly modeling the manifold structure of high-quality Supervised Fine-Tuning (SFT) data and encouraging video latents to lie on this manifold, we derive dense, differentiable, and nearly cost-free reward signals that significantly improve video quality, particularly in mitigating low-level distortions. Our modeling builds upon Local Coordinate Coding (LCC), which captures the `skeleton' of the manifold. However, directly applying LCC suffers from mean regression, pulling latents toward the geometric mean and losing high-frequency details. We therefore extend it to Shell Local Coordinate Coding (Shell-LCC), which models the manifold `surface' as an isotropic shell to align with the true high-density region. Experiments demonstrate that our approach improves realism, enhances high-frequency details, reduces over-smoothing artifacts, and alleviates motion blur.

cs.CV

Progress-SQL: Improving Reinforcement Learning for Text-to-SQL via Progressive Rewards

Reinforcement learning has recently shown promise in improving large language models for Text-to-SQL generation, yet existing methods typically optimize one-shot rewards defined over a single SQL state. Such rewards provide limited guidance for iterative SQL correction and are insufficient to capture the improvement of multi-turn SQL refinement. In this paper, we propose Progress-SQL, a multi-turn reinforcement learning framework with progressive rewards for Text-to-SQL. Our approach introduces an Oracle-guided Diagnostic Tree (ODT), which abstracts SQL queries into clause-level structural profiles and produces diagnostic feedback for next-turn refinement. To provide dense and robust reward signals, we combine ODT-based structural alignment with lexical alignment and define a progressive reward that measures the improvement from the initial SQL to the final SQL. We further incorporate a progression latency reward that favors earlier correctness and an execution status reward that encourages recovery from the invalid SQL. Experiments on BIRD, Spider, and Spider robustness variants demonstrate that our method consistently improves Text-to-SQL performance across both primary and robustness evaluations.

cs.CL

GPTQ-intrinsic LoRA: A Near-optimal Algorithm for Low-precision Quantization with Low-rank Adaptation

Post-training quantization is widely used for compressing large neural networks, but aggressive low-bit quantization can significantly degrade model quality. A common remedy is to augment the quantized weights with a low-rank correction, leading to approximations of the form $W\approx Q+LR$. In this paper, we study this low-precision plus low-rank representation through the layer-wise reconstruction objective $\|XW-X(Q+LR)\|_F^2$, where $X$ is a calibration matrix. We establish, to our knowledge, the first information-theoretic lower bounds for this problem under finite-alphabet and bounded low-rank compensation constraints. We then propose GPTQ-intrinsic LoRA, a training-free algorithm that incorporates the low-rank correction directly into a GPTQ-style quantization pass by appropriately augmenting the calibration Hessian. For the choice $L=V_r$, where $V_r$ contains the top right singular vectors of $X$, we prove layer-wise reconstruction error bounds in which the usual GPTQ dependence on $\|X\|_F^2$ is replaced by the rank-$r$ residual $\|X-X_r\|_F^2$, up to regularization terms. Under natural structural assumptions, these bounds match the information-theoretic lower bounds in their dominant scaling, up to constants and mild factors. We also introduce Bid-Up, a fixed-grid quantization refinement step that can be alternated with optimal low-rank compensation with guaranteed non-increasing layer-wise reconstruction error. Experiments on Qwen3 language models and DeiT vision transformers show that GPTQ-intrinsic LoRA improves over GPTQ and GPTQ followed by low-rank compensation, with additional gains from refinement loops.

cs.LG

Machine-learned atomistic simulations reveal the basis of hydrogen-induced crack-plane transition in alpha-Fe

Hydrogen-related fracture in body-centered cubic Fe and ferritic steels often appears as transgranular quasi-cleavage rather than purely intergranular failure, especially at low to moderate hydrogen contents. Fractography has suggested that hydrogen may change the dominant cleavage faceting from {100} toward {110}, but atomic-scale evidence for this possible crack-plane transition remains unclear. Here we construct an efficient neural-network potential for {\alpha}-Fe/H and combine large-scale, three-dimensional molecular dynamics with grand-canonical Monte Carlo (GCMC), allowing the near-tip crack-surface region and crack tip within a defined GCMC domain to exchange hydrogen with a reservoir at fixed chemical potential. A comparison of four crack systems identifies the controlling response: (100)[010], (100)[011], and (110)[001] remain cleavage-dominated, whereas the (110)[1-10] crack changes from dislocation emission in pure Fe to cleavage under hydrogen charging. The energetic origin is twofold. Hydrogen lowers the Griffith cleavage threshold of the {110} cleavage-plane family more strongly than that of {100}, and, for the controlling crack, a Rice-type energetic descriptor indicates that the surface-energy-controlled cleavage resistance decreases faster than the unstable-stacking-fault-controlled emission resistance, consistent with a weakened dislocation-emission shield. These results provide a thermodynamically consistent atomistic basis for a hydrogen-induced transgranular crack-plane transition in Fe.

physics.comp-ph

Rare-Earth-Tuned Evolution from d- to f-Orbital Dominance and Giant Anomalous Hall Effect in Topological RGaGe (R = Ce, Pr, Nd) Semimetals

The family of noncentrosymmetric rare-earth germanides RGaGe (R = Ce, Pr, Nd) provides a rich materials platform to explore the intertwined physics of strong magnetism, electronic correlations, and topological band structures. Through a combination of crystal growth, characterization, and first-principles calculations, we reveal that these compounds exhibit a pronounced uniaxial magnetic anisotropy, leading to distinct ground states: RGaGe orders ferromagnetically with moments along the crystallographic c-axis, and shows an antiferromagnetic-like structure in the ab-plane. A key finding is a significantly enhanced intrinsic anomalous Hall conductivity (AHC) compared to their well-known RAlGe counterparts, which even reaches as high as 948 {\Omega}-1 cm-1 at 2 K in PrGaGe. Our theoretical analysis predicts that this AHC originates from a robust Weyl semimetallic state driven by inversion symmetry breaking, where Weyl points near the Fermi level couple strongly to the magnetic order. Importantly, this topological state persists above the magnetic ordering temperature, confirming its intrinsic electronic origin. Our calculation also reveals that, while the near-Fermi-level states in CeGaGe and PrGaGe are dominated by d-orbital contributions, NdGaGe exhibits significant f-orbital involvement, signaling a progressive evolution from d- to f-orbital dominated topology. These results establish the RGaGe system as a tunable platform for systematically extending the RAlGe-related family, showcasing a large anomalous Hall response and orbital evolution near the Fermi level, and advancing the understanding of the interplay between topology and magnetism in quantum materials.

cond-mat.mtrl-sci

Nonlinear Hall quantum oscillations to probe topological Brown-Zak fermions in graphene moir\'e systems

Due to the deep connection with the quantum geometry of electronic Bloch wavefunctions, the second-order nonlinear Hall effect (NLHE) has been an attractive topic since its proposal. However, studies on NLHE under a magnetic field have been lacking. Given that quantum oscillations in the linear response regime have been proven to be useful tools in investigating electronic systems, searching for quantum oscillations in NLHE is of great interest and is expected to provide new avenues to unveil rich quantum geometric properties of novel quasiparticles. Here, we propose a new type of NLHE quantum oscillations and experimentally probe it in graphene moir\'e systems. It stems from the alternation of the dominant NLHE mechanisms with recurring Bloch states under magnetic field, which enables sensitive detection of Brown-Zak fermions, giving an onset field as low as 0.5 T. Most importantly, when the commensurability condition is satisfied, the nonlinear transport of Brown-Zak fermions is mainly governed by quantum geometric contributions. Our findings not only establish a new type of quantum oscillations, but also demonstrate the first experimental detection of the topological nature of Brown-Zak fermions, shedding light on the exploration of novel topological quasiparticles.

cond-mat.mes-hall

From Prediction to Justification: Aligning Sentiment Reasoning with Human Rationale via Reinforcement Learning

While Aspect-based Sentiment Analysis (ABSA) systems have achieved high accuracy in identifying sentiment polarities, they often operate as "black boxes," lacking the explicit reasoning capabilities characteristic of human affective cognition. Humans do not merely categorize sentiment; they construct causal explanations for their judgments. To bridge this gap, we propose ABSA-R1, a large language model framework designed to mimic this ``reason-before-predict" cognitive process. By leveraging reinforcement learning (RL), ABSA-R1 learns to articulate the why behind the what, generating natural language justifications that ground its sentiment predictions. We introduce a Cognition-Aligned Reward Model (formerly sentiment-aware reward model) that enforces consistency between the generated reasoning path and the final emotional label. Furthermore, inspired by metacognitive monitoring, we implement a performance-driven rejection sampling strategy that selectively targets hard cases where the model's internal reasoning is uncertain or inconsistent. Experimental results on four benchmarks demonstrate that equipping models with this explicit reasoning capability not only enhances interpretability but also yields superior performance in sentiment classification and triplet extraction compared to non-reasoning baselines.

cs.CL

Third-order optical response in d-wave altermagnets: Analytical and numerical results from microscopic model

Altermagnets represent a novel category of magnetic materials characterized by zero net magnetization yet featuring spin-split band structures, and they demonstrate distinctive orbital-spin locking phenomena. Commencing from the minimal multi-orbital tight-binding Hamiltonian of d-wave altermagnets, we conduct an analysis of the general formulas for the third-order injection and shift currents. These currents are solely determined by the quantum metric and quantum connection, being free from Berry curvature contamination. In the ideal scenario where the $\delta$-bond hopping $V_\delta$ approaches zero ($V_\delta = 0$), we derive closed-form analytical solutions for the third-order photoconductivities. For the general situation with a finite value of $V_\delta$, we present a perturbative analytical solution within the limit of $V_\delta \ll V_\pi$, and this solution is verified through numerical calculations. Our research establishes a comprehensive theoretical description of the third-order optospintronic responses in d-wave altermagnets based on a microscopic model. Moreover, it offers a viable approach for the experimental observation of pure quantum geometric effects.

cond-mat.mes-hall

Efficient Matrix Implementation for Rotary Position Embedding

Rotary Position Embedding (RoPE) has become a core component of modern Transformer architectures across language, vision, and 3D domains. However, existing implementations rely on vector-level split and merge operations that introduce non-negligible computational overhead, often overlooked in attention optimization. The problem is further amplified in multi-dimensional settings (e.g., 2D and 3D RoPE), where additional vector operations and uneven feature partitions degrade hardware utilization. To overcome these limitations, we propose RoME (Rotary Matrix position Embedding), a mathematically equivalent yet computationally efficient reformulation of RoPE that replaces vector operations with unified matrix transformations. RoME eliminates dimension-specific operations, simplifies implementation, and enables fused parallel execution across Cube and Vector units on modern NPUs. Experiments show that RoME delivers substantial acceleration at both the operator and full-model levels. The implementation is available at https://gitcode.com/cann/ops-transformer/blob/master/experimental/posembedding/rope_matrix/README.md.

cs.LG

Interlayer Coupling Driven Correlated and Charge-Ordered Electronic States in a Transition Metal Dichalcogenide Superlattice

4Hb-TaS_2, a van der Waals superlattice comprising alternate stacked Ising superconducting 1H-TaS_2 and cluster Mott insulating 1T-TaS_2, exhibits emergent properties beyond those of its constituent layers. Notable phenomena include time-reversal-symmetry-breaking superconductivity and spontaneous vortex phases, which are driven by nontrivial interlayer interactions that remain debated. Using area-selective angle-resolved photoemission spectroscopy, we provide direct spectroscopic evidence of such interaction by systematically probing the electronic structures of 1T- and 1H-terminted surfaces of 4Hb-TaS_2. The metallic states of subsurface 1H-layers are folded to the Brillouin zone center by the sqrt(13) by sqrt(13) modulation of the surface 1T-layer, forming chiral "windmill" Fermi surfaces via Umklapp scattering. These conducting states further hybridize with the incipient flat band of the surface 1T-layer, producing a Kondo-like peak at the Fermi level. Interlayer charge transfer induces distinct 3 by 3 and 2 by 2 charge orders on the surface and subsurface 1H-layers, respectively, which result in characteristic segmented Fermi surfaces and dichotomously shift the van Hove singularities. These findings reconcile the competing Kondo and Mott-Hubbard models in this material and emphasize the interplay of flat bands, van hove singularities, charge orders, and unconventional superconductivity in correlated superlattices.

cond-mat.str-el

A Simple and Effective Random Forest Modelling for Nonlinear Time Series Data

In this paper, we propose Random Forests by Random Weights (RF-RW), a theoretically grounded and practically effective alternative RF modelling for nonlinear time series data, where existing RF-based approaches struggle to adequately capture temporal dependence. RF-RW reconciles the strengths of classic RF with the temporal dependence inherent in time series forecasting. Specifically, it avoids the bootstrap resampling procedure, therefore preserves the serial dependence structure, whilst incorporates independent random weights to reduce correlations among trees. We establish non-asymptotic concentration bounds and asymptotic uniform consistency guarantees, for both fixed- and high-dimensional feature spaces, which extend beyond existing theoretical analyses of RF. Extensive simulation studies demonstrate that RF-RW outperforms existing RF-based approaches and other benchmarks such as SVM and LSTM. It also achieves the lowest error among competitors in our real-data example of predicting UK COVID-19 daily cases.

stat.ME