SearcharxivSearch

arXiv subjects

Tao Yao

Publications and source records attributed to Tao Yao.

At least 19 recordsLinked to original sources

Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization

Large language models (LLMs) are increasingly used to formulate optimization models from natural-language problem descriptions, yet realistic operations research (OR) requests are often incomplete: missing objectives, constraints, or business rules can change the resulting mathematical program. Existing evaluations largely assume a complete specification and therefore overlook whether an agent knows when clarification is needed before modeling. We introduce OR-Clarify, a benchmark for pre-formulation clarification. Each task presents a partial public problem description, withholds structured hidden slots, and evaluates agents through bounded interaction with a simulated user. The benchmark supports both openended and choice-based clarification, and measures slot recovery, stopping behavior, silent assumptions, and interaction cost. We further propose Interactive Optimization (InterOPT), a two-stage framework that identifies unresolved formulation-critical gaps and uses them to guide whether to ask the next question or to stop. In our choice-based experiments, InterOPT substantially outperforms all baselines in exact slot recovery; in the open-ended setting, it remains competitive with strong prior methods. Together, OR-Clarify and InterOPT reframe OR assistance as a selective completeness decision: clarify when needed, stop when ready, and quantify what remains missing.

math.OC

Multistate ferroelectricity and switchable layer-locked anomalous valley Hall effects in bilayer ReIrGe2Se6

Two-dimensional multiferroic materials, which combine magnetic and ferroelectric (FE) orders with strong magnetoelectric coupling, represent ideal platforms for high-density information storage and low-power multistate electronics. However, the intrinsic bistability of conventional ferroelectricity poses a substantial challenge to realizing multiple nonvolatile states and programmable Berry-curvature driven transport responses within a single material. Here, using first-principles calculations, we predict multistate ferroelectricity in AA0-stacked bilayer ReIrGe2Se6. The system hosts four energetically stable FE polarization configurations, among which three are connected through reversible switching pathways, while the fourth exhibits a unidirectional switching pathway. The distinct FE configurations further give rise to a cyclic semiconductor-metal-semiconductor evolution in the electronic structure. Notably, FE polarization switching is intimately coupled to layer degrees of freedom and Berry curvature. The layer-dependent electrostatic potential associated with different FE configurations controls the layer character of the band-edge states, thereby locking the Berry curvature to specific layer channels. As a result, bilayer ReIrGe2Se6 enables switching between an anomalous valley Hall effect and a layer-locked anomalous valley Hall effect, providing nonvolatile control of layer, valley, and spin-resolved transport responses. In addition, magnetization reversal switches the valley and spin channels while preserving the layer-resolved character. These results establish bilayer ReIrGe2Se6 as a multistate ferroelectric platform for programmable Berry-curvature related transport, offering microscopic insight into topology based multifunctional electronic and valleytronic devices.

cond-mat.mtrl-sci

Interlayer sliding direction as a symmetry selector in altermagnetic bilayer Fe2WS4: Switchable anomalous Hall and anomalous valley Hall effects

Altermagnets combine compensated collinear magnetic order with momentum-dependent spin splitting, offering a promising platform for coupling spin and valley degrees of freedom with ferroelectricity and Berry-curvature driven transport in the absence of net magnetization. However, achieving nonvolatile and selective control of these intertwined degrees of freedom remains a key challenge. Here, using first-principles calculations, we show that the direction of interlayer sliding serves as a symmetry selective control parameter in altermagnetic bilayer Fe2WS4. Diagonal sliding breaks inversion symmetry and produces two sliding ferroelectric states with opposite out-of-plane polarizations. Reversal of the ferroelectric polarization switches the momentum-dependent spin texture and reverses the anomalous Hall conductivity, revealing strong magnetoelectric coupling and enabling a ferroelectrically switchable anomalous Hall effect. In contrast, axial sliding preserves inversion symmetry but breaks the crystalline symmetry relating the X and Y valleys, leading to reversible valley polarization and a switchable anomalous valley Hall effect. These results establish the direction of interlayer sliding as a nonvolatile symmetry selector for controlling ferroelectricity, spin texture, valley polarization, and Hall transport responses in two-dimensional altermagnetic bilayers.

cond-mat.mtrl-sci

Fractional quantum ferroelectric control of spin-valley locking and valley Hall effects in altermagnetic monolayer Cr2S2

Fractional quantum multiferroics, arising from the coupling between fractional quantum ferroelectricity (FQFE) and altermagnetism (AM), provide a promising platform for nonvolatile control of momentum dependent spin splitting in systems with zero net magnetization. However, extending this FQFE-AM coupling to valley degrees of freedom and Berry curvature driven valley Hall effects remains largely unexplored. Here, using first-principles calculations, we demonstrate that monolayer Cr2S2 realizes a two dimensional FQFE-AM platform with two switchable FQFE states connected by composite symmetry operations combining a fractional lattice translation with time reversal or parity-time reversal. We show that FQFE switching reverses the AM spin-polarized band structure and interchanges the spin characters of the X and Y valleys without rotating the Néel vector, thereby enabling polarization switchable spin-valley locking. Moreover, the two FQFE states exhibit reversed Berry curvature distributions, which, together with the switched spin-valley locking, enable polarization controlled valley Hall effects under both electron and hole doping. These results demonstrate a symmetry based mechanism for nonvolatile electrical control of AM spin splitting, spin-valley locking, and valley Hall effects, offering a general route toward low-power valleytronic devices based on FQFE-AM coupling.

cond-mat.mtrl-sci

RFHNet: Relational and Frequency-Aware Hashing Network for Large-Scale Fine-Grained Food Image Retrieval

Fine-grained food image retrieval is a key task in computational gastronomy, with applications in food traceability, dietary monitoring, and smart catering systems. Although hashing-based retrieval is attractive for large-scale search due to its storage efficiency and fast Hamming-distance computation, existing methods often perform poorly in fine-grained food scenarios, where subtle local semantics and frequency-sensitive visual cues are essential. To address this challenge, we propose RFHNet, a cascaded hierarchical hashing network that captures both global structure and fine-grained local details through multi-level representations. RFHNet includes three components: (1) Fine-grained Relation Modeling (FRM) to capture subtle visual differences among similar food components; (2) Multi-Frequency Modulated Fusion (MFMF) to extract informative multi-frequency features; and (3) Hierarchical Semantic Synergy (HSS) to adaptively integrate multi-level representations and generate discriminative hash codes. Experiments on six food-specific benchmarks show that RFHNet consistently outperforms state-of-the-art hashing methods, with mAP gains of 4.44\% to 17.20\% at 12 bits. These results validate the effectiveness of RFHNet for large-scale visual food retrieval and smart catering applications. The source code will be released upon publication.

cs.CV

All-electrical switching of spin texture in a strain-tunable 2D Janus ferroelectric altermagnet

Altermagnetism (AM), a collinear magnetic phase with momentum-dependent spin splitting, is a promising candidate for strong magnetoelectric coupling. However, realizing direct and tunable coupling between ferroelectricity (FE) and AM within a single two-dimensional (2D) material remains an outstanding challenge. Here, based on first-principles calculations, we identify the distorted phase of monolayer Janus VOClBr as an intrinsic 2D FE-AM. This phase demonstrates robust magnetoelectric coupling, as evidenced by a complete reversal of momentum-space spin polarization upon FE switching, and further supported by spin texture analysis and the magneto-optical Kerr effect. Notably, the FE properties are highly strain-tunable: biaxial compression strain of -4% reduces the FE polarization switching barrier by approximately 87%, whereas a tensile strain of +3% induces a phase transition to an antiferromagnet. Leveraging the lock-in between the electrically controlled spin texture and the magneto-optical Kerr effect signal, we propose a non-volatile, polymorphic spintronic memory device featuring all-electrical writing and optical readout. This work establishes 2D FE-AMs as a versatile platform for coupled ferroic orders and paves the way for voltage-controlled, multifunctional spin-logic devices.

cond-mat.mtrl-sci

Strain-tunable multipiezo effects in Janus monolayer Cr2SSe: Selective reversal of valley polarization and single-spin-channel anomalous valley Hall effect

Altermagnetism, the third class of collinear magnetic order, uniquely combines a zero net magnetization with spin polarized bands in reciprocal space, opening new avenues for two dimensional valleytronics and spintronics. Here, using first principles calculations, we predict that the Janus monolayer Cr2SSe, which possesses intrinsic inversion symmetry breaking, hosts a strain tunable multipiezo effect and exhibits distinctive valleytronic properties. The system displays pronounced spin splitting and band inversion at the X and Y high symmetry points in the Brillouin zone, giving rise to robust spin-valley locking. The degeneracy of these valleys is protected by diagonal mirror symmetry. Application of uniaxial strain breaks this symmetry, concurrently inducing piezovalley, piezoelectric, and piezomagnetic responses, a manifestation of the multipiezo effect. Critically, strain applied along orthogonal crystallographic directions yields opposite valley polarization, while under small compressive strain, we achieve selective reversal of valley polarization, enabling independent control of valence and conduction band valleys and promoting a single-spin-channel anomalous valley Hall effect. These findings establish a pathway for low-power, non volatile manipulation of valley degrees of freedom and enhanced spin transport efficiency, providing a theoretical foundation for the design of energy-efficient valleytronic devices.

cond-mat.mtrl-sci

MISA: Memory-Efficient LLMs Optimization with Module-wise Importance Sampling

The substantial memory demands of pre-training and fine-tuning large language models (LLMs) require memory-efficient optimization algorithms. One promising approach is layer-wise optimization, which treats each transformer block as a single layer and optimizes it sequentially, while freezing the other layers to save optimizer states and activations. Although effective, these methods ignore the varying importance of the modules within each layer, leading to suboptimal performance. Moreover, layer-wise sampling provides only limited memory savings, as at least one full layer must remain active during optimization. To overcome these limitations, we propose Module-wise Importance SAmpling (MISA), a novel method that divides each layer into smaller modules and assigns importance scores to each module. MISA uses a weighted random sampling mechanism to activate modules, provably reducing gradient variance compared to layer-wise sampling. Additionally, we establish an \(\mathcal{O}(1/\sqrt{K})\) convergence rate under non-convex and stochastic conditions, where $K$ is the total number of block updates, and provide a detailed memory analysis showcasing MISA's superiority over existing baseline methods. Experiments on diverse learning tasks validate the effectiveness of MISA. Source code is available at https://github.com/pkumelon/MISA.

cs.LG

RoleRMBench & RoleRM: Towards Reward Modeling for Profile-Based Role Play in Dialogue Systems

Reward modeling has become a cornerstone of aligning large language models (LLMs) with human preferences. Yet, when extended to subjective and open-ended domains such as role play, existing reward models exhibit severe degradation, struggling to capture nuanced and persona-grounded human judgments. To address this gap, we introduce RoleRMBench, the first systematic benchmark for reward modeling in role-playing dialogue, covering seven fine-grained capabilities from narrative management to role consistency and engagement. Evaluation on RoleRMBench reveals large and consistent gaps between general-purpose reward models and human judgment, particularly in narrative and stylistic dimensions. We further propose RoleRM, a reward model trained with Continuous Implicit Preferences (CIP), which reformulates subjective evaluation as continuous consistent pairwise supervision under multiple structuring strategies. Comprehensive experiments show that RoleRM surpasses strong open- and closed-source reward models by over 24% on average, demonstrating substantial gains in narrative coherence and stylistic fidelity. Our findings highlight the importance of continuous preference representation and annotation consistency, establishing a foundation for subjective alignment in human-centered dialogue systems.

cs.CL

SimDiff: Simpler Yet Better Diffusion Model for Time Series Point Forecasting

Diffusion models have recently shown promise in time series forecasting, particularly for probabilistic predictions. However, they often fail to achieve state-of-the-art point estimation performance compared to regression-based methods. This limitation stems from difficulties in providing sufficient contextual bias to track distribution shifts and in balancing output diversity with the stability and precision required for point forecasts. Existing diffusion-based approaches mainly focus on full-distribution modeling under probabilistic frameworks, often with likelihood maximization objectives, while paying little attention to dedicated strategies for high-accuracy point estimation. Moreover, other existing point prediction diffusion methods frequently rely on pre-trained or jointly trained mature models for contextual bias, sacrificing the generative flexibility of diffusion models. To address these challenges, we propose SimDiff, a single-stage, end-to-end framework. SimDiff employs a single unified Transformer network carefully tailored to serve as both denoiser and predictor, eliminating the need for external pre-trained or jointly trained regressors. It achieves state-of-the-art point estimation performance by leveraging intrinsic output diversity and improving mean squared error accuracy through multiple inference ensembling. Key innovations, including normalization independence and the median-of-means estimator, further enhance adaptability and stability. Extensive experiments demonstrate that SimDiff significantly outperforms existing methods in time series point forecasting.

cs.AI

S$^3$Attention: Improving Long Sequence Attention with Smoothed Skeleton Sketching

Attention based models have achieved many remarkable breakthroughs in numerous applications. However, the quadratic complexity of Attention makes the vanilla Attention based models hard to apply to long sequence tasks. Various improved Attention structures are proposed to reduce the computation cost by inducing low rankness and approximating the whole sequence by sub-sequences. The most challenging part of those approaches is maintaining the proper balance between information preservation and computation reduction: the longer sub-sequences used, the better information is preserved, but at the price of introducing more noise and computational costs. In this paper, we propose a smoothed skeleton sketching based Attention structure, coined S$^3$Attention, which significantly improves upon the previous attempts to negotiate this trade-off. S$^3$Attention has two mechanisms to effectively minimize the impact of noise while keeping the linear complexity to the sequence length: a smoothing block to mix information over long sequences and a matrix sketching method that simultaneously selects columns and rows from the input matrix. We verify the effectiveness of S$^3$Attention both theoretically and empirically. Extensive studies over Long Range Arena (LRA) datasets and six time-series forecasting show that S$^3$Attention significantly outperforms both vanilla Attention and other state-of-the-art variants of Attention structures.

cs.LG

Online Influence Maximization under Decreasing Cascade Model

We study online influence maximization (OIM) under a new model of decreasing cascade (DC). This model is a generalization of the independent cascade (IC) model by considering the common phenomenon of market saturation. In DC, the chance of an influence attempt being successful reduces with previous failures. The effect is neglected by previous OIM works under IC and linear threshold models. We propose the DC-UCB algorithm to solve this problem, which achieves a regret bound of the same order as the state-of-the-art works on the IC model. Extensive experiments on both synthetic and real datasets show the effectiveness of our algorithm.

cs.SI

FiLM: Frequency improved Legendre Memory Model for Long-term Time Series Forecasting

Recent studies have shown that deep learning models such as RNNs and Transformers have brought significant performance gains for long-term forecasting of time series because they effectively utilize historical information. We found, however, that there is still great room for improvement in how to preserve historical information in neural networks while avoiding overfitting to noise presented in the history. Addressing this allows better utilization of the capabilities of deep learning models. To this end, we design a \textbf{F}requency \textbf{i}mproved \textbf{L}egendre \textbf{M}emory model, or {\bf FiLM}: it applies Legendre Polynomials projections to approximate historical information, uses Fourier projection to remove noise, and adds a low-rank approximation to speed up computation. Our empirical studies show that the proposed FiLM significantly improves the accuracy of state-of-the-art models in multivariate and univariate long-term forecasting by (\textbf{20.3\%}, \textbf{22.6\%}), respectively. We also demonstrate that the representation module developed in this work can be used as a general plug-in to improve the long-term prediction performance of other deep learning modules. Code is available at https://github.com/tianzhou2011/FiLM/

cs.LG

Double-heavy tetraquark states with heavy diquark-antiquark symmetry

We calculate the masses of the $QQ\bar{q}\bar{q}$ ($Q=c,b$; $q=u,d,s$) tetraquark states with the aid of heavy diquark-antiquark symmetry (HDAS) and the chromomagnetic interaction (CMI) model. The masses of the highest-spin ($J=2$) tetraquarks that have only the $(QQ)_{\bar{3}_c}(\bar{q}\bar{q})_{3_c}$ color structure are related with those of conventional hadrons using HDAS. Thereafter, the masses of their partner states are determined with the mass splittings in the CMI model. Our numerical results reveal that: (i) the lightest $cc\bar{n}\bar{n}$ ($n=u,d$) is an $I(J^P)=0(1^+)$ state around 3929 MeV (53 MeV above the $DD^*$ threshold) and none of the double-charm tetraquarks are stable; (ii) the stable double-bottom tetraquarks are the lowest $0(1^+)$ $bb\bar{n}\bar{n}$ around 10488 MeV ($\approx116$ MeV below the $BB^*$ threshold) and the lowest $1/2(1^+)$ $bb\bar{n}\bar{s}$ around 10671 MeV ($\approx20$ MeV below the $BB_s^*/B_sB^*$ threshold); and (iii) the two lowest $bc\bar{n}\bar{n}$ tetraquarks, namely the lowest $0(0^+)$ around 7167 MeV and the lowest $0(1^+)$ around 7223 MeV, are near-threshold states. Moreover, we discuss the constraints on the masses of double-heavy hadrons. Specifically, for the lowest nonstrange tetraquarks, we obtain $T_{cc}<3965$ MeV, $T_{bb}<10627$ MeV, and $T_{bc}<7199$ MeV.

hep-ph

Developing Univariate Neurodegeneration Biomarkers with Low-Rank and Sparse Subspace Decomposition

Cognitive decline due to Alzheimer's disease (AD) is closely associated with brain structure alterations captured by structural magnetic resonance imaging (sMRI). It supports the validity to develop sMRI-based univariate neurodegeneration biomarkers (UNB). However, existing UNB work either fails to model large group variances or does not capture AD dementia (ADD) induced changes. We propose a novel low-rank and sparse subspace decomposition method capable of stably quantifying the morphological changes induced by ADD. Specifically, we propose a numerically efficient rank minimization mechanism to extract group common structure and impose regularization constraints to encode the original 3D morphometry connectivity. Further, we generate regions-of-interest (ROI) with group difference study between common subspaces of $Aβ+$ AD and $Aβ-$ cognitively unimpaired (CU) groups. A univariate morphometry index (UMI) is constructed from these ROIs by summarizing individual morphological characteristics weighted by normalized difference between $Aβ+$ AD and $Aβ-$ CU groups. We use hippocampal surface radial distance feature to compute the UMIs and validate our work in the Alzheimer's Disease Neuroimaging Initiative (ADNI) cohort. With hippocampal UMIs, the estimated minimum sample sizes needed to detect a 25$\%$ reduction in the mean annual change with 80$\%$ power and two-tailed $P=0.05$ are 116, 279 and 387 for the longitudinal $Aβ+$ AD, $Aβ+$ mild cognitive impairment (MCI) and $Aβ+$ CU groups, respectively. Additionally, for MCI patients, UMIs well correlate with hazard ratio of conversion to AD ($4.3$, $95\%$ CI=$2.3-8.2$) within 18 months. Our experimental results outperform traditional hippocampal volume measures and suggest the application of UMI as a potential UNB.

cs.CV

Spectrum and rearrangement decays of tetraquark states with four different flavors

We have systematically investigated the mass spectrum and rearrangement decay properties of the exotic tetraquark states with four different flavors using a color-magnetic interaction model. Their masses are estimated by assuming that the $X(4140)$ is a $cs\bar{c}\bar{s}$ tetraquark state and their decay widths are obtained by assuming that the Hamiltonian for decay is a constant. According to the adopted method, we find that the most stable states are probably the isoscalar $bs\bar{u}\bar{d}$ and $cs\bar{u}\bar{d}$ with $J^P=0^+$ and $1^+$. The width for most unstable tetraquarks is about tens of MeVs, but that for unstable $cu\bar{s}\bar{d}$ and $cs\bar{u}\bar{d}$ can be around 100 MeV. For the $X(5568)$, our method cannot give consistent mass and width if it is a $bu\bar{s}\bar{d}$ tetraquark state. For the $I(J^P)=0(0^+),0(1^+)$ double-heavy $T_{bc}=bc\bar{u}\bar{d}$ states, their widths can be several MeVs.

hep-ph

Exclusive Production Ratio of Neutral over Charged Kaon Pair in $e^+e^-$ Annihilation Continuum via `Straton Model'

A completely relativistic quark model in the Bethe-Salpter framework is employed to calculate the exclusive production ratio of the neutral over charged Kaon pair in $e^+e^-$ annihilation continuum region for center of mass energies smaller than the $J/Ψ$ mass. The valence quark charge plays the key rôle. The cancellation of the diagrams for the same charge case (in $K_S + K_L$) and the non-cancellation of the diagrams for the different charge case (in $K^-+K^+$) lead to the ratio as $(m_s-m_d)^2/M_{Kaon}^2 \sim 1/10$.

hep-ph

Efficient Discrete Supervised Hashing for Large-scale Cross-modal Retrieval

Supervised cross-modal hashing has gained increasing research interest on large-scale retrieval task owning to its satisfactory performance and efficiency. However, it still has some challenging issues to be further studied: 1) most of them fail to well preserve the semantic correlations in hash codes because of the large heterogenous gap; 2) most of them relax the discrete constraint on hash codes, leading to large quantization error and consequent low performance; 3) most of them suffer from relatively high memory cost and computational complexity during training procedure, which makes them unscalable. In this paper, to address above issues, we propose a supervised cross-modal hashing method based on matrix factorization dubbed Efficient Discrete Supervised Hashing (EDSH). Specifically, collective matrix factorization on heterogenous features and semantic embedding with class labels are seamlessly integrated to learn hash codes. Therefore, the feature based similarities and semantic correlations can be both preserved in hash codes, which makes the learned hash codes more discriminative. Then an efficient discrete optimal algorithm is proposed to handle the scalable issue. Instead of learning hash codes bit-by-bit, hash codes matrix can be obtained directly which is more efficient. Extensive experimental results on three public real-world datasets demonstrate that EDSH produces a superior performance in both accuracy and scalability over some existing cross-modal hashing methods.

cs.LG