SearcharxivSearch

arXiv subjects

Tao Qin

Publications and source records attributed to Tao Qin.

At least 19 recordsLinked to original sources

HOPE: Heterophily-Aware Open-Set Node Classification with Pseudo-Extrapolation

Standard open-set node classification methods rely on the homophily assumption, where connected nodes share labels. However, real-world graphs are often heterophilic, exposing the limitations of current methods and posing new challenges to open-set node classification. On the one hand, cross-class connectivity causes representations from different known or unknown classes to become intertwined after aggregation, undermining their discriminative capacity. On the other hand, structural mixture invalidates threshold-based open-set methods and cross-class feature interpolation, leading to unreliable unknown-class rejection. To address these challenges, we propose HOPE, a Heterophily-aware Open-set node classification method with Pseudo-Extrapolation. To adapt open-set graph neural networks (GNNs) to heterophilic scenarios, HOPE uses a structure-augmented feature initialization layer to capture multi-hop structural patterns. Meanwhile, we design a trustworthy neighborhood aggregation mechanism for standard GNNs to dynamically filter noisy cross-class neighbors. To enhance unknown-class rejection, we introduce a heterophily-guided pseudo-extrapolation strategy. It dynamically maintains known-class centers and extrapolates along cross-class neighborhood displacement directions, synthesizing pseudo-unknown proxies near structurally ambiguous regions. Finally, we optimize the network with joint classification and logit margin regularization, routing synthetic proxies into a dedicated rejection slot without imposing geometric margin constraints in the representation space. Extensive experiments on multiple datasets show that HOPE consistently outperforms state-of-the-art models, validating its effectiveness, robustness, and efficiency.

cs.LG

Circular phonon dichroism in $d$-wave altermagnets

Altermagnets, a new class of collinear antiferromagnets, exhibit momentum-dependent spin splitting and offer compelling advantages for antiferromagnetic spintronics. However, the magnetic order is intrinsically difficult to read out, which hinders practical applications. We propose finite-momentum circular phonon dichroism as a direct probe of N\'eel vector in two-dimensional $d$-wave altermagnets. Combining Onsager reciprocity with $C_{2z}$ lattice symmetry, we find that the dichroic signal reverses sign when the N\'eel vector is flipped for the in-plane phonon wave vectors. Moreover, a channel-resolved decomposition identifies the circular phonon dichroism originates from the interband coherent transitions. Representative finite-momentum cuts show pronounced dichroic asymmetric ratio, with $|\eta_{\mathrm{CPD}}|=37.3\%$. Our work reveals that the circular-ultrasound absorption acts as a direct probe of the N\'eel vector of $d$-wave altermagnets.

cond-mat.mes-hall

Effect of Two-Body Interactions on Floquet topological phases

We study the circularly driven Falicov-Kimball model on a honeycomb lattice within real space Floquet dynamical mean field theory (DMFT). The noninteracting version of this model has been realized experimentally. The noninteracting system hosts an effective Haldane phase at large driving frequencies, while at intermediate frequencies it hosts an anomalous topological phase. We study the effect of two-body interactions $U$ on the stability of these phases. We find that charge pumping does not remain quantized upon increasing $U$, despite the presence of edge modes in the spectrum. This can be attributed to the broadening of the edge modes due to interaction. We also calculate the rate of energy dissipation into the bath and find remarkably different behaviour in the two regimes.

cond-mat.quant-gas

Electronic Hall viscosity: hidden indicator for antiferromagnets

The antiferromagnets with negligible stray fields and ultrafast spin dynamics play a crucial role in the fields of energy-efficient spintronics and topological electronics. However, the detection and control of the underlying nontrivial Berry curvature become extremely limited by the vanishing magnetization and anomalous Hall conductivity. Here, we show the electronic Hall viscosity is closely related to the quadruple Berry curvature of Bloch bands and is bounded by the $d$-orbit factor modulated second moment of the quantum volume. Moreover, we derive the symmetry requirement for nonzero electronic Hall viscosity that could characterize antiferromagnetic ordering even when the linear anomalous Hall response gets forbidden. We further examine our key findings in two archetypal antiferromagnets: $d$-wave altermagnet $\mathrm{RuO}_{2}$, and noncollinear $\mathrm{Mn_{3}Sn}$ through direct first-principle calculations. Thus, our work reveals a new and fundamental quantum geometry quantity of generic antiferromagnets and offers a broadly applicable way to design antiferromagnetic spintronics devices via unconventional Hall viscosity.

cond-mat.mes-hall

Accelerating Locality-Driven Integration in Quantum Chemistry with Block-Structured Matrix Multiplication

Locality-driven integration is a pervasive computational pattern in quantum chemistry, arising whenever spatially localized basis functions interact through numerical quadrature or integral screening. The dominant matrix multiplications in these tasks exhibit dynamic, structured sparsity driven by spatial locality, posing significant challenges for both dense batched kernels and generic sparse formats on GPUs. We present KerneLDI, a GPU-oriented framework that addresses this regime by co-designing data layout, screening logic, and matrix-computation operators to realize block-structured matrix multiplication for locality-driven integration. KerneLDI reorganizes operand matrices into a unified block-filtered representation that retains only spatially relevant blocks, and executes the resulting contractions with customized dense block multipliers that adapt proven dense-matmul optimizations to retained block pairs. We develop and evaluate KerneLDI on exchange--correlation (EXC) integration in Kohn--Sham density functional theory, a representative and computationally critical instance of this pattern. Across diverse molecular systems, KerneLDI preserves numerical accuracy while delivering up to 10$\times$ speedup for EXC evaluation over a dense GPU baseline, scales favorably with increasing system size and multi-GPU parallelism, accelerates end-to-end self-consistent field calculations, and yields nearly 6$\times$ throughput improvement for ab initio molecular dynamics.

physics.comp-ph

Stingray Patterns of Dominant Weights

We study the set $W_{r,e,w}\ $ of dominant weights of $\mathfrak{sl}_r$ arising from partitions of fixed $e$-weight $w$. For $e$-cores, we show that $W_{r,e,0}\ $ decomposes as a disjoint union of simplices indexed by compositions of $r$. For general $w$, we prove that $W_{r,e,w}\ $ is a disjoint union of copies of these simplices, with multiplicities determined by the corresponding quotient data, yielding in particular a closed counting formula for $|W_{r,e,w}\ |\ $. The geometry gives rise to the stingray patterns appearing in the title. More generally, it yields a natural labeling of the dominant $e$-alcoves meeting $W_{r,e,w}\ $ by weak compositions of $w$, together with a compatible partial action of the affine Weyl group via wall crossing. Finally, we give an explicit alcove-geometric proof of the empty runner removal theorem for Iwahori-Hecke algebras.

math.CO

Deciphering Scientific Reasoning Steps from Outcome Data for Molecule Optimization

Emerging reasoning models hold promise for automating scientific discovery. However, their training is hindered by a critical supervision gap: experimental outcomes are abundant, whereas intermediate reasoning steps are rarely documented at scale. To bridge this gap, we propose DESRO, a framework for deciphering scientific reasoning from outcomes. By analyzing shared patterns and key differences within grouped data, a large language model (LLM) can recover the underlying logic. We instantiate this framework in molecule optimization, a pivotal stage in drug discovery that traditionally relies on the iterative reasoning of medicinal chemists. Across 2.3 million molecular property records, our framework infers optimization rationales by grouping molecules with shared fragments, then using an LLM to analyze how structural variations correlate with property differences. Based on the derived data, we train a model that conducts molecule optimization through an interpretable reasoning process. DESRO achieves the highest success rates on 15 out of 18 tasks, spanning both single- and multi-property optimization of bioactivity and ADMET properties. The reasoning process enables robust generalization to out-of-distribution scenarios, including novel property combinations, unseen biological targets, and unseen properties defined solely by natural language descriptions. In retrospective case studies under strict temporal splits, the model autonomously reconstructs expert-level lead optimization trajectories. Additionally, our framework extends beyond molecule optimization to reaction ligand selection. Our results establish deciphering reasoning steps from outcome data as a viable paradigm for enabling scientific reasoning, providing a scalable approach to accelerate scientific discovery.

q-bio.BM

Subdivision and Runner Removal Theorems

We develop a combinatorial framework for the subdivision map -- introduced by Maksimau, Mathas and Tubbenhauer -- between the KLR(W) algebras of type $A^{(1)}_{e-1}$ and type $A^{(1)}_{e}$, which provides a partial categorification of the runner removal theorems.

math.RT

R2LED: Equipping Retrieval and Refinement in Lifelong User Modeling with Semantic IDs for CTR Prediction

Lifelong user modeling, which leverages users' long-term behavior sequences for CTR prediction, has been widely applied in personalized services. Existing methods generally adopted a two-stage "retrieval-refinement" strategy to balance effectiveness and efficiency. However, they still suffer from (i) noisy retrieval due to skewed data distribution and (ii) lack of semantic understanding in refinement. While semantic enhancement, e.g., LLMs modeling or semantic embeddings, offers potential solutions to these two challenges, these approaches face impractical inference costs or insufficient representation granularity. Obsorbing multi-granularity and lightness merits of semantic identity (SID), we propose a novel paradigm that equips retrieval and refinement in Lifelong User Modeling with SEmantic IDs (R2LED) to address these issues. First, we introduce a Multi-route Mixed Retrieval for the retrieval stage. On the one hand, it captures users' interests from various granularities by several parallel recall routes. On the other hand, a mixed retrieval mechanism is proposed to efficiently retrieve candidates from both collaborative and semantic views, reducing noise. Then, for refinement, we design a Bi-level Fusion Refinement, including a target-aware cross-attention for route-level fusion and a gate mechanism for SID-level fusion. It can bridge the gap between semantic and collaborative spaces, exerting the merits of SID. The comprehensive experimental results on two public datasets demonstrate the superiority of our method in both performance and efficiency. To facilitate the reproduction, we have released the code online https://github.com/abananbao/R2LED.

cs.IR

TIDE: Trajectory-based Diagnostic Evaluation of Test-Time Improvement in LLM Agents

Recent advances in autonomous LLM agents demonstrate their ability to improve performance through iterative interaction with the environment. We define this paradigm as Test-Time Improvement (TTI). However, the mechanisms under how and why TTI succeed or fail remain poorly understood, and existing evaluation metrics fail to capture their task optimization efficiency, behavior adaptation after erroneous actions, and the specific utility of working memory for task completion. To address these gaps, we propose Test-time Improvement Diagnostic Evaluation (TIDE), an agent-agnostic and environment-agnostic framework that decomposes TTI into three comprehensive and interconnected dimensions. The framework measures (1) the overall temporal dynamics of task completion and (2) identifies whether performance is primarily constrained by recursive looping behaviors or (3) by burdensome accumulated memory. Through extensive experiments across diverse agents and environments, TIDE highlights that improving agent performance requires more than scaling internal reasoning, calling for explicitly optimizing the interaction dynamics between the agent and the environment.

cs.AI

Scalable Machine Learning Force Fields for Macromolecular Systems Through Long-Range Aware Message Passing

Machine learning force fields (MLFFs) have revolutionized molecular simulations by providing quantum mechanical accuracy at the speed of molecular mechanical computations. However, a fundamental reliance of these models on fixed-cutoff architectures limits their applicability to macromolecular systems where long-range interactions dominate. We demonstrate that this locality constraint causes force prediction errors to scale monotonically with system size, revealing a critical architectural bottleneck. To overcome this, we establish the systematically designed MolLR25 ({Mol}ecules with {L}ong-{R}ange effect) benchmark up to 1200 atoms, generated using high-fidelity DFT, and introduce E2Former-LSR, an equivariant transformer that explicitly integrates long-range attention blocks. E2Former-LSR exhibits stable error scaling, achieves superior fidelity in capturing non-covalent decay, and maintains precision on complex protein conformations. Crucially, its efficient design provides up to 30% speedup compared to purely local models. This work validates the necessity of non-local architectures for generalizable MLFFs, enabling high-fidelity molecular dynamics for large-scale chemical and biological systems.

physics.chem-ph

Phonon Dichroisms Revealing Unusual Electronic Quantum Geometry

The quantum geometry tensor, intrinsic geometric characteristics of electronic states, plays a crucial role in the various nontrivial electromagnetic phenomena in quantum materials. Here, we reveal that quantum geometry significantly modifies phonon dichroisms through electron-phonon interactions in solids that break time-reversal and spatial inversion symmetries. Specifically, the circular phonon dichroism is primarily dominated by the heat magnetic moments, while the linear phonon dichroism depends on the heat Drude weight, a thermal analog of band Drude weight. Furthermore, we establish the f-sum rule for the heat magnetic moment that facilitates its experimental detections. We demonstrate our key findings in an archetypal model system: ferromagnetic two-dimensional electron gases with Rashba spin-orbit coupling. Our work uncovers the quantum-geometric origin of common phonon dichroisms and predicts the detectable signature of the heat magnetic moment of electrons in solids.

cond-mat.mes-hall

MolChord: Structure-Sequence Alignment for Protein-Guided Drug Design

Structure-based drug design (SBDD), which maps target proteins to candidate molecular ligands, is a fundamental task in drug discovery. Effectively aligning protein structural representations with molecular representations, and ensuring alignment between generated drugs and their pharmacological properties, remains a critical challenge. To address these challenges, we propose MolChord, which integrates two key techniques: (1) to align protein and molecule structures with their textual descriptions and sequential representations (e.g., FASTA for proteins and SMILES for molecules), we leverage NatureLM, an autoregressive model unifying text, small molecules, and proteins, as the molecule generator, alongside a diffusion-based structure encoder; and (2) to guide molecules toward desired properties, we curate a property-aware dataset by integrating preference data and refine the alignment process using Direct Preference Optimization (DPO). Experimental results on CrossDocked2020 demonstrate that our approach achieves state-of-the-art performance on key evaluation metrics, highlighting its potential as a practical tool for SBDD.

cs.AI

A Specht Filtration of Permutation Modules Over KLR Algebras

In type A, Kleshchev-Ram-Mathas realize Specht modules as quotient of Permutation modules, in this paper, we construct a Specht filtration of Permutation modules indexed by hook partition in affine type A; and construct a generalized Specht filtration of Permutation modules indexed by any partition in linear quiver case.

math.RT

Instructing Text-to-Image Diffusion Models via Classifier-Guided Semantic Optimization

Text-to-image diffusion models have emerged as powerful tools for high-quality image generation and editing. Many existing approaches rely on text prompts as editing guidance. However, these methods are constrained by the need for manual prompt crafting, which can be time-consuming, introduce irrelevant details, and significantly limit editing performance. In this work, we propose optimizing semantic embeddings guided by attribute classifiers to steer text-to-image models toward desired edits, without relying on text prompts or requiring any training or fine-tuning of the diffusion model. We utilize classifiers to learn precise semantic embeddings at the dataset level. The learned embeddings are theoretically justified as the optimal representation of attribute semantics, enabling disentangled and accurate edits. Experiments further demonstrate that our method achieves high levels of disentanglement and strong generalization across different domains of data.

cs.CV

Chain-of-Model Learning for Language Model

In this paper, we propose a novel learning paradigm, termed Chain-of-Model (CoM), which incorporates the causal relationship into the hidden states of each layer as a chain style, thereby introducing great scaling efficiency in model training and inference flexibility in deployment. We introduce the concept of Chain-of-Representation (CoR), which formulates the hidden states at each layer as a combination of multiple sub-representations (i.e., chains) at the hidden dimension level. In each layer, each chain from the output representations can only view all of its preceding chains in the input representations. Consequently, the model built upon CoM framework can progressively scale up the model size by increasing the chains based on the previous models (i.e., chains), and offer multiple sub-models at varying sizes for elastic inference by using different chain numbers. Based on this principle, we devise Chain-of-Language-Model (CoLM), which incorporates the idea of CoM into each layer of Transformer architecture. Based on CoLM, we further introduce CoLM-Air by introducing a KV sharing mechanism, that computes all keys and values within the first chain and then shares across all chains. This design demonstrates additional extensibility, such as enabling seamless LM switching, prefilling acceleration and so on. Experimental results demonstrate our CoLM family can achieve comparable performance to the standard Transformer, while simultaneously enabling greater flexiblity, such as progressive scaling to improve training efficiency and offer multiple varying model sizes for elastic inference, paving a a new way toward building language models. Our code will be released in the future at: https://github.com/microsoft/CoLM.

cs.CL

MoonCast: High-Quality Zero-Shot Podcast Generation

Recent advances in text-to-speech synthesis have achieved notable success in generating high-quality short utterances for individual speakers. However, these systems still face challenges when extending their capabilities to long, multi-speaker, and spontaneous dialogues, typical of real-world scenarios such as podcasts. These limitations arise from two primary challenges: 1) long speech: podcasts typically span several minutes, exceeding the upper limit of most existing work; 2) spontaneity: podcasts are marked by their spontaneous, oral nature, which sharply contrasts with formal, written contexts; existing works often fall short in capturing this spontaneity. In this paper, we propose MoonCast, a solution for high-quality zero-shot podcast generation, aiming to synthesize natural podcast-style speech from text-only sources (e.g., stories, technical reports, news in TXT, PDF, or Web URL formats) using the voices of unseen speakers. To generate long audio, we adopt a long-context language model-based audio modeling approach utilizing large-scale long-context speech data. To enhance spontaneity, we utilize a podcast generation module to generate scripts with spontaneous details, which have been empirically shown to be as crucial as the text-to-speech modeling itself. Experiments demonstrate that MoonCast outperforms baselines, with particularly notable improvements in spontaneity and coherence.

eess.AS

Probing the Limit of Heat Transfer in Inorganic Crystals with Deep Learning

Heat transfer is a fundamental property of matter. Research spanning decades has attempted to discover materials with exceptional thermal conductivity, yet the upper limit remains unknown. Using deep learning accelerated crystal structure prediction and first-principles calculation, we systematically explore the thermal conductivity landscape of inorganic crystals. We brute-force over half a million ordered crystalline structures, encompassing an extensive coverage of local energy minima in binary compounds with up to four atoms per primitive cell. We confirm diamond sets the upper bound of thermal conductivity within our search space, very likely also among all stable crystalline solids at ambient conditions. We also identify over 20 novel crystals surpassing silicon in thermal conductivity, validated by density functional theory. These include a semiconductor TaN with ultrahigh thermal conductivity (~900 $\mathrm{W\cdot m^{-1}\cdot K^{-1}}$), and metallic compounds such as MnV that exhibit high lattice and electronic thermal conductivity simultaneously, a distinctive feature not observed before. These results as well as the deep learning-driven screening method, redefine the landscape of thermal transport and establish a large open-access database for future materials discovery.

cond-mat.mtrl-sci