SearcharxivSearch

arXiv subjects

Xiaoyang Tan

Publications and source records attributed to Xiaoyang Tan.

At least 19 recordsLinked to original sources

HiQA: A Hierarchical Contextual Augmentation RAG for Multi-Documents QA

Retrieval-Augmented Generation (RAG) significantly improves document-based question answering by integrating external documents during generation. However, retrieval accuracy can degrade when the knowledge base contains many semantically and structurally similar documents. We introduce HiQA, a practical hierarchical contextual augmentation framework for multi-document question answering (MDQA). HiQA enriches text chunks with cascading document metadata, such as document titles and section paths, so that retrieval can use both local content and document structure. The framework also uses a multi-route retriever that combines semantic, lexical, and keyword/entity signals. We further introduce MasQA, a benchmark designed to evaluate MDQA systems in realistic similar-document settings. Experiments show that HiQA improves retrieval and answer quality on MasQA and remains competitive on public MDQA benchmarks, while its benefits are strongest for structured, domain-specific, highly similar document collections.

cs.CL

Saturated and Anisotropic Magnetostriction in an Altermagnet

Magnetostriction, a fundamental phenomenon bridging magnetism and mechanics, has enabled a broad spectrum of applications. For almost two centuries, it has been mainly investigated for ferromagnets. Regarding the magnetostriction of antiferromagnets (AFMs), limitedly known examples for both conventional collinear AFMs and noncollinear AFMs predominantly exhibit non-saturating magnetic-field dependence. Herein, we report an easily saturated magnetostriction effect in a prototypical altermagnet - MnTe, which is an emerging class of collinear AFMs with special crystal symmetries. For high-quality MnTe single crystals, the magnetostriction saturates under a moderate field of ~0.7 T with an intriguing two-fold-symmetry anisotropy. First-principles calculations reveal that the saturated and anisotropic magnetostriction originates from symmetry-allowed coupling between elastic strain and its Néel order parameter. These findings break the traditional wisdom on antiferromagnetic magnetostriction.

cond-mat.mtrl-sci

Giant Room-Temperature Third-Order Electrical Transport in a Thin-Film Altermagnet Candidate

Quantum geometry, a quantum mechanical quantity comprised of Berry curvature and quantum metric, describes the geometric structure of the electronic bands in solids. The correlation between nontrivial quantum geometry and quantum materials leads to new findings in condensed matter systems. Here we demonstrate that altermagnets, with spontaneously broken time-reversal (T)- half-lattice-translation and parity-time symmetry, host both T-odd and T-even quantum geometric quantities that simultaneously manifest themselves despite the vanishing net magnetization. Consequently, giant room-temperature third-order electrical transport responses with sizable quantum geometric contributions are observed in (101)-oriented RuO2 thin films, an altermagnetic candidate; in particular, the third-order Hall effect is intimately correlated with altermagnetic order and can serve as a promising tool for detecting the Neel vector. Our work not only supports the existence of altermagnetism in 8-nm-thick RuO2 thin films, but also shows altermagnets as a versatile platform for exploring quantum geometry and constructing quantum electronic and spintronic devices.

cond-mat.mes-hall

Bulk OsO2 Single Crystals: Superior Catalysts for Water Oxidation

Although rutile RuO2 has been a well-known and almost the best oxygen evolution reaction (OER) catalyst, the OER properties for the similar rutile oxide OsO2 with the same group element with Ru have been unknown, mainly due to long-standing synthesis difficulties. In this work, we report the successful synthesis of high-quality OsO2 single crystals, and the ground micrometer-size single crystals are chemically stable in alkaline solutions and exhibit robust OER performance. In sharp contrast, OsO2 nanopowder reacts quickly with KOH solutions and cannot work for OER. Compared with commercial RuO2 nanopowder, the OsO2 single crystals show comparable catalytic current densities, remarkably lower overpotentials at high current densities and better stability. These findings question the universal applicability of nanoscaling and highlight crystal integrity as a key descriptor for achieving stable and efficient OER electrocatalysis.

cond-mat.mtrl-sci

Pressure-Induced Metal-Insulator and Paramagnet-Altermagnet Transitions in Rutile OsO2 Single Crystals

Altermagnets with compensated spin structures and nonrelativistic spin splitting have emerged as a new class of magnetic materials. Rutile OsO2 has been theoretically predicted to be altermagnetic, but experimental studies have been limited by synthesis challenges. We have succeeded in synthesizing high-quality single crystals of rutile OsO2. Electrical transport studies reveal that OsO2 is highly conductive and exhibits clear Fermi liquid behavior, indicating strong electron-electron scattering. Magnetic measurements show that the crystals are isotropically paramagnetic. Density-functional theory calculations indicate that bulk OsO2 is semimetallic with coexisting electron and hole pockets, with its magnetic ground state strongly dependent on the on-site Coulomb correlation U. Angle-resolved photoemission spectroscopy studies unveil that the bulk bands do not yet show altermagnetic spin splitting. Interestingly, resistivity is rather pressure sensitive: at 44 GPa, a clear metal-insulator transition occurs. Hybrid functional calculations reveal that applying pressure significantly increases the Hubbard U value, driving a phase transition from a paramagnetic metal to an altermagnetic metal, and eventually to an altermagnetic insulator. These findings suggest that tuning external pressure effectively modulates the magnetic ground state of OsO2, providing a pathway to realize altermagnetism in this material.

cond-mat.mes-hall

Observation of Orbit-Orbit Torques: Highly Efficient Torques on Orbital Moments Induced by Orbital Currents

We study the current-induced torques in bilayers composed of a light 3d metal, chromium, and a rare-earth ferromagnet with finite orbital moments, terbium, utilizing second-harmonic Hall-response measurements. The dampinglike torque efficiency of chromium is found to be positive and reaches ~3.66 in this system, in sharp contrast to the negative and subtle dampinglike torque efficiency in general Cr/ferromagnet heterostructures with quenched orbital moment. We suggest that the orbital currents generated by the orbital Hall effect in Cr can be injected into Tb with negligible loss at the interface and then efficiently interact with the orbital moments. We term such an exotic effect as the orbit-orbit torque (OOT). Our work implies that orbital currents could be harnessed to manipulate the orbital magnetization of materials, which would advance the development of orbitronics.

cond-mat.mes-hall

Electric-Field-Controlled Altermagnetic Transition for Neuromorphic Computing

Altermagnets represent a novel magnetic phase with transformative potential for ultrafast spintronics, yet efficient control of their magnetic states remains challenging. We demonstrate an ultra-low-power electric-field control of altermagnetism in MnTe through strain-mediated coupling in MnTe/PMN-PT heterostructures with negligible Joule heating. Application of +6 kV/cm electric fields induces piezoelectric strain in PMN-PT, modulating the Néel temperature from 310 to 328 K. As a result, around the magnetic phase transition, the altermagnetic spin splitting of MnTe is reversibly switched "on" and "off" by the electric fields. Meanwhile, the piezoelectric strain generates lattice distortions and magnetic structure changes in MnTe, enabling up to 9.7% resistance modulation around the magnetic phase transition temperature. Leveraging this effect, we implement programmable resistance states in a Hopfield neuromorphic network, achieving 100% pattern recognition accuracy at <=40% noise levels. This approach establishes the electric-field control as a low-power strategy for altermagnetic manipulation while demonstrating the viability of altermagnetic materials for energy-efficient neuromorphic computing beyond conventional charge-based architectures.

cond-mat.mtrl-sci

Freestanding Thin-Film Materials

Freestanding thin films, a class of low-dimensional materials capable of maintaining structural integrity without substrates, have emerged as a forefront research focus. Their unique advantages-circumventing substrate clamping, liberating intrinsic material properties, and enabling cross-platform heterogeneous integration-underpin this prominence. This review systematically summarizes core fabrication techniques, including physical delamination (e.g., laser lift-off, mechanical exfoliation) and chemical etching, alongside associated transfer strategies. It further explores the induced strain modulation mechanisms, extreme mechanical properties and interface decoupling effects enabled by these films. Representative case studies demonstrate breakthrough applications in flexible/ultrathin electronics, ultrahigh-sensitivity sensors and the exploration of novel quantum states. Critical challenges regarding scalable fabrication, precise interface control, and long-term stability are analyzed, concluding with prospects for emerging applications in bio-inspired intelligent devices, quantum precision sensing, and brain-inspired neural networks.

cond-mat.mtrl-sci

Variational OOD State Correction for Offline Reinforcement Learning

The performance of Offline reinforcement learning is significantly impacted by the issue of state distributional shift, and out-of-distribution (OOD) state correction is a popular approach to address this problem. In this paper, we propose a novel method named Density-Aware Safety Perception (DASP) for OOD state correction. Specifically, our method encourages the agent to prioritize actions that lead to outcomes with higher data density, thereby promoting its operation within or the return to in-distribution (safe) regions. To achieve this, we optimize the objective within a variational framework that concurrently considers both the potential outcomes of decision-making and their density, thus providing crucial contextual information for safe decision-making. Finally, we validate the effectiveness and feasibility of our proposed method through extensive experimental evaluations on the offline MuJoCo and AntMaze suites.

cs.LG

RoGA: Towards Generalizable Deepfake Detection through Robust Gradient Alignment

Recent advancements in domain generalization for deepfake detection have attracted significant attention, with previous methods often incorporating additional modules to prevent overfitting to domain-specific patterns. However, such regularization can hinder the optimization of the empirical risk minimization (ERM) objective, ultimately degrading model performance. In this paper, we propose a novel learning objective that aligns generalization gradient updates with ERM gradient updates. The key innovation is the application of perturbations to model parameters, aligning the ascending points across domains, which specifically enhances the robustness of deepfake detection models to domain shifts. This approach effectively preserves domain-invariant features while managing domain-specific characteristics, without introducing additional regularization. Experimental results on multiple challenging deepfake detection datasets demonstrate that our gradient alignment strategy outperforms state-of-the-art domain generalization techniques, confirming the efficacy of our method. The code is available at https://github.com/Lynn0925/RoGA.

cs.CV

Contrastive Desensitization Learning for Cross Domain Face Forgery Detection

In this paper, we propose a new cross-domain face forgery detection method that is insensitive to different and possibly unseen forgery methods while ensuring an acceptable low false positive rate. Although existing face forgery detection methods are applicable to multiple domains to some degree, they often come with a high false positive rate, which can greatly disrupt the usability of the system. To address this issue, we propose an Contrastive Desensitization Network (CDN) based on a robust desensitization algorithm, which captures the essential domain characteristics through learning them from domain transformation over pairs of genuine face images. One advantage of CDN lies in that the learnt face representation is theoretical justified with regard to the its robustness against the domain changes. Extensive experiments over large-scale benchmark datasets demonstrate that our method achieves a much lower false alarm rate with improved detection accuracy compared to several state-of-the-art methods.

cs.CV

Beyond Non-Expert Demonstrations: Outcome-Driven Action Constraint for Offline Reinforcement Learning

We address the challenge of offline reinforcement learning using realistic data, specifically non-expert data collected through sub-optimal behavior policies. Under such circumstance, the learned policy must be safe enough to manage distribution shift while maintaining sufficient flexibility to deal with non-expert (bad) demonstrations from offline data.To tackle this issue, we introduce a novel method called Outcome-Driven Action Flexibility (ODAF), which seeks to reduce reliance on the empirical action distribution of the behavior policy, hence reducing the negative impact of those bad demonstrations.To be specific, a new conservative reward mechanism is developed to deal with distribution shift by evaluating actions according to whether their outcomes meet safety requirements - remaining within the state support area, rather than solely depending on the actions' likelihood based on offline data.Besides theoretical justification, we provide empirical evidence on widely used MuJoCo and various maze benchmarks, demonstrating that our ODAF method, implemented using uncertainty quantification techniques, effectively tolerates unseen transitions for improved "trajectory stitching," while enhancing the agent's ability to learn from realistic non-expert data.

cs.LG

Transductive Off-policy Proximal Policy Optimization

Proximal Policy Optimization (PPO) is a popular model-free reinforcement learning algorithm, esteemed for its simplicity and efficacy. However, due to its inherent on-policy nature, its proficiency in harnessing data from disparate policies is constrained. This paper introduces a novel off-policy extension to the original PPO method, christened Transductive Off-policy PPO (ToPPO). Herein, we provide theoretical justification for incorporating off-policy data in PPO training and prudent guidelines for its safe application. Our contribution includes a novel formulation of the policy improvement lower bound for prospective policies derived from off-policy data, accompanied by a computationally efficient mechanism to optimize this bound, underpinned by assurances of monotonic improvement. Comprehensive experimental results across six representative tasks underscore ToPPO's promising performance.

cs.LG

Highway Reinforcement Learning

Learning from multi-step off-policy data collected by a set of policies is a core problem of reinforcement learning (RL). Approaches based on importance sampling (IS) often suffer from large variances due to products of IS ratios. Typical IS-free methods, such as $n$-step Q-learning, look ahead for $n$ time steps along the trajectory of actions (where $n$ is called the lookahead depth) and utilize off-policy data directly without any additional adjustment. They work well for proper choices of $n$. We show, however, that such IS-free methods underestimate the optimal value function (VF), especially for large $n$, restricting their capacity to efficiently utilize information from distant future time steps. To overcome this problem, we introduce a novel, IS-free, multi-step off-policy method that avoids the underestimation issue and converges to the optimal VF. At its core lies a simple but non-trivial \emph{highway gate}, which controls the information flow from the distant future by comparing it to a threshold. The highway gate guarantees convergence to the optimal VF for arbitrary $n$ and arbitrary behavioral policies. It gives rise to a novel family of off-policy RL algorithms that safely learn even when $n$ is very large, facilitating rapid credit assignment from the far future to the past. On tasks with greatly delayed rewards, including video games where the reward is given only at the end of the game, our new methods outperform many existing multi-step off-policy algorithms.

cs.LG

ProxyFormer: Proxy Alignment Assisted Point Cloud Completion with Missing Part Sensitive Transformer

Problems such as equipment defects or limited viewpoints will lead the captured point clouds to be incomplete. Therefore, recovering the complete point clouds from the partial ones plays an vital role in many practical tasks, and one of the keys lies in the prediction of the missing part. In this paper, we propose a novel point cloud completion approach namely ProxyFormer that divides point clouds into existing (input) and missing (to be predicted) parts and each part communicates information through its proxies. Specifically, we fuse information into point proxy via feature and position extractor, and generate features for missing point proxies from the features of existing point proxies. Then, in order to better perceive the position of missing points, we design a missing part sensitive transformer, which converts random normal distribution into reasonable position information, and uses proxy alignment to refine the missing proxies. It makes the predicted point proxies more sensitive to the features and positions of the missing part, and thus make these proxies more suitable for subsequent coarse-to-fine processes. Experimental results show that our method outperforms state-of-the-art completion networks on several benchmark datasets and has the fastest inference speed. Code is available at https://github.com/I2-Multimedia-Lab/ProxyFormer.

cs.CV

Contextual Conservative Q-Learning for Offline Reinforcement Learning

Offline reinforcement learning learns an effective policy on offline datasets without online interaction, and it attracts persistent research attention due to its potential of practical application. However, extrapolation error generated by distribution shift will still lead to the overestimation for those actions that transit to out-of-distribution(OOD) states, which degrades the reliability and robustness of the offline policy. In this paper, we propose Contextual Conservative Q-Learning(C-CQL) to learn a robustly reliable policy through the contextual information captured via an inverse dynamics model. With the supervision of the inverse dynamics model, it tends to learn a policy that generates stable transition at perturbed states, for the fact that pertuebed states are a common kind of OOD states. In this manner, we enable the learnt policy more likely to generate transition that destines to the empirical next state distributions of the offline dataset, i.e., robustly reliable transition. Besides, we theoretically reveal that C-CQL is the generalization of the Conservative Q-Learning(CQL) and aggressive State Deviation Correction(SDC). Finally, experimental results demonstrate the proposed C-CQL achieves the state-of-the-art performance in most environments of offline Mujoco suite and a noisy Mujoco setting.

cs.LG

Smoothing Advantage Learning

Advantage learning (AL) aims to improve the robustness of value-based reinforcement learning against estimation errors with action-gap-based regularization. Unfortunately, the method tends to be unstable in the case of function approximation. In this paper, we propose a simple variant of AL, named smoothing advantage learning (SAL), to alleviate this problem. The key to our method is to replace the original Bellman Optimal operator in AL with a smooth one so as to obtain more reliable estimation of the temporal difference target. We give a detailed account of the resulting action gap and the performance bound for approximate SAL. Further theoretical analysis reveals that the proposed value smoothing technique not only helps to stabilize the training procedure of AL by controlling the trade-off between convergence rate and the upper bound of the approximation errors, but is beneficial to increase the action gap between the optimal and sub-optimal action value as well.

cs.LG

Robust Action Gap Increasing with Clipped Advantage Learning

Advantage Learning (AL) seeks to increase the action gap between the optimal action and its competitors, so as to improve the robustness to estimation errors. However, the method becomes problematic when the optimal action induced by the approximated value function does not agree with the true optimal action. In this paper, we present a novel method, named clipped Advantage Learning (clipped AL), to address this issue. The method is inspired by our observation that increasing the action gap blindly for all given samples while not taking their necessities into account could accumulate more errors in the performance loss bound, leading to a slow value convergence, and to avoid that, we should adjust the advantage value adaptively. We show that our simple clipped AL operator not only enjoys fast convergence guarantee but also retains proper action gaps, hence achieving a good balance between the large action gap and the fast convergence. The feasibility and effectiveness of the proposed method are verified empirically on several RL benchmarks with promising performance.

cs.LG