Searcharxiv⌕ Search

arXiv subjects

Zhiheng Huang

Publications and source records attributed to Zhiheng Huang.

At least 19 recordsLinked to original sources

Programmable, Spontaneous Superlattice Memory in a Monolayer Topological Insulator

Memory is a foundational concept across disciplines, from neurobiology and electronics to artificial intelligence and quantum gravity. In materials, memory effects typically arise from ferroic orders, such as ferroelectricity and ferromagnetism, where information is stored in charge or spin degrees of freedom. Here, we report a surprising discovery of a nonvolatile superlattice memory effect in monolayer TaIrTe4, a dual quantum spin Hall insulator, where information is encoded through sharply contrasting lattice periodicities. In particular, in a pristine monolayer, we observe the spontaneous emergence of a long-period superlattice that can be programmed ON and OFF in a nonvolatile manner by electrostatic tuning of low-energy electronic states. This switching toggles the system between two structural configurations with unit cell areas differing by nearly two orders of magnitude. Mechanistically, our results reveal two independent and distinct instabilities, one in the lattice and the other in the QSH electrons, which are coupled, leading to electrostatic control of lattice configurations with nonvolatile memory. This finding is enabled by combining linear and nonlinear transport measurements, Raman spectroscopy, and scanning tunneling microscopy, which probe complementary aspects of the underlying orders. Remarkably, this nonvolatile memory effect stabilizes a spontaneous superlattice with a periodicity on the few-nanometer scale that remains robust across a wide doping range, persists over days, and survives above 70 K. Combined with the QSH topology, this stability offers a promising route to nonvolatile memory control of topological flat bands and their filling enabled quantum states. Our preliminary data indeed show the emergence of new insulating states at fractional superlattice fillings, which can be clearly switched ON and OFF together with the superlattice.

cond-mat.mes-hall↗

Comprehensive Study of Phonon Chirality under Symmetry Constraints

Phonons are quanta of lattice vibrations, and their modes (linear, circular, or stationary) are symmetry-determined. Circularly polarized phonons, possessing nonzero angular momentum (AM), have drawn widespread attention recently. Despite widespread use of pseudo-angular momentum (PAM) and circularly polarized light polarization flips to identify chiral phonons in Raman scattering, their reliability is debated due to symmetry dependence, and experimental verification standards remain lacking. Here, we systematically study phonon chirality and associated phenomena across magnetic point groups. We establish that the AM-PAM correlation is governed by both crystalline symmetry and Wyckoff positions, dictating conditions where nonzero AM manifests in PAM signatures. Crucially, phonons belonging to distinct irreducible representations exhibit distinct experimental benchmarks, enabling direct determination of crystalline chirality and symmetry classification. Furthermore, we report the discovery of a signature for symmetry-induced phenomena, notably a half-wave plate-analogous effect induced by mirror-odd phonons. Meanwhile, we conducted five experiments to validate our theory.

cond-mat.mtrl-sci↗

Dancing in Chains: Reconciling Instruction Following and Faithfulness in Language Models

Modern language models (LMs) need to follow human instructions while being faithful; yet, they often fail to achieve both. Here, we provide concrete evidence of a trade-off between instruction following (i.e., follow open-ended instructions) and faithfulness (i.e., ground responses in given context) when training LMs with these objectives. For instance, fine-tuning LLaMA-7B on instruction following datasets renders it less faithful. Conversely, instruction-tuned Vicuna-7B shows degraded performance at following instructions when further optimized on tasks that require contextual grounding. One common remedy is multi-task learning (MTL) with data mixing, yet it remains far from achieving a synergic outcome. We propose a simple yet effective method that relies on Rejection Sampling for Continued Self-instruction Tuning (ReSet), which significantly outperforms vanilla MTL. Surprisingly, we find that less is more, as training ReSet with high-quality, yet substantially smaller data (three-fold less) yields superior results. Our findings offer a better understanding of objective discrepancies in alignment training of LMs.

cs.CL↗

Sliding ferroelectric memories and synapses

Ferroelectric materials with switchable electric polarization hold great promise for a plethora of emergent applications, such as post-Moore's law nanoelectronics, beyond-Boltzmann transistors, non-volatile memories, and above-bandgap photovoltaic devices. Recent advances have uncovered an exotic sliding ferroelectric mechanism, which endows to design atomically thin ferroelectrics from non-ferroelectric parent monolayers. Although notable progress has been witnessed in understanding its fundamental properties, functional devices based on sliding ferroelectrics, the key touchstone toward applications, remain elusive. Here, we demonstrate the rewritable, non-volatile memory devices at room-temperature utilizing a two-dimensional (2D) sliding ferroelectric semiconductor of rhombohedral-stacked bilayer molybdenum disulfide. The 2D sliding ferroelectric memories (SFeMs) show superior performances with a large memory window of >8V, a high conductance ratio of above 106, a long retention time of >10 years, and a programming endurance greater than 104 cycles. Remarkably, flexible SFeMs are achieved with state-of-the-art performances competitive to their rigid counterparts and maintain their performances post bending over 103 cycles. Furthermore, synapse-specific Hebbian forms of plasticity and image recognition with a high accuracy of 97.81% are demonstrated based on flexible SFeMs. Our work demonstrates the sliding ferroelectric memories and synaptic plasticity on both rigid and flexible substrates, highlighting the great potential of sliding ferroelectrics for emerging technological applications in brain-inspired in-memory computing, edge intelligence and energy-efficient wearable electronics.

cond-mat.mes-hall↗

Room-temperature correlated states in twisted bilayer MoS$_2$

Moiré superlattices have emerged as an exciting condensed-matter quantum simulator for exploring the exotic physics of strong electronic correlations. Notable progress has been witnessed, but such correlated states are achievable usually at low temperatures. Here, we report the transport evidences of room-temperature correlated electronic states and layer-hybridized SU(4) Hubbard model simulator in AB-stacked MoS$_2$ homo-bilayer moiré superlattices. Correlated insulating states at moiré band filling factors v = 1, 2, 3 are unambiguously established in twisted bilayer MoS$_2$. Remarkably, the correlated electronic states can persist up to a record-high critical temperature of over 285 K. The realization of room-temperature correlated states in twisted bilayer MoS$_2$ can be understood as the cooperation effects of the stacking-specific atomic reconstruction and the resonantly enhanced interlayer hybridization, which largely amplify the moiré superlattice effects on electronic correlations. Furthermore, extreme large non-linear Hall responses up to room-temperature are uncovered near correlated insulating states, demonstrating the quantum geometry of moiré flat conduction band.

cond-mat.mtrl-sci↗

Sample size calculation based on the difference in restricted mean time lost for clinical trials with competing risks

Computation of sample size is important when designing clinical trials. The presence of competing risks makes the design of clinical trials with time-to-event endpoints cumbersome. A model based on the subdistribution hazard ratio (SHR) is commonly used for trials under competing risks. However, this approach has some limitations related to model assumptions and clinical interpretation. Considering such limitations, the difference in restricted mean time lost (RMTLd) is recommended as an alternative indicator. In this paper, we propose a sample size calculation method based on the RMTLd for the Weibull distribution (RMTLdWeibull) for clinical trials, which considers experimental conditions such as equal allocation, uniform accrual, uniform loss to follow-up, and administrative censoring. Simulation results show that sample size calculation based on the RMTLdWeibull can generally achieve a predefined power level and maintain relative robustness. Moreover, the performance of the sample size calculation based on the RMTLdWeibull is similar or superior to that based on the SHR. Even if the event time does not follow the Weibull distribution, the sample size calculation based on the RMTLdWeibull still performs well. The results also verify the performance of the sample size calculation method based on the RMTLdWeibull. From the perspective of the results of this study, clinical interpretation, application conditions and statistical performance, we recommend that when designing clinical trials in the presence of competing risks, the RMTLd indicator be applied for sample size calculation and subsequent effect size measurement.

stat.ME↗

Tokenization Consistency Matters for Generative Models on Extractive NLP Tasks

Generative models have been widely applied to solve extractive tasks, where parts of the input is extracted to form the desired output, and achieved significant success. For example, in extractive question answering (QA), generative models have constantly yielded state-of-the-art results. In this work, we identify the issue of tokenization inconsistency that is commonly neglected in training these models. This issue damages the extractive nature of these tasks after the input and output are tokenized inconsistently by the tokenizer, and thus leads to performance drop as well as hallucination. We propose a simple yet effective fix to this issue and conduct a case study on extractive QA. We show that, with consistent tokenization, the model performs better in both in-domain and out-of-domain datasets, with a notable average of +1.7 F2 gain when a BART model is trained on SQuAD and evaluated on 8 QA datasets. Further, the model converges faster, and becomes less likely to generate out-of-context answers. With these findings, we would like to call for more attention on how tokenization should be done when solving extractive tasks and recommend applying consistent tokenization during training.

cs.CL↗

Personalized Search Via Neural Contextual Semantic Relevance Ranking

Existing neural relevance models do not give enough consideration for query and item context information which diversifies the search results to adapt for personal preference. To bridge this gap, this paper presents a neural learning framework to personalize document ranking results by leveraging the signals to capture how the document fits into users' context. In particular, it models the relationships between document content and user query context using both lexical representations and semantic embeddings such that the user's intent can be better understood by data enrichment of personalized query context information. Extensive experiments performed on the search dataset, demonstrate the effectiveness of the proposed method.

cs.IR↗

Weyl phonons in chiral crystals

Chirality is an indispensable concept that pervades fundamental science and nature, manifesting itself in diverse forms such as chiral quasiparticles and chiral structures. Of particular interest are Weyl phonons carrying specific Chern numbers and chiral phonons doing circular motions in crystals. Up to now, Weyl and chiral phonons have been studied independently and the interpretations of chirality seem to be different in these two concepts, impeding our understanding. Here, we demonstrate that Weyl and chiral phonons are entangled in chiral crystals. Employing a typical chiral crystal of elementary tellurium (Te) as a case study, we expound on the intrinsic relationship between Chern number of Weyl phonons and pseudo-angular momentum (PAM) of chiral phonons. In light of the mutual coupling, we propose Raman scattering as a new technique to demonstrate the existence of Weyl phonons in Te, by detecting the chirality-induced energy splitting between the two constituent chiral phonon branches for Weyl phonons. By using the same experimental approach, we also observe the obstructed phonon surface states for the first time.

cond-mat.mtrl-sci↗

Language Agnostic Multilingual Information Retrieval with Contrastive Learning

Multilingual information retrieval (IR) is challenging since annotated training data is costly to obtain in many languages. We present an effective method to train multilingual IR systems when only English IR training data and some parallel corpora between English and other languages are available. We leverage parallel and non-parallel corpora to improve the pretrained multilingual language models' cross-lingual transfer ability. We design a semantic contrastive loss to align representations of parallel sentences that share the same semantics in different languages, and a new language contrastive loss to leverage parallel sentence pairs to remove language-specific information in sentence representations from non-parallel corpora. When trained on English IR data with these losses and evaluated zero-shot on non-English data, our model demonstrates significant improvement to prior work on retrieval performance, while it requires much less computational effort. We also demonstrate the value of our model for a practical setting when a parallel corpus is only available for a few languages, but a lack of parallel corpora resources persists for many other low-resource languages. Our model can work well even with a small number of parallel sentences, and be used as an add-on module to any backbones and other tasks.

cs.IR↗

Electron-infrared phonon coupling in ABC trilayer graphene

Stacking order plays a crucial role in determining the crystal symmetry and has significant impacts on electronic, optical, magnetic, and topological properties. Electron-phonon coupling, which is central to a wide range of intriguing quantum phenomena, is expected to be intricately connected with stacking order. Understanding the stacking order-dependent electron-phonon coupling is essential for understanding peculiar physical phenomena associated with electron-phonon coupling, such as superconductivity and charge density waves. In this study, we investigate the effect of stacking order on electron-infrared phonon coupling in graphene trilayers. By using gate-tunable Raman spectroscopy and excitation frequency-dependent near-field infrared nanoscopy, we show that rhombohedral ABC-stacked trilayer graphene has a significantly stronger electron-infrared phonon coupling strength than the Bernal ABA-stacked trilayer graphene. Our findings provide novel insights into the superconductivity and other fundamental physical properties of rhombohedral ABC-stacked trilayer graphene, and can enable nondestructive and high-throughput imaging of trilayer graphene stacking order using Raman scattering.

physics.optics↗

STREET: A Multi-Task Structured Reasoning and Explanation Benchmark

We introduce STREET, a unified multi-task and multi-domain natural language reasoning and explanation benchmark. Unlike most existing question-answering (QA) datasets, we expect models to not only answer questions, but also produce step-by-step structured explanations describing how premises in the question are used to produce intermediate conclusions that can prove the correctness of a certain answer. We perform extensive evaluation with popular language models such as few-shot prompting GPT-3 and fine-tuned T5. We find that these models still lag behind human performance when producing such structured reasoning steps. We believe this work will provide a way for the community to better train and test systems on multi-step reasoning and explanations in natural language.

cs.CL↗

Improving Cross-task Generalization of Unified Table-to-text Models with Compositional Task Configurations

There has been great progress in unifying various table-to-text tasks using a single encoder-decoder model trained via multi-task learning (Xie et al., 2022). However, existing methods typically encode task information with a simple dataset name as a prefix to the encoder. This not only limits the effectiveness of multi-task learning, but also hinders the model's ability to generalize to new domains or tasks that were not seen during training, which is crucial for real-world applications. In this paper, we propose compositional task configurations, a set of prompts prepended to the encoder to improve cross-task generalization of unified models. We design the task configurations to explicitly specify the task type, as well as its input and output types. We show that this not only allows the model to better learn shared knowledge across different tasks at training, but also allows us to control the model by composing new configurations that apply novel input-output combinations in a zero-shot manner. We demonstrate via experiments over ten table-to-text tasks that our method outperforms the UnifiedSKG baseline by noticeable margins in both in-domain and zero-shot settings, with average improvements of +0.5 and +12.6 from using a T5-large backbone, respectively.

cs.CL↗

Entailment Tree Explanations via Iterative Retrieval-Generation Reasoner

Large language models have achieved high performance on various question answering (QA) benchmarks, but the explainability of their output remains elusive. Structured explanations, called entailment trees, were recently suggested as a way to explain and inspect a QA system's answer. In order to better generate such entailment trees, we propose an architecture called Iterative Retrieval-Generation Reasoner (IRGR). Our model is able to explain a given hypothesis by systematically generating a step-by-step explanation from textual premises. The IRGR model iteratively searches for suitable premises, constructing a single entailment step at a time. Contrary to previous approaches, our method combines generation steps and retrieval of premises, allowing the model to leverage intermediate conclusions, and mitigating the input size limit of baseline encoder-decoder models. We conduct experiments using the EntailmentBank dataset, where we outperform existing benchmarks on premise retrieval and entailment tree generation, with around 300% gain in overall correctness.

cs.CL↗

Contrastive Document Representation Learning with Graph Attention Networks

Recent progress in pretrained Transformer-based language models has shown great success in learning contextual representation of text. However, due to the quadratic self-attention complexity, most of the pretrained Transformers models can only handle relatively short text. It is still a challenge when it comes to modeling very long documents. In this work, we propose to use a graph attention network on top of the available pretrained Transformers model to learn document embeddings. This graph attention network allows us to leverage the high-level semantic structure of the document. In addition, based on our graph document model, we design a simple contrastive learning strategy to pretrain our models on a large amount of unlabeled corpus. Empirically, we demonstrate the effectiveness of our approaches in document classification and document retrieval tasks.

cs.CL↗

Attention-guided Generative Models for Extractive Question Answering

We propose a novel method for applying Transformer models to extractive question answering (QA) tasks. Recently, pretrained generative sequence-to-sequence (seq2seq) models have achieved great success in question answering. Contributing to the success of these models are internal attention mechanisms such as cross-attention. We propose a simple strategy to obtain an extractive answer span from the generative model by leveraging the decoder cross-attention patterns. Viewing cross-attention as an architectural prior, we apply joint training to further improve QA performance. Empirical results show that on open-domain question answering datasets like NaturalQuestions and TriviaQA, our method approaches state-of-the-art performance on both generative and extractive inference, all while using much fewer parameters. Furthermore, this strategy allows us to perform hallucination-free inference while conferring significant improvements to the model's ability to rerank relevant passages.

cs.CL↗

Multiplicative Position-aware Transformer Models for Language Understanding

Transformer models, which leverage architectural improvements like self-attention, perform remarkably well on Natural Language Processing (NLP) tasks. The self-attention mechanism is position agnostic. In order to capture positional ordering information, various flavors of absolute and relative position embeddings have been proposed. However, there is no systematic analysis on their contributions and a comprehensive comparison of these methods is missing in the literature. In this paper, we review major existing position embedding methods and compare their accuracy on downstream NLP tasks, using our own implementations. We also propose a novel multiplicative embedding method which leads to superior accuracy when compared to existing methods. Finally, we show that our proposed embedding method, served as a drop-in replacement of the default absolute position embedding, can improve the RoBERTa-base and RoBERTa-large models on SQuAD1.1 and SQuAD2.0 datasets.

cs.CL↗

Spatially indirect intervalley excitons in bilayer WSe2

Spatially indirect excitons with displaced wavefunctions of electrons and holes play a pivotal role in a large portfolio of fascinating physical phenomena and emerging optoelectronic applications, such as valleytronics, exciton spin Hall effect, excitonic integrated circuit and high-temperature superfluidity. Here, we uncover three types of spatially indirect excitons (including their phonon replicas) and their quantum-confined Stark effects in hexagonal boron nitride encapsulated bilayer WSe2, by performing electric field-tunable photoluminescence measurements. Because of different out-of-plane electric dipole moments, the energy order between the three types of spatially indirect excitons can be switched by a vertical electric field. Remarkably, we demonstrate, assisted by first-principles calculations, that the observed spatially indirect excitons in bilayer WSe2 are also momentum-indirect, involving electrons and holes from Q and K/Γ valleys in the Brillouin zone, respectively. This is in contrast to the previously reported spatially indirect excitons with electrons and holes localized in the same valley. Furthermore, we find that the spatially indirect intervalley excitons in bilayer WSe2 can exhibit considerable, doping-sensitive circular polarization. The spatially indirect excitons with momentum-dark nature and highly tunable circular polarization open new avenues for exotic valley physics and technological innovations in photonics and optoelectronics.

cond-mat.mes-hall↗