Searcharxiv⌕ Search

arXiv subjects

Mingda Li

Publications and source records attributed to Mingda Li.

At least 55 records · Page 3Linked to original sources

Unraveling and Mitigating Retriever Inconsistencies in Retrieval-Augmented Large Language Models

Although Retrieval-Augmented Large Language Models (RALMs) demonstrate their superiority in terms of factuality, they do not consistently outperform the original retrieval-free Language Models (LMs). Our experiments reveal that this example-level performance inconsistency exists not only between retrieval-augmented and retrieval-free LM but also among different retrievers. To understand this phenomenon, we investigate the degeneration behavior of RALMs and theoretically decompose it into four categories. Further analysis based on our decomposition reveals that the innate difference in knowledge sources and the unpredictable degeneration of the reader model contribute most to the inconsistency. Drawing from our analysis, we introduce Ensemble of Retrievers (EoR), a trainable framework that can adaptively retrieve from different knowledge sources and effectively decrease unpredictable reader errors. Our experiments on Open Domain Question Answering show that EoR substantially improves performance over the RALM with a single retriever by considerably reducing inconsistent behaviors.

cs.AI↗

TRACE: Temporal Grounding Video LLM via Causal Event Modeling

Video Temporal Grounding (VTG) is a crucial capability for video understanding models and plays a vital role in downstream tasks such as video browsing and editing. To effectively handle various tasks simultaneously and enable zero-shot prediction, there is a growing trend in employing video LLMs for VTG tasks. However, current video LLM-based methods rely exclusively on natural language generation, lacking the ability to model the clear structure inherent in videos, which restricts their effectiveness in tackling VTG tasks. To address this issue, this paper first formally introduces causal event modeling framework, which represents video LLM outputs as sequences of events, and predict the current event using previous events, video inputs, and textural instructions. Each event consists of three components: timestamps, salient scores, and textual captions. We then propose a novel task-interleaved video LLM called TRACE to effectively implement the causal event modeling framework in practice. The TRACE process visual frames, timestamps, salient scores, and text as distinct tasks, employing various encoders and decoding heads for each. Task tokens are arranged in an interleaved sequence according to the causal event modeling framework's formulation. Extensive experiments on various VTG tasks and datasets demonstrate the superior performance of TRACE compared to state-of-the-art video LLMs. Our model and code are available at https://github.com/gyxxyg/TRACE.

cs.CV↗

Real-time interpretation of neutron vibrational spectra with symmetry-equivariant Hessian matrix prediction

The vibrational behavior of molecules serves as a crucial fingerprint of their structure, chemical state, and surrounding environment. Neutron vibrational spectroscopy provides comprehensive measurements of vibrational modes without selection rule restrictions. However, analyzing and interpreting the resulting spectra remains a computationally formidable task. Here, we introduce a symmetry-aware neural network that directly predicts Hessian matrices from molecular structures, thereby enabling rapid vibrational spectral reconstruction. Unlike traditional approaches that focus on eigenvalue prediction, the Hessian matrix provides richer, more fundamental information with broader applications and superior extrapolation. This approach also paves the way for predicting other properties, such as reaction pathways. Trained on small molecules, our model achieves spectroscopic-level accuracy, allowing real-time, unambiguous peak assignment. Moreover, it maintains high accuracy for larger molecules, demonstrating strong transferability. This adaptability unlocks new capabilities, including on-the-fly spectral interpretation for future autonomous laboratories, and offers insights into molecular design for targeted chemical pathways.

physics.chem-ph↗

AI-driven materials design: a mini-review

Materials design is an important component of modern science and technology, yet traditional approaches rely heavily on trial-and-error and can be inefficient. Computational techniques, enhanced by modern artificial intelligence (AI), have greatly accelerated the design of new materials. Among these approaches, inverse design has shown great promise in designing materials that meet specific property requirements. In this mini-review, we summarize key computational advancements for materials design over the past few decades. We follow the evolution of relevant materials design techniques, from high-throughput forward machine learning (ML) methods and evolutionary algorithms, to advanced AI strategies like reinforcement learning (RL) and deep generative models. We highlight the paradigm shift from conventional screening approaches to inverse generation driven by deep generative models. Finally, we discuss current challenges and future perspectives of materials inverse design. This review may serve as a brief guide to the approaches, progress, and outlook of designing future functional materials with technological relevance.

cond-mat.mtrl-sci↗

VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding

Video Temporal Grounding (VTG) strives to accurately pinpoint event timestamps in a specific video using linguistic queries, significantly impacting downstream tasks like video browsing and editing. Unlike traditional task-specific models, Video Large Language Models (video LLMs) can handle multiple tasks concurrently in a zero-shot manner. Consequently, exploring the application of video LLMs for VTG tasks has become a burgeoning research area. However, despite considerable advancements in video content understanding, video LLMs often struggle to accurately pinpoint timestamps within videos, limiting their effectiveness in VTG tasks. To address this, we introduce VTG-LLM, a model designed to enhance video LLMs' timestamp localization abilities. Our approach includes: (1) effectively integrating timestamp knowledge into visual tokens; (2) incorporating absolute-time tokens to manage timestamp knowledge without concept shifts; and (3) introducing a lightweight, high-performance, slot-based token compression technique designed to accommodate the demands of a large number of frames to be sampled for VTG tasks. Additionally, we present VTG-IT-120K, a collection of publicly available VTG datasets that we have re-annotated to improve upon low-quality annotations. Our comprehensive experiments demonstrate the superior performance of VTG-LLM in comparison to other video LLM methods across a variety of VTG tasks.

cs.CV↗

Theory of the Photomolecular Effect

It is well-known that water in both liquid and vapor phases exhibits exceptionally weak absorption of light in the visible range. Recent experiments, however, have demonstrated that at the liquid-air interface, absorption in the visible range is drastically increased. This increased absorption results in a rate of evaporation that exceeds the theoretical thermal limit by between two and five times. Curiously, the evaporation rate peaks at green wavelengths of light, while no corresponding absorptance peak has been observed. Experiments suggest that photons can cleave off clusters of water molecules at the surface, but no clear theoretical model has yet been proposed to explain how this is possible. This paper aims to present such a model and explain this surprising and important phenomenon.

cond-mat.soft↗

Quantum Theory of X-ray Photon Correlation Spectroscopy

Characterizing quantum materials is essential for understanding their microscopic interactions and advancing quantum technology. X-ray photon correlation spectroscopy (XPCS) with coherent X-ray sources offers access to higher-order correlations, but its theoretical basis, the Siegert relation, is derived from dynamical light scattering with independent classical scatterers, and its validity for XPCS remains unexamined. Here we present a microscopic quantum theory of XPCS derived from elecron-photon interaction Hamiltonians, introducing four configurations tied to distinct fourth-order electron-density correlation functions. We examine the validity of the Siegert relation and derive a generalized Siegert relation. Notably, the Siegert relation breaks down even in non-interacting Fermi gas due to exchange interactions. Furthermore, density matrix renormalization group calculations on 1D Kitaev chain reveal oscillatary signatures that can distinguish topologically trivial phases from topological phases with Majorana zero modes. Our work provides a robust theoretical foundation for XPCS and highlights the value of higher-order correlations in advanced X-ray and neutron sources for probing quantum materials.

cond-mat.mtrl-sci↗

Unstructured Text Enhanced Open-domain Dialogue System: A Systematic Survey

Incorporating external knowledge into dialogue generation has been proven to benefit the performance of an open-domain Dialogue System (DS), such as generating informative or stylized responses, controlling conversation topics. In this article, we study the open-domain DS that uses unstructured text as external knowledge sources (\textbf{U}nstructured \textbf{T}ext \textbf{E}nhanced \textbf{D}ialogue \textbf{S}ystem, \textbf{UTEDS}). The existence of unstructured text entails distinctions between UTEDS and traditional data-driven DS and we aim to analyze these differences. We first give the definition of the UTEDS related concepts, then summarize the recently released datasets and models. We categorize UTEDS into Retrieval and Generative models and introduce them from the perspective of model components. The retrieval models consist of Fusion, Matching, and Ranking modules, while the generative models comprise Dialogue and Knowledge Encoding, Knowledge Selection, and Response Generation modules. We further summarize the evaluation methods utilized in UTEDS and analyze the current models' performance. At last, we discuss the future development trends of UTEDS, hoping to inspire new research in this field.

cs.CL↗

Large Language Model-Guided Prediction Toward Quantum Materials Synthesis

The synthesis of inorganic crystalline materials is essential for modern technology, especially in quantum materials development. However, designing efficient synthesis workflows remains a significant challenge due to the precise experimental conditions and extensive trial and error. Here, we present a framework using large language models (LLMs) to predict synthesis pathways for inorganic materials, including quantum materials. Our framework contains three models: LHS2RHS, predicting products from reactants; RHS2LHS, predicting reactants from products; and TGT2CEQ, generating full chemical equations for target compounds. Fine-tuned on a text-mined synthesis database, our model raises accuracy from under 40% with pretrained models, to under 80% using conventional fine-tuning, and further to around 90% with our proposed generalized Tanimoto similarity, while maintaining robust to additional synthesis steps. Our model further demonstrates comparable performance across materials with varying degrees of quantumness quantified using quantum weight, indicating that LLMs offer a powerful tool to predict balanced chemical equations for quantum materials discovery.

cond-mat.mtrl-sci↗

Policy-driven Knowledge Selection and Response Generation for Document-grounded Dialogue

Document-grounded dialogue (DGD) uses documents as external knowledge for dialogue generation. Correctly understanding the dialogue context is crucial for selecting knowledge from the document and generating proper responses. In this paper, we propose using a dialogue policy to help the dialogue understanding in DGD. Our dialogue policy consists of two kinds of guiding signals: utterance function and topic transfer intent. The utterance function reflects the purpose and style of an utterance, and the topic transfer intent reflects the topic and content of an utterance. We propose a novel framework exploiting our dialogue policy for two core tasks in DGD, namely knowledge selection (KS) and response generation (RG). The framework consists of two modules: the Policy planner leverages policy-aware dialogue representation to select knowledge and predict the policy of the response; the generator uses policy/knowledge-aware dialogue representation for response generation. Our policy-driven model gets state-of-the-art performance on three public benchmarks and we provide a detailed analysis of the experimental results. Our code/data will be released on GitHub.

cs.CL↗

CRUcialG: Reconstruct Integrated Attack Scenario Graphs by Cyber Threat Intelligence Reports

Cyber Threat Intelligence (CTI) reports are factual records compiled by security analysts through their observations of threat events or their own practical experience with attacks. In order to utilize CTI reports for attack detection, existing methods have attempted to map the content of reports onto system-level attack provenance graphs to clearly depict attack procedures. However, existing studies on constructing graphs from CTI reports suffer from problems such as weak natural language processing (NLP) capabilities, discrete and fragmented graphs, and insufficient attack semantic representation. Therefore, we propose a system called CRUcialG for the automated reconstruction of attack scenario graphs (ASGs) by CTI reports. First, we use NLP models to extract systematic attack knowledge from CTI reports to form preliminary ASGs. Then, we propose a four-phase attack rationality verification framework from the tactical phase with attack procedure to evaluate the reasonability of ASGs. Finally, we implement the relation repair and phase supplement of ASGs by adopting a serialized graph generation model. We collect a total of 10,607 CTI reports and generate 5,761 complete ASGs. Experimental results on CTI reports from 30 security vendors and DARPA show that the similarity of ASG reconstruction by CRUcialG can reach 84.54%. Compared with SOTA (EXTRACTOR and AttackG), the recall of CRUcialG (extraction of real attack events) can reach 88.13% and 94.46% respectively, which is 40% higher than SOTA on average. The F1-score of attack phase verification is able to reach 90.04%.

cs.CR↗

PclGPT: A Large Language Model for Patronizing and Condescending Language Detection

Disclaimer: Samples in this paper may be harmful and cause discomfort! Patronizing and condescending language (PCL) is a form of speech directed at vulnerable groups. As an essential branch of toxic language, this type of language exacerbates conflicts and confrontations among Internet communities and detrimentally impacts disadvantaged groups. Traditional pre-trained language models (PLMs) perform poorly in detecting PCL due to its implicit toxicity traits like hypocrisy and false sympathy. With the rise of large language models (LLMs), we can harness their rich emotional semantics to establish a paradigm for exploring implicit toxicity. In this paper, we introduce PclGPT, a comprehensive LLM benchmark designed specifically for PCL. We collect, annotate, and integrate the Pcl-PT/SFT dataset, and then develop a bilingual PclGPT-EN/CN model group through a comprehensive pre-training and supervised fine-tuning staircase process to facilitate implicit toxic detection. Group detection results and fine-grained detection from PclGPT and other models reveal significant variations in the degree of bias in PCL towards different vulnerable groups, necessitating increased societal attention to protect them.

cs.CL↗

Thickness-Dependent Polaron Crossover in Tellurene

Polarons, quasiparticles arising from electron-phonon coupling, are crucial in understanding material properties such as high-temperature superconductivity and colossal magnetoresistance. However, scarce studies have been performed to investigate the formation of polarons in low-dimensional materials with phonon polarity and electronic structure transitions. In this work, we studied polarons of tellurene that are composed of chiral chains of tellurium atoms. The frequency and linewidth of the A1 phonon, which becomes increasingly polar for thinner tellurene, exhibit an abrupt change when the thickness of tellurene is below 10 nm. Meanwhile, the field effect mobility of tellurene drops rapidly as the thickness is smaller than 10 nm. These phonon and transport signatures, combined with the calculated phonon polarity and band structure, suggest a crossover from large polarons for bulk tellurium to small polarons for few-layer tellurene. Effective field theory considers the phonon renormalization in the strong coupling (small polaron) regime, and semi-quantitatively reproduces the observed phonon hardening and broadening effects in few-layer tellurene. This polaron crossover stems from the quasi-1D nature of tellurene where modulation of the interchain distance reduces the dielectric screening and promotes electron-phonon coupling. Our work provides valuable insights into the influence of polarons on phononic, electronic, and structural properties in low-dimensional materials.

cond-mat.mtrl-sci↗

Enhancing Long Video Understanding via Hierarchical Event-Based Memory

Recently, integrating visual foundation models into large language models (LLMs) to form video understanding systems has attracted widespread attention. Most of the existing models compress diverse semantic information within the whole video and feed it into LLMs for content comprehension. While this method excels in short video understanding, it may result in a blend of multiple event information in long videos due to coarse compression, which causes information redundancy. Consequently, the semantics of key events might be obscured within the vast information that hinders the model's understanding capabilities. To address this issue, we propose a Hierarchical Event-based Memory-enhanced LLM (HEM-LLM) for better understanding of long videos. Firstly, we design a novel adaptive sequence segmentation scheme to divide multiple events within long videos. In this way, we can perform individual memory modeling for each event to establish intra-event contextual connections, thereby reducing information redundancy. Secondly, while modeling current event, we compress and inject the information of the previous event to enhance the long-term inter-event dependencies in videos. Finally, we perform extensive experiments on various video understanding tasks and the results show that our model achieves state-of-the-art performances.

cs.CV↗

Magneto-optical trapping of a heavy polyatomic molecule for precision measurement

We report a magneto-optical trap of strontium monohydroxide (SrOH) containing 2000(600) molecules at a temperature of 1.2(3) mK. The lifetime is 91(9) ms, which is limited by decay to optically unaddressed vibrational states. This provides the foundation for future sub-Doppler cooling and optical trapping of SrOH, a polyatomic molecule suited for precision searches for physics beyond the Standard Model including new CP violating particles and ultralight dark matter. We also identify important features in this system that guide cooling and trapping of complex and heavy polyatomic molecules into the ultracold regime.

physics.atom-ph↗

TC-LLaVA: Rethinking the Transfer from Image to Video Understanding with Temporal Considerations

Multimodal Large Language Models (MLLMs) have significantly improved performance across various image-language applications. Recently, there has been a growing interest in adapting image pre-trained MLLMs for video-related tasks. However, most efforts concentrate on enhancing the vision encoder and projector components, while the core part, Large Language Models (LLMs), remains comparatively under-explored. In this paper, we propose two strategies to enhance the model's capability in video understanding tasks by improving inter-layer attention computation in LLMs. Specifically, the first approach focuses on the enhancement of Rotary Position Embedding (RoPE) with Temporal-Aware Dual RoPE, which introduces temporal position information to strengthen the MLLM's temporal modeling capabilities while preserving the relative position relationships of both visual and text tokens. The second approach involves enhancing the Attention Mask with the Frame-wise Block Causal Attention Mask, a simple yet effective method that broadens visual token interactions within and across video frames while maintaining the causal inference mechanism. Based on these proposed methods, we adapt LLaVA for video understanding tasks, naming it Temporal-Considered LLaVA (TC-LLaVA). Our TC-LLaVA achieves new state-of-the-art performance across various video understanding benchmarks with only supervised fine-tuning (SFT) on video-related datasets.

cs.CV↗

Giant Uniaxial Magnetocrystalline Anisotropy in SmCrGe$_3$

Magnetic anisotropy is a crucial characteristic for enhancing spintronic device performance. The synthesis of SmCrGe$_3$ single crystals through a high-temperature solution method has led to the determination of uniaxial magnetocrystalline anisotropy. Phase verification was achieved using scanning transmission electron microscopy (STEM), powder, and single-crystal X-ray diffraction techniques. Electrical transport and specific heat measurements indicate a Curie temperature ($T_C$) of approximately 160 K, while magnetization measurements were utilized to determine the anisotropy fields and constants. Curie-Weiss fitting applied to magnetization data suggests the contribution of both Sm and Cr in the paramagnetic phase. Additionally, density functional theory (DFT) calculations explored the electronic structures and magnetic properties of SmCrGe$_3$, revealing a significant easy-axis single-ion Sm magnetocrystalline anisotropy of 16 meV/f.u.. Based on the magnetization measurements, easy-axis magnetocrystalline anisotropy at 20 K is 13 meV/f.u..

cond-mat.mtrl-sci↗

Structural Constraint Integration in Generative Model for Discovery of Quantum Material Candidates

Billions of organic molecules are known, but only a tiny fraction of the functional inorganic materials have been discovered, a particularly relevant problem to the community searching for new quantum materials. Recent advancements in machine-learning-based generative models, particularly diffusion models, show great promise for generating new, stable materials. However, integrating geometric patterns into materials generation remains a challenge. Here, we introduce Structural Constraint Integration in the GENerative model (SCIGEN). Our approach can modify any trained generative diffusion model by strategic masking of the denoised structure with a diffused constrained structure prior to each diffusion step to steer the generation toward constrained outputs. Furthermore, we mathematically prove that SCIGEN effectively performs conditional sampling from the original distribution, which is crucial for generating stable constrained materials. We generate eight million compounds using Archimedean lattices as prototype constraints, with over 10% surviving a multi-staged stability pre-screening. High-throughput density functional theory (DFT) on 26,000 survived compounds shows that over 50% passed structural optimization at the DFT level. Since the properties of quantum materials are closely related to geometric patterns, our results indicate that SCIGEN provides a general framework for generating quantum materials candidates.

cond-mat.mtrl-sci↗