SearcharxivSearch

arXiv subjects

Hongfei Li

Publications and source records attributed to Hongfei Li.

10 recordsLinked to original sources

LentEx: Generalizable Latent Entity Extraction via Synthetic Data and Instruction-Tuned LLMs

Latent entity extraction (LEE) tackles the challenge of identifying implicit, contextually inferred entities within free text-an area where traditional entity extraction methods fall short. In this paper, we introduce LentEx, a novel framework for latent entity extraction that leverages synthetic data generation and instruction fine-tuning to optimize smaller, efficient large language models (LLMs). Latent entities, which are often abstract and thematic, are crucial for applications such as retrieval-augmented generation (RAG), customer persona analysis, and knowledge graph enrichment. LentEx addresses the scarcity of labeled datasets by employing a template-based approach to generate diverse, contextually rich synthetic data, ensuring high variability and alignment with real-world distributions. To our knowledge, LentEx is the first to systematically approach LEE through the lens of LLMs. LentEx demonstrates significant performance improvements across multiple tasks, notably surpassing state-of-the-art models on the MTEB Clustering Benchmark. Furthermore, our methodology enables robust generalization to unseen domains, making LentEx highly applicable in real-world NLP tasks, including RAG and clustering, thereby establishing a new paradigm for latent entity understanding and extraction in natural language processing.

cs.CL

A Calibrated Reflection Approach for Enhancing Confidence Estimation in LLMs

A critical challenge in deploying Large Language Models (LLMs) is developing reliable mechanisms to estimate their confidence, enabling systems to determine when to trust model outputs versus seek human intervention. We present a Calibrated Reflection approach for enhancing confidence estimation in LLMs, a framework that combines structured reasoning with distance-aware calibration technique. Our approach introduces three key innovations: (1) a Maximum Confidence Selection (MCS) method that comprehensively evaluates confidence across all possible labels, (2) a reflection-based prompting mechanism that enhances reasoning reliability, and (3) a distance-aware calibration technique that accounts for ordinal relationships between labels. We evaluate our framework on diverse datasets, including HelpSteer2, Llama T-REx, and a proprietary conversational dataset, demonstrating its effectiveness across both conversational and fact-based classification tasks. This work contributes to the broader goal of developing reliable and well-calibrated confidence estimation methods for LLMs, enabling informed decisions about model trust and human judgement.

cs.CL

Real-World Cooperative Bimanual Dexterous Grasp of Large Objects from Single-View Observations

Bimanual dexterous grasping of large objects is a critical challenge in robotic manipulation. However, most existing studies focus on sequential manipulation rather than cooperative grasping, and methods addressing such bimanual tasks have largely been limited to simulation. These limitations stem from the difficulty of acquiring full 3D object models and generating physically plausible grasping actions. To fill this gap, we propose a real-world bimanual grasping framework that includes: a multimodal dataset capturing joint angles, visual observations and force signals; a Denoising Diffusion Probabilistic Model (DDPM)-based module that generates joint-level grasp configurations from segmented point clouds; and an execution strategy that integrates motion planning with online grasp refinement to ensure physical stability and feasibility. Our approach enables the synthesis of executable bimanual grasps from single-view inputs, reducing dependence on complete 3D object models and ensuring stable real-world performance. Experiments on a dual-arm robot demonstrate high success rates across unseen objects with varying geometries and poses, and ablation studies confirm the contributions of key components of our system.

cs.RO

Anisotropic Tensile Strength and Fracture Mechanism of $\theta$-TaN: A Machine-Learning Potential Molecular Dynamics Study

theta-phase tantalum nitride (theta-TaN) combines metallic conductivity with exceptionally high thermal conductivity, making it a potential material for device thermal management and interconnect applications. However, its tensile strength and fracture behavior remain unclear. Here, we investigate the anisotropic tensile response and fracture mechanism of theta-TaN using neuroevolution-potential molecular dynamics simulations. Size-convergence tests show that a 20 nm long model is sufficient for reliable prediction, and the mechanical parameters vary by less than 3.5% over the strain-rate range of 10^7 to 10^9 s^-1. The results reveal strong tensile anisotropy. The c-axis direction ([0001]) shows a higher strength of 80.10 GPa and modulus of 748.63 GPa, but a lower fracture strain of 15.02%. In contrast, the a-axis direction ([2-1-10]) shows a lower strength of 56.87 GPa and modulus of 570.74 GPa, but a higher fracture strain of 17.71%. From 300 to 900 K, the mechanical properties decrease nearly linearly, while more than 73% of the 300 K strength is retained at 900 K. Fracture occurs without observable dislocation activity and is governed by cleavage-plane selection: {10-10} prismatic planes under a-axis tension and the (0001) basal plane under c-axis tension. Atomic displacement analysis shows that local separation and microvoid formation precede macroscopic crack growth, indicating a brittle fracture process driven by local bond-network instability. These results provide atomic-scale mechanical data for assessing the reliability of theta-TaN in thermal management applications.

cond-mat.mtrl-sci

Inference Time Feature Injection: A Lightweight Approach for Real-Time Recommendation Freshness

Many recommender systems in long-form video streaming reply on batch-trained models and batch-updated features, where user features are updated daily and served statically throughout the day. While efficient, this approach fails to incorporate a user's most recent actions, often resulting in stale recommendations. In this work, we present a lightweight, model-agnostic approach for intra-day personalization that selectively injects recent watch history at inference time without requiring model retraining. Our approach selectively overrides stale user features at inference time using the recent watch history, allowing the system to adapt instantly to evolving preferences. By reducing the personalization feedback loop from daily to intra-day, we observed a statistically significant 0.47% increase in key user engagement metrics which ranked among the most substantial engagement gains observed in recent experimentation cycles. To our knowledge, this is the first published evidence that intra-day personalization can drive meaningful impact in long-form video streaming service, providing a compelling alternative to full real-time architectures where model retraining is required.

cs.LG

Bias in Large Language Models: Origin, Evaluation, and Mitigation

Large Language Models (LLMs) have revolutionized natural language processing, but their susceptibility to biases poses significant challenges. This comprehensive review examines the landscape of bias in LLMs, from its origins to current mitigation strategies. We categorize biases as intrinsic and extrinsic, analyzing their manifestations in various NLP tasks. The review critically assesses a range of bias evaluation methods, including data-level, model-level, and output-level approaches, providing researchers with a robust toolkit for bias detection. We further explore mitigation strategies, categorizing them into pre-model, intra-model, and post-model techniques, highlighting their effectiveness and limitations. Ethical and legal implications of biased LLMs are discussed, emphasizing potential harms in real-world applications such as healthcare and criminal justice. By synthesizing current knowledge on bias in LLMs, this review contributes to the ongoing effort to develop fair and responsible AI systems. Our work serves as a comprehensive resource for researchers and practitioners working towards understanding, evaluating, and mitigating bias in LLMs, fostering the development of more equitable AI technologies.

cs.CL

High-performance magnesium/sodium hybrid ion battery based on sodium vanadate oxide for reversible storage of Na+ and Mg2+

Magnesium ion batteries (MIBs) are a potential field for the energy storage of the future but are restricted by insufficient rate capability and rapid capacity degradation. Magnesium-sodium hybrid ion batteries (MSHBs) are an effective way to address these problems. Here, we report a new type of MSHBs that use layered sodium vanadate ((Na, Mn)V8O20 5H2O, Mn-NVO) cathodes coupled with an organic 3,4,9,10-perylenetetracarboxylic diimide (PTCDI) anode in Mg2+/Na+ hybrid electrolytes. During electrochemical cycling, Mg2+ and Na+ co-participate in the cathode reactions, and the introduction of Na+ promotes the structural stability of the Mn-NVO cathode, as cleared by several ex-situ characterizations. Consequently, the Mn-NVO cathode presents great specific capacity (249.9 mAh g-1 at 300 mA g-1) and cycling (1500 cycles at 1500 mA g-1) in the Mg2+/Na+ hybrid electrolytes. Besides, full battery displays long lifespan with 10,000 cycles at 1000 mA g-1. The rate performance and cycling stability of MSHBs have been improved by an economical and scalable method, and the mechanism for these improvements was discussed.

cond-mat.mtrl-sci

Au12@Au30: Core-Shell Molecule Constituted of an Icosahedron and an Icosidodecahedron

A stable core-shell structure with Ih symmetry, Au12@Au30, has been investigated by first-principles calculations. It is composed of an icosahedron core and an icosidodecahedron shell. The stability of the core-shell Au42 structure is verified by vibrational frequency analysis and molecular dynamics NVT simulations. Both the frontier molecular orbitals and the spin density of states show obvious s-d hybridization characteristics. The adaptive natural density partitioning analysis demonstrate multi-center bonds, twenty 6-center {\sigma} bonds ,and one 12-center {\sigma} bond, which are of great importance for the core-shell structural stability. In this core-shell nanostructure, there are also a large number of one-center valence lone electron pairs with the characteristics of d-like orbitals, so that the proposed Au12@Au30 could be used in medicine and catalysis fields.

cond-mat.mtrl-sci

Electrospun Conjugated Polymer/Fullerene Hybrid Fibers: Photoactive Blends, Conductivity through Tunnelling-AFM, Light-Scattering, and Perspective for Their Use in Bulk-Heterojunction Organic Solar Cells

Hybrid conjugated polymer/fullerene filaments based on MEH-PPV/PVP/PCBM are prepared by electrospinning, and their properties assessed by scanning electron, atomic and lateral force, tunnelling, and confocal microscopy, as well as by attenuated total reflection Fourier transform-infrared spectroscopy, photoluminescence quantum yield and spatially-resolved fluorescence. Highlighted features include ribbon-shape of the realized fibers, and the persistence of a network serving as a template for heterogeneous active layers in solar cell devices. A set of favorable characteristics is evidenced in this way in terms of homogeneous charge transport behavior and formation of effective interfaces for diffusion and dissociation of photogenerated excitons. The interaction of the organic filaments with light, exhibiting specific light-scattering properties of the nanofibrous mat, might also contribute to spreading incident radiation across the active layers, thus potentially enhancing photovoltaic performance. This method might be applied to other electron donor-electron acceptor material systems for the fabrication of solar cell devices enhanced by nanofibrillar morphologies embedding conjugated polymers and fullerene compounds.

physics.app-ph

Sparse-GEV: Sparse Latent Space Model for Multivariate Extreme Value Time Serie Modeling

In many applications of time series models, such as climate analysis and social media analysis, we are often interested in extreme events, such as heatwave, wind gust, and burst of topics. These time series data usually exhibit a heavy-tailed distribution rather than a Gaussian distribution. This poses great challenges to existing approaches due to the significantly different assumptions on the data distributions and the lack of sufficient past data on extreme events. In this paper, we propose the Sparse-GEV model, a latent state model based on the theory of extreme value modeling to automatically learn sparse temporal dependence and make predictions. Our model is theoretically significant because it is among the first models to learn sparse temporal dependencies among multivariate extreme value time series. We demonstrate the superior performance of our algorithm to the state-of-art methods, including Granger causality, copula approach, and transfer entropy, on one synthetic dataset, one climate dataset and two Twitter datasets.

stat.ME