SearcharxivSearch

arXiv subjects

Lu Han

Publications and source records attributed to Lu Han.

At least 19 recordsLinked to original sources

GeoStore: Finding Small Storefronts in Large Scenes -- A Fine-Grained POI Localization Benchmark with Global-to-Local Asymmetric Matching

Point-of-interest (POI) localization -- matching a user's close-up storefront photograph against large-scale geo-tagged street-view imagery -- underpins map construction, POI verification, and location-based services. Its closest existing paradigm, visual place recognition (VPR), assumes symmetric, whole-image matching of the same scene at a comparable scale; POI localization instead must match a close-up query, in which the target fills the frame, against wide references in which the same POI occupies only a small, off-center region among visually similar shops, under a substantial capture-domain gap. We introduce GeoStore, to our knowledge the first benchmark dedicated to this asymmetric, fine-grained, open-set formulation, and show that global-descriptor methods tuned for symmetric VPR are systematically limited on it, since a single global vector dilutes the small target. We further propose GLAM (Global-to-Local Asymmetric Matching), which couples a retrieval-anchoring global descriptor with an asymmetric local pathway: each reference is kept as a compact set of pooled region tokens and matched against a single query probe through a learnable soft late interaction; at inference, the same tokens enable a lightweight mutual-nearest-neighbor re-ranking. GLAM surpasses strong global and two-stage baselines on Recall@1/5/10 and mAP, with ~5x smaller re-ranking features and ~two orders of magnitude lower per-pair matching cost than prior local re-ranking. The benchmark and code will be publicly released.

cs.CV

Quasi-Sinusoidal Single Diamond Structure in Royal Jewel Butterfly: An Angle-Independent Photonic Structure

Structural colouration with narrow spectral photonic bandwidth and high reflectivity is of critical importance for modern optical applications, including displays, laser systems, and optical sensing, etc. Achieving such angle independent colouration typically relies on polycrystalline or inherent structural disorder. However, balancing angular uniformity with high brightness and strong colour contrast remains challenging. Herein, we uncover the structural origin of the spectacular bright, angle-independent blue colouration of Hypochrysops polycletus, a sapphire-like Royal Jewel butterfly. Three-dimensional (3D) electron microscopy reveals that the dorsal wing scale has a single diamond structure, a 3D photonic crystal previously documented only in beetles and weevils. The crystal domains form an extraordinary quasi sinusoidal surface geometry with a distinct template morphology-guided arrangement. Unlike typically thicker biophotonic structures that support multiple high symmetry stopbands, this design contains only 3-4 unit cells in the propagation direction. Its optical response is dominated by the fundamental stopband, with two dominant scattering mechanisms: specular reflection at the {111} inclined sidewalls of the hierarchical structure, and funnelling into localised quasi-normal modes enabled by a strongly anisotropic Bloch transport. By mimicking these features with two-photon polymerisation, we artificially reproduced the optical response in the infrared region. The study opens a pathway towards bioinspired brilliant diffuse colouration and angle-robust photonic devices.

physics.optics

Complete Structural Determination of Mesostructural Dodecagonal Quasicrystalline Particles

Quasicrystals have revolutionized our understanding of order in solids by demonstrating exotic structural and physicochemical properties with diverse potential applications. Despite the development of various theoretical models and experimental techniques to describe quasicrystal structures, the precise determination of local three dimensional (3D) arrangements of constituent atoms, or of secondary building units such as clusters or micelles, remains elusive. This challenge is particularly acute in self assembled soft matter quasicrystalline systems, where the complex assembly of molecular groups introduces additional defects and structural modulations. Herein, we report the first complete structural determination of self-assembled mesostructural dodecagonal quasicrystalline particles. Employing advanced electron tomography, combined with dedicated structural tracing and processing workflows, the 3D coordinates of all nodal sites were extracted. This approach reveals that the actual structure deviates from the conventionally assumed tetrahedral close packing geometry, exhibiting diverse coordination environments and displacive fluctuations. We identified and quantified rotational intergrowths arising from node exchange, as well as various defects and disorder, with these features discernible only through 3D analysis. Additionally, we propose a simplified two-layer stacking of isomorphic hexagonal model to form dodecagonal quasicrystal. This work advances our understanding of soft-matter dodecagonal quasicrystals and paves the way for detailed structural elucidation of self-assembled systems.

cond-mat.mes-hall

SepSeq: A Training-Free Framework for Long Numerical Sequence Processing in LLMs

While transformer-based Large Language Models (LLMs) theoretically support massive context windows, they suffer from severe performance degradation when processing long numerical sequences. We attribute this failure to the attention dispersion in the Softmax mechanism, which prevents the model from concentrating attention. To overcome this, we propose Separate Sequence (SepSeq), a training-free, plug-and-play framework to mitigate dispersion by strategically inserting separator tokens. Mechanistically, we demonstrate that separator tokens act as an attention sink, recalibrating attention to focus on local segments while preserving global context. Extensive evaluations on 9 widely-adopted LLMs confirm the effectiveness of our approach: SepSeq yields an average relative accuracy improvement of 35.6% across diverse domains while reducing total inference token consumption by 16.4% on average.

cs.CL

Rethinking ANN-based Retrieval: Multifaceted Learnable Index for Large-scale Recommendation System

Approximate nearest neighbor (ANN) search is widely used in the retrieval stage of large-scale recommendation systems. In this stage, candidate items are indexed using their learned embedding vectors, and ANN search is executed for each user (or item) query to retrieve a set of relevant items. However, ANN-based retrieval has two key limitations. First, item embeddings and their indices are typically learned in separate stages: indexing is often performed offline after embeddings are trained, which can yield suboptimal retrieval quality-especially for newly created items. Second, although ANN offers sublinear query time, it must still be run for every request, incurring substantial computation cost at industry scale. In this paper, we propose MultiFaceted Learnable Index (MFLI), a scalable, real-time retrieval paradigm that learns multifaceted item embeddings and indices within a unified framework and eliminates ANN search at serving time. Specifically, we construct a multifaceted hierarchical codebook via residual quantization of item embeddings and co-train the codebook with the embeddings. We further introduce an efficient multifaceted indexing structure and mechanisms that support real-time updates. At serving time, the learned hierarchical indices are used directly to identify relevant items, avoiding ANN search altogether. Extensive experiments on real-world data with billions of users show that MFLI improves recall on engagement tasks by up to 11.8\%, cold-content delivery by up to 57.29\%, and semantic relevance by 13.5\% compared with prior state-of-the-art methods. We also deploy MFLI in the system and report online experimental results demonstrating improved engagement, less popularity bias, and higher serving efficiency.

cs.IR

Spin-polarized chiral ZnIn2S4 for targeted solar-driven CO2 reduction to acetic acid

Acetic acid, an important industrial chemical, is a key target product for CO2 reduction due to its dual role in carbon utilization and chemical feedstock supply. Although photocatalytic CO2 reduction (PCCR) can generate acetic acid alongside other multicarbon products, its yield is typically low, limited by competing reactions and inefficient C-C coupling. Herein, we report a chiral mesostructured ZnIn2S4 (CMZI) photocatalyst that achieves a remarkable acetic acid yield of 962 {umol g-1 h-1 with a high selectivity of 97.3 %. This yield is ten times higher than the current highest reported value, while attaining state-of-the-art selectivity10. The remarkable productivity arises from synergistic effect between chiral structure and sulfur (S) sites of CMZI. Chirality-induced spin polarization in CMZI stabilizes the key triplet OCCO intermediate, significantly promoting C-C coupling efficiency. Theoretical calculations reveal that the S sites on {102} crystal facets of ZnIn2S4 exhibit thermodynamic and kinetic preferences for acetic acid formation. This work offers critical insights into catalytic strategies for CO2 reduction toward the efficient and scalable synthesis of various multicarbon products.

cond-mat.mtrl-sci

One-Embedding-Fits-All: Efficient Zero-Shot Time Series Forecasting by a Model Zoo

The proliferation of Time Series Foundation Models (TSFMs) has significantly advanced zero-shot forecasting, enabling predictions for unseen time series without task-specific fine-tuning. Extensive research has confirmed that no single TSFM excels universally, as different models exhibit preferences for distinct temporal patterns. This diversity suggests an opportunity: how to take advantage of the complementary abilities of TSFMs. To this end, we propose ZooCast, which characterizes each model's distinct forecasting strengths. ZooCast can intelligently assemble current TSFMs into a model zoo that dynamically selects optimal models for different forecasting tasks. Our key innovation lies in the One-Embedding-Fits-All paradigm that constructs a unified representation space where each model in the zoo is represented by a single embedding, enabling efficient similarity matching for all tasks. Experiments demonstrate ZooCast's strong performance on the GIFT-Eval zero-shot forecasting benchmark while maintaining the efficiency of a single TSFM. In real-world scenarios with sequential model releases, the framework seamlessly adds new models for progressive accuracy gains with negligible overhead.

cs.LG

Advanced spectral clustering for heterogeneous data in credit risk monitoring systems

Heterogeneous data, which encompass both numerical financial variables and textual records, present substantial challenges for credit monitoring. To address this issue, we propose Advanced Spectral Clustering (ASC), a method that integrates financial and textual similarities through an optimized weight parameter and selects eigenvectors using a novel eigenvalue-silhouette optimization approach. Evaluated on a dataset comprising 1,428 small and medium-sized enterprises (SMEs), ASC achieves a Silhouette score that is 18% higher than that of a single-type data baseline method. Furthermore, the resulting clusters offer actionable insights; for instance, 51% of low-risk firms are found to include the term 'social recruitment' in their textual records. The robustness of ASC is confirmed across multiple clustering algorithms, including k-means, k-medians, and k-medoids, with {\Delta}Intra/Inter < 0.13 and {\Delta}Silhouette Coefficient < 0.02. By bridging spectral clustering theory with heterogeneous data applications, ASC enables the identification of meaningful clusters, such as recruitment-focused SMEs exhibiting a 30% lower default risk, thereby supporting more targeted and effective credit interventions.

cs.LG

Integrated Multivariate Segmentation Tree for Heterogeneous Credit Data Analysis in Small- and Medium-Sized Enterprises

Traditional decision tree models, which rely exclusively on numerical variables, often face challenges in handling high-dimensional data and are limited in their ability to incorporate textual information effectively. To address these limitations, we propose the integrated multivariate segmentation tree (IMST), a comprehensive framework designed to improve credit evaluation for small- and medium-sized enterprises (SMEs) by integrating financial data with textual sources. This method comprises three core stages: (1) transforming textual data into numerical matrices through matrix factorization, (2) selecting salient financial features using Lasso regression, and (3) constructing a multivariate segmentation tree based on either the Gini index or entropy, with weakest-link pruning applied to control model complexity. Experimental results based on a dataset of 1,428 Chinese SMEs demonstrated that IMST achieved an accuracy rate of 88.9%, surpassing both baseline decision trees (87.4%) and conventional models such as support vector machines and neural networks. Furthermore, the proposed model demonstrated superior interpretability and computational efficiency, featuring a more streamlined architecture and improved risk detection capabilities.

cs.LG

WTFormer: A Wavelet Conformer Network for MIMO Speech Enhancement with Spatial Cues Peservation

Current multi-channel speech enhancement systems mainly adopt single-output architecture, which face significant challenges in preserving spatio-temporal signal integrity during multiple-input multiple-output (MIMO) processing. To address this limitation, we propose a novel neural network, termed WTFormer, for MIMO speech enhancement that leverages the multi-resolution characteristics of wavelet transform and multi-dimensional collaborative attention to effectively capture globally distributed spatial features, while using Conformer for time-frequency modeling. A multi task loss strategy accompanying MUSIC algorithm is further proposed for optimization training to protect spatial information to the greatest extent. Experimental results on the LibriSpeech dataset show that WTFormer can achieve comparable denoising performance to advanced systems while preserving more spatial information with only 0.98M parameters.

eess.AS

UniCA: Unified Covariate Adaptation for Time Series Foundation Model

Time Series Foundation Models (TSFMs) have achieved remarkable success through large-scale pretraining. However, their design primarily targets real-valued series, limiting their ability to handle general forecasting tasks involving diverse and often heterogeneous covariates -- such as categorical variables and multimodal data (e.g., images, text) -- which are typically task-specific and difficult to leverage during pretraining. To address this gap, we propose Unified Covariate Adaptation (UniCA), a framework to bridge TSFMs with general covariate-aware forecasting. UniCA first performs covariate homogenization to transform heterogeneous covariates into high-level homogeneous series representations and then fuses them via a unified attention-based fusion mechanism. UniCA is compatible and universal for adaptation with both homogeneous and heterogeneous covariates, incorporating extra covariate information while preserving the generalization ability of TSFMs.Extensive experiments on multiple unimodal and multimodal covariate-aware forecasting benchmarks demonstrate the superiority of UniCA, highlighting the promise of covariate-aware TSFM adaptation in real-world forecasting scenarios.Code: https://github.com/hanlu-nju/UniCA.

cs.LG

$La_3Pd_2NaO_9$: A High-Valent Insulating Palladate

A high-valent palladate, $La_3Pd_2NaO_9$, has been synthesized for the first time. Single crystals with dimensions of 20 ${\mu}$m on edge were successfully grown using the flux method at 420 $^o$C and 70 bar oxygen pressure. Energy dispersive spectroscopy (EDS) and inductively coupled plasma mass spectroscopy (ICP) measurements show that the atomic ratio of La: (Pd+Na) is 3: 3 and Pd: Na is 2: 1. X-ray photoelectron spectroscopy (XPS) measurements show that the oxidation state of Pd is dominated by +4. Synchrotron X-ray single-crystal diffraction measurements revealed that this material crystallizes in the monoclinic $P2_1/c$ space group with charge ordering of Na and Pd. Real-space imaging via scanning transmission electron microscopy (STEM) confirmed the crystal structure and revealed excellent sample homogeneity. Electrical resistivity measurements show an insulating behavior. Magnetic measurements show an unexpected paramagnetic behavior, which probably originate from a small fraction of high-spin Pd$^{2+}$ evidenced by XPS. The successful growth of $La_3Pd_2NaO_9$ single crystals with a high-valent oxidation state of Pd offers an approach for exploring interesting palladates, including potential bilayer Ruddlesden-Popper palladates analogous to the high temperature superconducting $La_3Ni_2O_7$.

cond-mat.str-el

OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning

While anomaly detection has made significant progress, generating detailed analyses that incorporate industrial knowledge remains a challenge. To address this gap, we introduce OmniAD, a novel framework that unifies anomaly detection and understanding for fine-grained analysis. OmniAD is a multimodal reasoner that combines visual and textual reasoning processes. The visual reasoning provides detailed inspection by leveraging Text-as-Mask Encoding to perform anomaly detection through text generation without manually selected thresholds. Following this, Visual Guided Textual Reasoning conducts comprehensive analysis by integrating visual perception. To enhance few-shot generalization, we employ an integrated training strategy that combines supervised fine-tuning (SFT) with reinforcement learning (GRPO), incorporating three sophisticated reward functions. Experimental results demonstrate that OmniAD achieves a performance of 79.1 on the MMAD benchmark, surpassing models such as Qwen2.5-VL-7B and GPT-4o. It also shows strong results across multiple anomaly detection benchmarks. These results highlight the importance of enhancing visual perception for effective reasoning in anomaly understanding. All codes and models will be publicly available.

cs.CV

AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion

Diffusion models have significantly improved the quality and diversity of audio generation but are hindered by slow inference speed. Rectified flow enhances inference speed by learning straight-line ordinary differential equation (ODE) paths. However, this approach requires training a flow-matching model from scratch and tends to perform suboptimally, or even poorly, at low step counts. To address the limitations of rectified flow while leveraging the advantages of advanced pre-trained diffusion models, this study integrates pre-trained models with the rectified diffusion method to improve the efficiency of text-to-audio (TTA) generation. Specifically, we propose AudioTurbo, which learns first-order ODE paths from deterministic noise sample pairs generated by a pre-trained TTA model. Experiments on the AudioCaps dataset demonstrate that our model, with only 10 sampling steps, outperforms prior models and reduces inference to 3 steps compared to a flow-matching-based acceleration model.

cs.SD

Superconductivity and phase diagram in Sr-doped La$_{3-x}$Sr$_{x}$Ni$_2$O$_7$ thin films

Recent studies have demonstrated ambient pressure superconductivity in compressively strained La$_{3}$Ni$_{2}$O$_{7}$ thin films, yet the phase diagram of heterovalent doping$-$critical for advancing the field$-$remains unexplored. Here, we report superconductivity in Sr$^{2+}$-doped La$_{3-x}$Sr$_{x}$Ni$_2$O$_7$ films synthesized via molecular beam epitaxy with ozone-assisted post-annealing. The superconducting transition temperature ($T_{\mathrm{c}}$) follows an asymmetric dome-like profile, persisting across a wide doping range ($0 \leq x \leq 0.21$) before diminishing at $x \approx 0.38$. Optimally doped films ($x = 0.09$) achieve $T_{\mathrm{c}}$ of $\sim$ 42 K, with high critical current ($J_{\mathrm{c}} > 1.4$ $\mathrm{kA/cm^{2}}$ at 2 K) and upper critical fields ($\mu_{0}H_{\mathrm{c,\parallel}}(0)= 83.7$ $\mathrm{T}$, $\mu_{0}H_{\mathrm{c,\perp}}(0)= 110.3$ $\mathrm{T}$), comparable to reported La$_{3-x}$Pr$_{x}$Ni$_2$O$_7$ films. Scanning transmission electron microscopy reveals oxygen vacancies predominantly occupy at planar NiO$_{2}$ sites$-$unlike apical-site vacancies in bulk samples$-$due to Coulomb repulsion destabilizing planar oxygen under compressive strain. Additionally, the elongated out-of-plane Ni-O bonds, exceeding those in pressurized bulk samples by $4\%$, likely weaken the interlayer $d_{z^2}$ coupling, thus contributing to the reduced $T_{\mathrm{c}}$ in strained films. This work establishes heterovalent Sr$^{2+}$ doping as a robust tuning parameter for nickelate superconductivity, unveiling a unique phase diagram topology.

cond-mat.supr-con

DongbaMIE: A Multimodal Information Extraction Dataset for Evaluating Semantic Understanding of Dongba Pictograms

Dongba pictographic is the only pictographic script still in use in the world. Its pictorial ideographic features carry rich cultural and contextual information. However, due to the lack of relevant datasets, research on semantic understanding of Dongba hieroglyphs has progressed slowly. To this end, we constructed \textbf{DongbaMIE} - the first dataset focusing on multimodal information extraction of Dongba pictographs. The dataset consists of images of Dongba hieroglyphic characters and their corresponding semantic annotations in Chinese. It contains 23,530 sentence-level and 2,539 paragraph-level high-quality text-image pairs. The annotations cover four semantic dimensions: object, action, relation and attribute. Systematic evaluation of mainstream multimodal large language models shows that the models are difficult to perform information extraction of Dongba hieroglyphs efficiently under zero-shot and few-shot learning. Although supervised fine-tuning can improve the performance, accurate extraction of complex semantics is still a great challenge at present.

cs.CV

Deep Learning-Assisted Fourier Analysis for High-Efficiency Structural Design: A Case Study on Three-Dimensional Photonic Crystals Enumeration

The geometric design of structures with optimized physical and chemical properties is one of the core topics in materials science. However, designing new functional materials is challenging due to the vast number of existing and the possible unknown structures to be enumerated and difficulties in mining the underlying correlations between structures and their properties. Here, we propose a universal method for periodic structural design and property optimization. The key in our approach is a deep-learning assisted inverse Fourier transform, which enables the creation of arbitrary geometries within crystallographic space groups. It effectively explores extensive parameter spaces to identify ideal structures with desired properties. Taking the research of three-dimensional (3D) photonic structures as a case study, this method is capable of modelling numerous structures and identifying their photonic bandgaps in just a few hours. We confirmed the established knowledge that the widest photonic bandgaps exist in network morphologies, among which the single diamond (dia net) reigns supreme. Additionally, this method identified a rarely-known lcs topology with excellent photonic properties, highlighting the infinitely extensible application boundaries of our approach. This work demonstrates the high efficiency and effectiveness of the Fourier-based method, advancing material design and providing insights for next-generation functional materials.

physics.optics

Molly: Making Large Language Model Agents Solve Python Problem More Logically

Applying large language models (LLMs) as teaching assists has attracted much attention as an integral part of intelligent education, particularly in computing courses. To reduce the gap between the LLMs and the computer programming education expert, fine-tuning and retrieval augmented generation (RAG) are the two mainstream methods in existing researches. However, fine-tuning for specific tasks is resource-intensive and may diminish the model`s generalization capabilities. RAG can perform well on reducing the illusion of LLMs, but the generation of irrelevant factual content during reasoning can cause significant confusion for learners. To address these problems, we introduce the Molly agent, focusing on solving the proposed problem encountered by learners when learning Python programming language. Our agent automatically parse the learners' questioning intent through a scenario-based interaction, enabling precise retrieval of relevant documents from the constructed knowledge base. At generation stage, the agent reflect on the generated responses to ensure that they not only align with factual content but also effectively answer the user's queries. Extensive experimentation on a constructed Chinese Python QA dataset shows the effectiveness of the Molly agent, indicating an enhancement in its performance for providing useful responses to Python questions.

cs.CL