Searcharxiv⌕ Search

arXiv subjects

Li Li

Publications and source records attributed to Li Li.

At least 37 records · Page 2Linked to original sources

Adaptive Preference Modeling via Explicit Indirect Relational Learning for Personalized Fashion Matching

Personalized fashion complementary recommendation requires jointly modeling user preferences and item compatibility under sparse and multimodal data conditions. Existing approaches often capture higher-order relational signals implicitly through graph propagation or rely on direct interaction data, limiting their ability to explicitly model indirect preference and compatibility relationships. To address this limitation, we propose an Adaptive Preference with Contrastive Learning framework (APCL) that explicitly models both direct and indirect relational signals within a unified recommendation architecture. Specifically, APCL constructs indirect user-item and item-item relationships through a correlation-guided adaptive aggregation mechanism and represents them as dedicated personalization and compatibility views. To improve representation learning, we further introduce a functional view contrastive learning strategy that aligns direct and indirect preference representations and direct and indirect compatibility representations, encouraging consistency across relational contexts. By integrating multimodal visual and textual information with explicit indirect relational modeling, APCL captures richer semantic characteristics while improving robustness in sparse-interaction settings. Experiments on two benchmark fashion recommendation datasets demonstrate that APCL consistently outperforms representative baseline methods.

cs.IR↗

HapCiD: Detecting API-related Compatibility Issues in OpenHarmony Apps

OpenHarmony, an emerging open-source mobile platform, is rapidly gaining attention in the mobile development community. Its fast-evolving Software Development Kit (SDK) introduces numerous Application Programming Interfaces (APIs) to boost developer productivity but inevitably introduces compatibility challenges, an issue well known from platforms like Android. However, existing compatibility analysis tools are ineffective for OpenHarmony due to its newly introduced ArkTS programming language and the lack of a mature behavioral model to capture app execution semantics. To bridge this gap, we present HapCiD, an open-source tool for automatically detecting API-related compatibility issues in OpenHarmony apps. Applied to 4,478 apps, HapCiD successfully detects 2,040 compatibility issues across 163 apps with 100 percent accuracy. Based on these findings, we conduct a comprehensive study of compatibility issues, categorizing their types, examining existing mitigation strategies, and proposing future improvements.

cs.SE↗

RoofSeg: An edge-aware transformer-based network for end-to-end roof plane segmentation

Roof plane segmentation is one of the key procedures for reconstructing three-dimensional (3D) building models at levels of detail (LoD) 2 and 3 from airborne light detection and ranging (LiDAR) point clouds. The majority of current approaches for roof plane segmentation rely on the manually designed or learned features followed by some specifically designed geometric clustering strategies. Because the learned features are more powerful than the manually designed features, the deep learning-based approaches usually perform better than the traditional approaches. However, the current deep learning-based approaches have three unsolved problems. The first is that most of them are not truly end-to-end, the plane segmentation results may be not optimal. The second is that the point feature discriminability near the edges is relatively low, leading to inaccurate planar edges. The third is that the planar geometric characteristics are not sufficiently considered to constrain the network training. To solve these issues, a novel edge-aware transformer-based network, named RoofSeg, is developed for segmenting roof planes from LiDAR point clouds in a truly end-to-end manner. In the RoofSeg, we leverage a transformer encoder-decoder-based framework to hierarchically predict the plane instance masks with the use of a set of learnable plane queries. To further improve the segmentation accuracy of edge regions, we also design an Edge-Aware Mask Module (EAMM) that sufficiently incorporates planar geometric prior of edges to enhance its discriminability for plane instance mask refinement. In addition, we propose an adaptive weighting strategy in the mask loss to reduce the influence of misclassified points, and also propose a new plane geometric loss to constrain the network training.

cs.CV↗

At Most One Inner Killing Horizon in Stationary Black Holes

Stationary-axisymmetric black holes are the standard theoretical framework for astrophysical black holes, yet their interiors remain largely inaccessible: the absence of a hypersurface-orthogonal timelike Killing vector renders the metric non-diagonal, and frame dragging obstructs the usual techniques for probing the interior. We prove that any stationary-axisymmetric black hole satisfying the classical energy conditions contains at most one nondegenerate inner Killing horizon per connected interior branch, irrespective of its matter content. For compact horizon sections, we apply the Raychaudhuri equation to the congruence normal to constant-time hypersurfaces: the strong energy condition makes the squared lapse subharmonic, and the maximum principle excludes a second inner horizon. For noncompact planar and hyperbolic sections central to holography, the null energy condition instead yields a monotonicity obstruction through a geometric flow we introduce here, the ``inverse lapse flow,'' whose weak continuation via viscosity solutions we also construct. The attractive nature of gravity, encoded in these energy conditions, thus both drives horizon formation and caps the number of inner horizons, revealing a striking dual role of the same focusing mechanism.

gr-qc↗

Channel-Resolved High-Overtone Quasinormal Modes Decode Kasner Scaling inside Black Holes

The high-overtone quasinormal-mode spectrum encodes the near-singularity Kasner scaling of black-hole interiors. We show that distinct structures in this spectrum independently determine the temporal and spatial Kasner grades, allowing the corresponding metric exponents to be reconstructed without imposing Kasner constraints or invoking holography. We demonstrate this independent reconstruction at the fixed Kasner point of five-dimensional asymptotically flat Schwarzschild and planar Schwarzschild--AdS$_5$, and then extend it to a continuously varying Gibbons--Maeda black-hole family. The reconstructed exponents satisfy the Kasner relation a posteriori, providing a nontrivial geometric consistency check. More generally, the high-overtone spectrum separates into global spectral data and local algebraic corrections whose fractional powers are fixed by the near-singularity Kasner geometry. These fractional corrections therefore provide a direct spectral probe of the local geometry behind the horizon.

hep-th↗

VRL-Bench: Benchmarking agents on computer control tasks under finite trial budgets

Learning from trial and error is a promising way to improve language agents on complex tasks such as computer control. Reflexion introduced verbal reinforcement learning, which turns failed trials into text that guides later attempts without updating model parameters. We introduce VRL-Bench, a harness for fair evaluation of trial-and-error learning under finite trial budgets. Across three models on MiniWoB and WebShop, we evaluate updates from several prominent verbal-memory methods spanning Reflexion and later work: each improves observed success over memory-free retry in some settings but reduces it in others. Replay experiments show that using reflection can reduce success rates, revealing a trade-off between exploiting experience and continued exploration. We propose VEX$^2$, a verbal exploration--exploitation scheduler that uses a language model to jointly select policies and allocate the remaining trial budget. VEX$^2$ is the only evaluated update to achieve positive observed success-rate gains over retry in all six settings.

cs.AI↗

Self-powered InAs nanowire detector arrays for extended-SWIR spectrometry at room temperature

Spectral sensing in the extended shortwave infrared (e-SWIR) is important for molecular analysis, infrared imaging, and machine vision, motivating the development of compact spectrometers for broader applications. However, conventional commercial off-the-shelf spectrometers in this wavelength region are expensive and bulky due to their reliance on external dispersive optics/filters and/or cryogenic accessories. Other emerging computational spectrometers are based on Si and InGaAs photodetectors that remain focused on the visible and near-infrared, with few detector platforms operating in the e-SWIR regime that simultaneously provide broadband sensitivity, low-noise room-temperature operation, and diverse spectral signatures for accurate identification and reconstruction. Here, we report a room-temperature e-SWIR computational spectrometer based on InAs/InP core-shell nanowire photodetector arrays with geometry-encoded spectral responses. The detectors exhibit self-powered broadband photoresponse across the 1--3 $μ$m range, with responsivity up to 0.215 A W$^{-1}$, detectivity up to $1.6 \times 10^{9}$ cm Hz$^{1/2}$ W$^{-1}$, and microsecond response times. The excellent detector performance is leveraged to demonstrate filter-free spectral reconstruction using a compact multipixel photodetector array device. This enables high-accuracy molecular absorption spectrum reconstruction and hyperspectral imaging. Our results indicate that InAs nanowire arrays are a promising platform for compact computational spectrometry and imaging in the e-SWIR at room temperature.

physics.optics↗

Graded Betti numbers of general curves of large degree

Let $C$ be a smooth projective complex curve of genus $g$ and gonality $k$, and $L$ be a very ample line bundle on $C$. When $L$ has sufficiently large degree, the vanishing and nonvanishing of the Koszul cohomology groups $K_{p,q}(C,L)$ have been determined previously, but the exact values of the graded Betti numbers $κ_{p,q}(C, L)$ remain largely unknown. In this paper, we give explicit closed formulas for all graded Betti numbers $κ_{p,q}(C, L)$ when the Brill--Noether locus $W_k^1(C)$ has the expected dimension and $H^1(C, L \otimes ω_C^{-1})=0$. Consequently, we determine the complete Betti table for a general curve when $°L \geq 4g-3$ or when $°L \geq 3g-3$ and $L$ is general. We also explicitly compute the Boij--Söderberg coefficient of the section ring $R(C, L)$ governing asymptotic purity, and show eventual monotonicity of the remaining coefficients: they decrease for hyperelliptic curves and increase under a natural generic reducedness assumption on the relevant Brill--Noether loci.

math.AG↗

Clinician-Friendly Foundation Models for Ophthalmic Image Diagnostics without Fine-Tuning or Technical Barriers

Artificial intelligence (AI) shows remarkable potential in medical imaging diagnostics, yet most current models require retraining when applied across different clinical settings, limiting their scalability. We developed GlobeReady, a deployment-oriented platform powered by the RetiGlobe foun- dation model and local feature augmentation. RetiGlobe was pretrained in two stages: 1) self-supervised learning using DINOv2 on 38 million synthetic ophthalmic images, and 2) contrastive learning using CLIP on 475,845 real image-text pairs spanning diverse ethnicities, imaging devices, and geographic regions worldwide. We evaluate GlobeReady on 488,448 ophthalmic images, including color fundus photographs (CFPs) and optical coherence tomography scans, from multi-centres in China, Singapore, Vietnam and the UK. Prospective testing included usability assessment with 31 ophthalmologists. Exploratory analyses evaluated domain generalisability, Bayesian uncertainty quantification, out-of-distribution (OOD) detection, and feature-based case retrieval.

cs.CV↗

Signless Laplacian spectral conditions for rainbow matchings in a collection of bipartite graphs

Let ${\cal G}=\{G_1,\ldots,G_k\}$ be a collection of (not necessarily distinct) bipartite graphs on the same vertex bipartition $(X,Y)$, where $|X|=a$, $|Y|=b$ and $2\le k\le a\le b$. A \emph{rainbow matching} of ${\cal G}$ is a set of pairwise disjoint edges that can be chosen from distinct members of ${\cal G}$. Denote by $q(G)$ the signless Laplacian spectral radius of a graph $G$. In this paper, we prove that if $q(G_i)\ge b+k-1$ for each $i\in\{1,2,\ldots,k\}$, then ${\cal G}$ admits a rainbow matching of size $k$ unless $G_1=\cdots=G_k\cong K_{k-1,b}\cup\overline{K_{a-k+1}}$, and show that the threshold is sharp and attained by the exceptional collection. The condition is also extended to larger collections for a prescribed level $t$. In addition, we obtain a lower bound for the rainbow matching number in terms of the ordered signless Laplacian spectral radii of the members, and provide a stability version of the extremal characterization. In the proofs, we use the shifting technique and a quotient matrix arising from an equitable partition of a signless Laplacian matrix.

math.CO↗

Extending Fill-In-the-Middle with Instructions for Steerable Code Completion

Code completion models often fail when the developer's intent is under-specified in the code context. To mitigate this, developers frequently use natural language comments to clarify objectives. However, current code completion models fail to prioritize these directives effectively since they are merely pre-trained using the Fill-In-the-Middle (FIM) objective. On the one hand, the natural language instructions, mixed with the noisy code comments, are just treated as part of the background context within the prefix. On the other hand, the pre-training datasets for the FIM objective are mostly sourced from open-source repositories, which results in a scarcity of high-intent instruction-to-code pairings that reflect the developers' workflow in code completion. To bridge this gap, we propose Instruction-aware Fill-In-the-Middle (IFIM), a fine-tuning method that extends the FIM structure with a dedicated, structurally separated instruction section. Our evaluation shows that IFIM substantially improves adherence to developer intent, while leaving infilling performance unchanged when no instruction is given. The gains hold on an in-the-wild benchmark of 100 instructions written by real developers and across model scales from 1.5B to 7B. IFIM thus offers a backward-compatible upgrade path for existing FIM-based code completion systems at a modest training cost.

cs.SE↗

ProGVC: Progressive-based Generative Video Compression via Auto-Regressive Context Modeling

Perceptual video compression leverages generative priors to reconstruct realistic textures and motions at low bitrates. However, existing perceptual codecs often lack native support for variable bitrate and progressive delivery, and their generative modules are weakly coupled with entropy coding, limiting bitrate reduction. Inspired by the next-scale prediction in the Visual Auto-Regressive (VAR) models, we propose ProGVC, a Progressive-based Generative Video Compression framework that unifies progressive transmission, efficient entropy coding, and detail synthesis within a single codec. ProGVC encodes videos into hierarchical multi-scale residual token maps, enabling flexible rate adaptation by transmitting a coarse-to-fine subset of scales in a progressive manner. A Transformer-based multi-scale autoregressive context model estimates token probabilities, utilized both for efficient entropy coding of the transmitted tokens and for predicting truncated fine-scale tokens at the decoder to restore perceptual details. Extensive experiments demonstrate that as a new coding paradigm, ProGVC delivers promising perceptual compression performance at low bitrates while offering practical scalability at the same time.

cs.CV↗

Understanding Deep Learning via Entropy Space Theory

Deep learning is often criticized for its theoretical research lagging behind practice. To make deep learning easier to understand, the entropy space theory is first introduced here. The entropy space can cover all the possibilities of any deep learning model by topological structure. It is independent of network parameters. Through the designed fundamental operations and norm, entropy space is proven to be a normed space within the formal axiomatic framework. Based on the theory, a unified coordinate system is proposed. It can coordinatize every state of a model and rank them by compression of the maximal value of information entropy. The theory offers a novel priori framework for mathematical fundamentals of deep learning.

cs.AI↗

Transgression Completion for Chern-Simons Black Hole Thermodynamics

The quantum statistical relation, which identifies the thermodynamic free energy with the on-shell Euclidean action, is a cornerstone of black hole thermodynamics. In this Letter, we demonstrate that it fails for charged magnetic black holes in five-dimensional Einstein-Maxwell-Chern-Simons theory unless the action is formulated in a gauge-invariant manner. In particular, the free energy computed via the quantum statistical relation violates both the first law of thermodynamics and hydrodynamics. We identify the missing finite contribution---a nonlocal radial integral---whose inclusion restores exact thermodynamic consistency. Remarkably, this term is not an ad hoc correction: combined with the local Chern-Simons representative, it forms the transgression form, yielding a strictly gauge-invariant action. The correct action is therefore not the local Chern-Simons term alone but its transgression completion---a physical requirement rather than a mere mathematical refinement. The same mechanism persists for helical, rotating, scalarized, higher-derivative, and higher-dimensional black holes. Our work establishes that topological interactions can fundamentally alter the relation between the Euclidean action and the thermodynamic potential and identifies the gauge-invariant transgression form as the correct framework for black hole thermodynamics in Chern-Simons-coupled theories.

gr-qc↗

FLM: Frequency-Aware Language Models for Generative Image Compression

Generative models have significantly improved the performance ceiling of image lossy compression at low bitrates by exploiting learned priors. However, the generated textures and semantic details may deviate from the source content, thereby affecting the fidelity of image reconstruction. To solve these challenges, we propose FLM, a frequency-aware language model that improves compression efficiency through frequency-domain probabilistic modeling while retaining deterministic reconstruction. At the encoder, the input image is transformed into quantized DCT coefficients, which are organized into discrete sequences using macroblock-based coefficient tokenization. FLM then performs next-coefficient prediction to autoregressively estimate token-wise conditional probability distributions for arithmetic coding, thereby generating a compact bitstream. At the decoder, the LLM and arithmetic decoder jointly recover the frequency-domain data, followed by inverse transformations for image reconstruction. A task-specific frequency-domain dataset and a two-stage fine-tuning strategy are further developed to enable the model to operate across multiple bitrate settings. FLM is a versatile compressor that is compatible with both lossy compression and lossless JPEG recompression frameworks. Experiments show that FLM exceeds conventional and generative lossy compression methods in rate-distortion performance. FLM achieves BD-PSNR gains of 3.30 dB, 3.83 dB, and 3.80 dB than JPEG baseline on Kodak, Tecnick, and CLIC2020, respectively. Better qualitative quality of FLM can be achieved in improving semantically high fidelity and suppressing blocking artifacts. FLM is also validated to be applicable to the lossless recompression task with competitive performance.

cs.CV↗

Frequency-Aware Continual Learning for Smart Contract Vulnerability Detection with Large Language Models

Smart contract vulnerability detection with Large Language Models (LLMs) faces three causally linked challenges. First, new vulnerability categories demand parameter-efficient adaptation, since full retraining is prohibitive for sequentially arriving tasks. Second, training per-task adapters on a shared backbone causes catastrophic forgetting of previously learned vulnerabilities. Third, the resulting multiplicity of adapters must be consolidated into a single model, since task identity is unknown at inference time. Each challenge arises directly from the solution to its predecessor, making an integrated framework essential. We propose a three-stage pipeline in which each stage addresses one challenge and feeds into the next. The adaptation stage uses Frequency-Aware Low-Rank Adaptation (FA-LoRA), which performs adaptation in the Fourier domain with per-frequency importance gates, requiring only 0.4% trainable parameters while outperforming standard LoRA and QLoRA. The continual learning stage applies Forget-Aware Replay (FAR), which uses these frequency gates to estimate per-sample forgetting risk via loss dynamics and prioritizes vulnerable knowledge for rehearsal, achieving an average Micro-F1 of 0.8022 across sequential tasks. The deployment stage employs Anchor-Protected Progressive Merging (APPM), which exploits the asymmetric generalization produced by FAR training to identify the strongest-generalizing adapter as an anchor and consolidates all adapters into a single model via anchor-protected weighted merging with frequency-domain gate competition. APPM achieves a Micro-F1 of 0.8085, within 2.7% of the independent per-task upper bound, at a merge cost of 156 ms and no additional runtime memory. Experiments on DIVE confirm the framework effectively addresses all three challenges for evolving blockchain ecosystems.

cs.AI↗

OpenHarmony Bench: Evaluating LLMs and Coding Agents on OpenHarmony App Development

We present OPENHARMONY BENCH, an app-level coding benchmark for evaluating LLM-based coding agents on OpenHarmony ArkTS applications. Unlike function-level benchmarks, it evaluates complete app-level changes: each task requires an agent to modify a buildable ArkTS project so that a requested behavior works end to end, involving UI state, data persistence, build configuration, and platform APIs. The benchmark installs and drives the delivered application on a device to check whether the behavior is observable. It covers three input sources: natural-language feature requests (new-feature), structured scenario specifications (spec-driven), and bug descriptions (bug-fix). The benchmark contains 153 top-level tasks and 242 Feature points (F-points), where an F-point is one executable behavior check. The snapshot includes 32 new-feature tasks, 50 spec-driven tasks with 139 F-points, and 71 bug-fix tasks. The main leaderboard is scored over top-level tasks rather than independently weighted F-points. We describe the benchmark construction, statistics, and build-and-test evaluation pipeline, and evaluate DevEco Code with eight LLMs across three independent full-suite runs per configuration. Three findings emerge. First, newer generations complete more tasks than their predecessors within evaluated model-family pairs. Second, buildability is close to saturated while behavioral correctness is not: mean Final Build Success Rate is 94.77% to 100.00%, whereas mean Task Completion is 48.36% to 58.39%. Third, spec-driven tasks have the lowest Task Completion under all-checks task scoring, with no configuration exceeding 35%. The code, data, tasks, reference solutions, tests, evaluation scripts, and leaderboard are released through the official OPENHARMONY BENCH website at https://bench.matrix.openharmony.cn/.

cs.SE↗

Statistically Steady Holographic Quantum Turbulence: Hyperuniform Vortex Matter and Crossover

A long-standing obstacle in quantum turbulence has been the difficulty of sustaining robust statistical steady states, preventing unambiguous identification of universal vortex organization and kinetic scaling. We construct such a steady state in two-dimensional holographic superfluid turbulence by continuous Landau-instability driving, sustaining ${\sim}2500$ vortices free from transient artifacts. The topological charge structure factor $S_c(k)$ reveals Class I disordered hyperuniformity with $S_c(k)\propto k^{α>1}$ as $k\to0$, where $k$ is the wavenumber. This constitutes the strongest long-range order of its kind and its first observation in a strongly driven, far-from-equilibrium quantum fluid with topological defects as the organizing principle, establishing a novel non-equilibrium vortex phase. Exploiting this platform, we resolve the scaling controversy: the apparent $k^{-5/3}$ signature in the kinetic energy spectrum is a narrow crossover between the $k^{-1}$ single-vortex and $k^{-3}$ core regimes, not a genuine Kolmogorov inertial range. The real-space second-order structure function provides decisive evidence via $S_2(r)\propto \ln r$ where $r$ is the spatial separation, with no $r^{2/3}$ Kolmogorov scaling, ruling out a true inertial cascade. These findings reveal that strongly coupled quantum turbulence lacks an inverse cascade due to the absence of macroscopic Onsager clusters, demonstrating energy transport fundamentally distinct from weakly coupled superfluids.

hep-th↗