SearcharxivSearch

arXiv subjects

Lei Jia

Publications and source records attributed to Lei Jia.

12 recordsLinked to original sources

Unlimited OCR Works

Recently, end-to-end OCR models, exemplified by DeepSeek OCR, have once again thrust OCR into the spotlight. A widely held view is that employing a large language model (LLM) as the decoder allows the model to leverage the prior distribution of language, leading to improved OCR performance. However, the downside is equally evident: as the output sequence lengthens, the accumulated KV cache drives up memory consumption and progressively slows down generation. This stands in stark contrast to humans, who exhibit no such decline in efficiency during long-horizon copying tasks. In this technical report, we propose Unlimited OCR, a model designed to emulate human parsing working memory. Taking DeepSeek OCR as the baseline, we replace all attention layers in the decoder with our proposed Reference Sliding Window Attention (R-SWA), which reduces attention computation costs while maintaining a constant KV cache throughout the entire decoding process. By combining the high compression rate of DeepSeek OCR's encoder with our constant KV cache design, Unlimited OCR can transcribe dozens of pages of documents in a single forward pass under a standard maximum length of 32K. More importantly, R-SWA is a general-purpose parsing attention mechanism - beyond OCR, it is equally applicable to tasks such as ASR, translation, etc. Codes and model weights are publicly available at http://github.com/baidu/Unlimited-OCR.

cs.CV

Structural Change Detection in High-Dimensional Transformed Factor Models via Canonical Correlation Analysis

This paper develops a canonical-correlation-based method for detecting structural changes in high-dimensional transformed factor models. The proposed approach exploits the low-rank canonical-correlation structure induced by dynamically dependent common factors, while serially uncorrelated idiosyncratic components correspond to a noise subspace with zero canonical correlations. We construct an eigenvalue-ratio criterion that measures residual dynamic dependence in the estimated noise subspace and identifies the true change point under sufficient separation of the regime-specific loading spaces or dynamic canonical correlation structures. Since the change-point location and the regime-specific factor numbers are both unknown, we further propose an alternating iterative estimation procedure that updates them sequentially until convergence. Under suitable mixing and moment conditions, we establish asymptotic properties of the proposed estimators, with convergence rates depending explicitly on factor strength, cross-sectional dimension, and sample size. Monte Carlo experiments and empirical applications to intraday stock returns and U.S. temperature series demonstrate the finite-sample

stat.ME

SVOM/C-GFT: Instrumentation and Performances on the SVOM Alerts

The Chinese Ground Follow-up Telescope (C-GFT) is an optical facility upgraded to support the Space Variable Objects Monitor mission (\textit{SVOM}). Located at the Jilin Observation Station, it is capable of rapidly identifying and monitoring the optical counterparts of Gamma-Ray Bursts (GRBs). The 1.2-m telescope is equipped with two switchable focal-plane instruments: the prime-focus wide-field LATIOS camera and the Cassegrain-focus three-channel CATCH camera. In this paper, we present a system overview, including the observatory, the telescope, the instrumentation, the automated operational framework managed by the Operations Center, and the data processing pipelines. We also report the performance results obtained during over one year of \textit{SVOM}'s post-launch operations. The results demonstrate that the system meets its design specifications and delivers robust observational and operational performance.

astro-ph.IM

Output-Constrained Controller with Fuzzy-Tuned Parameters for Overhead Cranes

This study proposes a fuzzy-adjusted nonlinear control method based on torque jitter output limit constraints for overhead crane systems with double pendulum effects. The proposed control method can effectively suppress swing and achieve precise positioning. Firstly, by enhancing the coupling relationship between the trolley displacement and swing angle, a composite signal with an error term was designed. Then, an energy-based Lyapunov function was constructed using the composite error signal, which incorporated a new formulation of the inertia matrix and potential energy function. Subsequently, using the backstepping method in conjunction with the hyperbolic tangent function, a controller with partial performance constraints was designed. In addition, to further enhance the system's dynamic performance, a fuzzy control scheme with online adjustable system parameters was designed. Finally, the stability of the system is proven using Lyapunov theory combined with LaSalle's invariance principle. Simulation results demonstrate that the proposed controller exhibits superior performance and robustness.

eess.SY

Lightning-Induced Faults in Low-Voltage Distribution Networks via Hybrid VTS-PEEC Method

As a critical component of power supply systems, low-voltage distribution net-works directly affect grid stability and user power supply reliability, yet they face significant threats from lightning-induced faults. Transient simulations are more economical and adaptable for investigating lightning-induced faults in low-voltage distribution networks than experiments. A hybrid Variable Time Step (VTS)-Partial Element Equivalent Circuit (PEEC) method, has been validat-ed in previous study, is used for Lightning-induced Electromagnetic Pulse (LEMP) simulation and fault analysis. The lightning-induced faults in ex-tended unequal-length double-circuit low-voltage distribution networks are ana-lyzed in this paper. The impact of lightning stroke location on overvoltage and fault risk is the primary focus of this study. Key findings indicate that, for ground strokes in front of the center of one double circuit, similar three-phase negative and bipolar oscillatory waveforms that are linked to fault initiation emerge. Closer strokes promote bipolar waveforms with the main peak negative as well as higher overvoltages and fault risk. These results provide essential insights for under-standing lightning-induced fault mechanisms, thereby laying a foundation for formulating more targeted and effective lightning protection measures.

physics.plasm-ph

ITFormer: Bridging Time Series and Natural Language for Multi-Modal QA with Large-Scale Multitask Dataset

Time-series data are critical in diverse applications, such as industrial monitoring, medical diagnostics, and climate research. However, effectively integrating these high-dimensional temporal signals with natural language for dynamic, interactive tasks remains a significant challenge. To address this, we introduce the Time-Series Question Answering (Time-Series QA) task and release EngineMT-QA, the first large-scale, multi-task, temporal-textual QA dataset designed to capture complex interactions between time-series signals and natural language. Building on this resource, we propose the Instruct Time Transformer (ITFormer), a novel framework that bridges time-series encoders with frozen large language models (LLMs). ITFormer effectively extracts, aligns, and fuses temporal and textual features, achieving a strong improvement in QA accuracy over strong baselines with fewer than 1\% additional trainable parameters. By combining computational efficiency with robust cross-modal modeling, our work establishes a adaptable paradigm for integrating temporal data with natural language, paving the way for new research and applications in multi-modal AI. More details about the project, including datasets and code, are available at: https://pandalin98.github.io/itformer_site/

cs.CL

Photometric Stellar Parameters for 195,478 Kepler Input Catalog (KIC) Stars

The stellar atmospheric parameters and physical properties of stars in the Kepler Input Catalog (KIC) are of great significance for the study of exoplanets, stellar activity, and asteroseismology. However, despite extensive effort over the past decades, accurate spectroscopic estimates of these parameters are available for only about half of the stars in the full KIC catalog. In our work, by training relationships between photometric colors and spectroscopic stellar parameters from Gaia DR3, the Kepler Issac-Newton Survey, LAMOST DR10, and APOGEE DR17, we have obtained atmospheric-parameter estimates for over 195,000 stars, accounting for 97$\%$ of the total sample of KIC stars. We obtain 1$\sigma$ uncertainties of 0.1 dex on metallicity [Fe/H], 100 K on effective temperature $T_{\rm eff}$, and 0.2 dex on surface gravity log $g$. In addition, based on these atmospheric parameters, we estimated the ages, masses, radii, and surface gravities of these stars using the commonly adopted isochrone-fitting approach. External comparisons indicate that the resulting precision for turn-off stars is 20$\%$ in age; for dwarf stars, it is 0.07 $M_{\odot}$ in mass, 0.05 $R_{\odot}$ in radius, and 0.12 dex in surface gravity; and for giant stars, it is 0.14 $M_{\odot}$ in mass, 0.73 $R_{\odot}$ in radius, and 0.11 dex in surface gravity.

astro-ph.SR

Reducing Memory Contention and I/O Congestion for Disk-based GNN Training

Graph neural networks (GNNs) gain wide popularity. Large graphs with high-dimensional features become common and training GNNs on them is non-trivial on an ordinary machine. Given a gigantic graph, even sample-based GNN training cannot work efficiently, since it is difficult to keep the graph's entire data in memory during the training process. Leveraging a solid-state drive (SSD) or other storage devices to extend the memory space has been studied in training GNNs. Memory and I/Os are hence critical for effectual disk-based training. We find that state-of-the-art (SoTA) disk-based GNN training systems severely suffer from issues like the memory contention between a graph's topological and feature data, and severe I/O congestion upon loading data from SSD for training. We accordingly develop GNNDrive. GNNDrive 1) minimizes the memory footprint with holistic buffer management across sampling and extracting, and 2) avoids I/O congestion through a strategy of asynchronous feature extraction. It also avoids costly data preparation on the critical path and makes the most of software and hardware resources. Experiments show that GNNDrive achieves superior performance. For example, when training with the Papers100M dataset and GraphSAGE model, GNNDrive is faster than SoTA PyG+, Ginex, and MariusGNN by 16.9x, 2.6x, and 2.7x, respectively.

cs.DC

Pome: Parallelizing I/Os and Computations for Efficient LSM-tree-based Data Storage

CPU computations and I/O operations are fundamental to data storage systems. Storage systems conduct computations with their user threads, such as sorting data for orderliness. They handle I/Os mainly through system calls (syscalls) including file write, read, and fsync, which the OS's kernel threads perform with storage devices. Today, LSM-tree-based storage systems are widely deployed in production environments. Compaction is an essential operation that LSM-tree employs to maintain its tiered tree-like structure by re-sorting and re-storing data through computations and I/Os,respectively. In this paper, we first overhaul the procedure of a compaction. We find that computations and I/Os execute in sequential order. After re-sorting data, the user thread waits for a kernel thread to complete file write and fsync I/Os. These costly synchronous I/Os create a severely long critical path that affects the performance of LSM-tree. To address this issue, we propose parallelizing I/Os and computations for efficient LSM-tree-based data storage (Pome). Pome decouples computations from I/Os within each compaction by referring to its new protocol that moves I/O operations out of the critical path. To this end, it leverages the io_uring to perform asynchronous I/Os. Furthermore, regarding the potential I/O congestion caused by accelerated compactions, Pome incorporates an adaptive I/O rate limiter to achieve smooth execution. We prototype Pome on top of RocksDB. Experimental results demonstrate that Pome significantly improves the performance of RocksDB and outperforms several state-of-the-art LSM-tree variants.

cs.DB

Kernel-Induced Label Propagation by Mapping for Semi-Supervised Classification

Kernel methods have been successfully applied to the areas of pattern recognition and data mining. In this paper, we mainly discuss the issue of propagating labels in kernel space. A Kernel-Induced Label Propagation (Kernel-LP) framework by mapping is proposed for high-dimensional data classification using the most informative patterns of data in kernel space. The essence of Kernel-LP is to perform joint label propagation and adaptive weight learning in a transformed kernel space. That is, our Kernel-LP changes the task of label propagation from the commonly-used Euclidean space in most existing work to kernel space. The motivation of our Kernel-LP to propagate labels and learn the adaptive weights jointly by the assumption of an inner product space of inputs, i.e., the original linearly inseparable inputs may be mapped to be separable in kernel space. Kernel-LP is based on existing positive and negative LP model, i.e., the effects of negative label information are integrated to improve the label prediction power. Also, Kernel-LP performs adaptive weight construction over the same kernel space, so it can avoid the tricky process of choosing the optimal neighborhood size suffered in traditional criteria. Two novel and efficient out-of-sample approaches for our Kernel-LP to involve new test data are also presented, i.e., (1) direct kernel mapping and (2) kernel mapping-induced label reconstruction, both of which purely depend on the kernel matrix between training set and testing set. Owing to the kernel trick, our algorithms will be applicable to handle the high-dimensional real data. Extensive results on real datasets demonstrate the effectiveness of our approach.

cs.CV

Chemi-net: a graph convolutional network for accurate drug property prediction

Absorption, distribution, metabolism, and excretion (ADME) studies are critical for drug discovery. Conventionally, these tasks, together with other chemical property predictions, rely on domain-specific feature descriptors, or fingerprints. Following the recent success of neural networks, we developed Chemi-Net, a completely data-driven, domain knowledge-free, deep learning method for ADME property prediction. To compare the relative performance of Chemi-Net with Cubist, one of the popular machine learning programs used by Amgen, a large-scale ADME property prediction study was performed on-site at Amgen. The results showed that our deep neural network method improved current methods by a large margin. We foresee that the significantly increased accuracy of ADME prediction seen with Chemi-Net over Cubist will greatly accelerate drug discovery.

cs.LG

The LAMOST Survey of Background Quasars in the Vicinity of the Andromeda and Triangulum Galaxies -- II. Results from the Commissioning Observations and the Pilot Surveys

We present new quasars discovered in the vicinity of the Andromeda and Triangulum galaxies with the LAMOST during the 2010 and 2011 observational seasons. Quasar candidates are selected based on the available SDSS, KPNO 4 m telescope, XSTPS optical, and WISE near infrared photometric data. We present 509 new quasars discovered in a stripe of ~135 sq. deg from M31 to M33 along the Giant Stellar Stream in the 2011 pilot survey datasets, and also 17 new quasars discovered in an area of ~100 sq. deg that covers the central region and the southeastern halo of M31 in the 2010 commissioning datasets. These 526 new quasars have i magnitudes ranging from 15.5 to 20.0, redshifts from 0.1 to 3.2. They represent a significant increase of the number of identified quasars in the vicinity of M31 and M33. There are now 26, 62 and 139 known quasars in this region of the sky with i magnitudes brighter than 17.0, 17.5 and 18.0 respectively, of which 5, 20 and 75 are newly-discovered. These bright quasars provide an invaluable collection with which to probe the kinematics and chemistry of the ISM/IGM in the Local Group of galaxies. A total of 93 quasars are now known with locations within 2.5 deg of M31, of which 73 are newly discovered. Tens of quasars are now known to be located behind the Giant Stellar Stream, and hundreds behind the extended halo and its associated substructures of M31. The much enlarged sample of known quasars in the vicinity of M31 and M33 can potentially be utilized to construct a perfect astrometric reference frame to measure the minute PMs of M31 and M33, along with the PMs of substructures associated with the Local Group of galaxies. Those PMs are some of the most fundamental properties of the Local Group.

astro-ph.GA