SearcharxivSearch

arXiv subjects

Kun Shi

Publications and source records attributed to Kun Shi.

16 recordsLinked to original sources

Reinforcement Learning on Pre-Training Data

The growing disparity between the exponential scaling of computational resources and the finite growth of high-quality text data now constrains conventional scaling approaches for large language models (LLMs). To address this challenge, we introduce Reinforcement Learning on Pre-Training data (RLPT), a new training-time scaling paradigm for optimizing LLMs. In contrast to prior approaches that scale training primarily through supervised learning, RLPT enables the policy to autonomously explore meaningful trajectories to learn from pre-training data and improve its capability through reinforcement learning (RL). While existing RL strategies such as reinforcement learning from human feedback (RLHF) and reinforcement learning with verifiable rewards (RLVR) rely on human annotation for reward construction, RLPT eliminates this dependency by deriving reward signals directly from pre-training data. Specifically, it adopts a next-segment reasoning objective, rewarding the policy for accurately predicting subsequent text segments conditioned on the preceding context. This formulation allows RL to be scaled on pre-training data, encouraging the exploration of richer trajectories across broader contexts and thereby fostering more generalizable reasoning skills. Extensive experiments on both general-domain and mathematical reasoning benchmarks across multiple models validate the effectiveness of RLPT. For example, when applied to Qwen3-4B-Base, RLPT yields absolute improvements of $3.0$, $5.1$, $8.1$, $6.0$, $6.6$, and $5.3$ on MMLU, MMLU-Pro, GPQA-Diamond, KOR-Bench, AIME24, and AIME25, respectively. The results further demonstrate favorable scaling behavior, suggesting strong potential for continued gains with more compute. In addition, RLPT provides a solid foundation, extending the reasoning boundaries of LLMs and enhancing RLVR performance.

cs.CL

WRT-SAM: Foundation Model-Driven Segmentation for Generalized Weld Radiographic Testing

Radiographic testing is a fundamental non-destructive evaluation technique for identifying weld defects and assessing quality in industrial applications due to its high-resolution imaging capabilities. Over the past decade, deep learning techniques have significantly advanced weld defect identification in radiographic images. However, conventional approaches, which rely on training small-scale, task-specific models on single-scenario datasets, exhibit poor cross-scenario generalization. Recently, the Segment Anything Model (SAM), a pre-trained visual foundation model trained on large-scale datasets, has demonstrated exceptional zero-shot generalization capabilities. Fine-tuning SAM with limited domain-specific data has yielded promising results in fields such as medical image segmentation and anomaly detection. To the best of our knowledge, this work is the first to introduce SAM-based segmentation for general weld radiographic testing images. We propose WRT-SAM, a novel weld radiographic defect segmentation model that leverages SAM through an adapter-based integration with a specialized prompt generator architecture. To improve adaptability to grayscale weld radiographic images, we introduce a frequency prompt generator module, which enhances the model's sensitivity to frequency-domain information. Furthermore, to address the multi-scale nature of weld defects, we incorporate a multi-scale prompt generator module, enabling the model to effectively extract and encode defect information across varying scales. Extensive experimental evaluations demonstrate that WRT-SAM achieves a recall of 78.87%, a precision of 84.04%, and an AUC of 0.9746, setting a new state-of-the-art (SOTA) benchmark. Moreover, the model exhibits superior zero-shot generalization performance, highlighting its potential for practical deployment in diverse radiographic testing scenarios.

cs.CV

Dependence of Electrostatic Patch Force Evaluation on the Lateral Resolution of Kelvin Probe Force Microscopy

Kelvin Probe Force Microscopy (KPFM) is widely used to measure the surface potential on samples, from which electrostatic patch force can be calculated. However, since the KPFM measurements represent a weighted average of local potentials on the sample, the accuracy of the evaluation critically depends on the precision and lateral resolution of the method. In this paper, we investigate the influence of this averaging effect on patch force estimations using both analytic and numerical methods. First, we derive the correlation functions of patch potential and establish the formulas for calculating the electrostatic patch forces in the parallel-plate geometry, with and without consideration of the KPFM measurement effect. Thus, an analytic method is established to determine the accuracy of patch force evaluation when the statistical parameters of the patch potential and the lateral resolution of the KPFM are given. Second, numerical simulations are employed to explore the dependence of estimated patch forces on the KPFM's lateral resolution under more realistic conditions. Both analytic and numerical results show a similar dependence of the patch force estimation on the patch characteristic size, potential fluctuation and the lateral resolution of the KPFM. It is also found that the underestimation of the patch force becomes less sensitive to the KPFM's resolution as the separation between plates increases. The results of this study could provide useful guidance for the accurate evaluation of electrostatic patch forces using KPFM.

cond-mat.mes-hall

Radar and Camera Fusion for Object Detection and Tracking: A Comprehensive Survey

Multi-modal fusion is imperative to the implementation of reliable object detection and tracking in complex environments. Exploiting the synergy of heterogeneous modal information endows perception systems the ability to achieve more comprehensive, robust, and accurate performance. As a nucleus concern in wireless-vision collaboration, radar-camera fusion has prompted prospective research directions owing to its extensive applicability, complementarity, and compatibility. Nonetheless, there still lacks a systematic survey specifically focusing on deep fusion of radar and camera for object detection and tracking. To fill this void, we embark on an endeavor to comprehensively review radar-camera fusion in a holistic way. First, we elaborate on the fundamental principles, methodologies, and applications of radar-camera fusion perception. Next, we delve into the key techniques concerning sensor calibration, modal representation, data alignment, and fusion operation. Furthermore, we provide a detailed taxonomy covering the research topics related to object detection and tracking in the context of radar and camera technologies.Finally, we discuss the emerging perspectives in the field of radar-camera fusion perception and highlight the potential areas for future research.

cs.CV

AdaptiveFusion: Adaptive Multi-Modal Multi-View Fusion for 3D Human Body Reconstruction

Recent advancements in sensor technology and deep learning have led to significant progress in 3D human body reconstruction. However, most existing approaches rely on data from a specific sensor, which can be unreliable due to the inherent limitations of individual sensing modalities. Additionally, existing multi-modal fusion methods generally require customized designs based on the specific sensor combinations or setups, which limits the flexibility and generality of these methods. Furthermore, conventional point-image projection-based and Transformer-based fusion networks are susceptible to the influence of noisy modalities and sensor poses. To address these limitations and achieve robust 3D human body reconstruction in various conditions, we propose AdaptiveFusion, a generic adaptive multi-modal multi-view fusion framework that can effectively incorporate arbitrary combinations of uncalibrated sensor inputs. By treating different modalities from various viewpoints as equal tokens, and our handcrafted modality sampling module by leveraging the inherent flexibility of Transformer models, AdaptiveFusion is able to cope with arbitrary numbers of inputs and accommodate noisy modalities with only a single training network. Extensive experiments on large-scale human datasets demonstrate the effectiveness of AdaptiveFusion in achieving high-quality 3D human body reconstruction in various environments. In addition, our method achieves superior accuracy compared to state-of-the-art fusion methods.

cs.CV

Hydrodynamics of polydisperse gas-solid flows: Kinetic theory and multifluid simulation

Polydisperse gas-solid flows, which is notoriously difficult to model due to the complex gas-particle and particle-particle interactions, are widely encountered in industry. In this article, a refined kinetic theory for polydisperse flow is developed, which features single-parameter Chapman-Enskog expansion (the Knudsen number) and exact calculation of the integrations related to pair distribution function of particle velocity without any mathematical approximations. The Navier-Stokes order constitutive relations for multifluid modeling of polydisperse gas-solid flow are then obtained analytically, including the solid stress tensor, the solid-solid drag force, the granular heat flux and the energy dissipation rate. Finally, the model is preliminarily validated by comparing to the discrete element simulation data of one-dimensional granular shear flow and by showing that the hydrodynamic characteristics of gas-solid flows in a bubbling fluidized bed containing bidisperse particles can be successfully predicted.

physics.flu-dyn

HDNet: Hierarchical Dynamic Network for Gait Recognition using Millimeter-Wave Radar

Gait recognition is widely used in diversified practical applications. Currently, the most prevalent approach is to recognize human gait from RGB images, owing to the progress of computer vision technologies. Nevertheless, the perception capability of RGB cameras deteriorates in rough circumstances, and visual surveillance may cause privacy invasion. Due to the robustness and non-invasive feature of millimeter wave (mmWave) radar, radar-based gait recognition has attracted increasing attention in recent years. In this research, we propose a Hierarchical Dynamic Network (HDNet) for gait recognition using mmWave radar. In order to explore more dynamic information, we propose point flow as a novel point clouds descriptor. We also devise a dynamic frame sampling module to promote the efficiency of computation without deteriorating performance noticeably. To prove the superiority of our methods, we perform extensive experiments on two public mmWave radar-based gait recognition datasets, and the results demonstrate that our model is superior to existing state-of-the-art methods.

cs.CV

ImmFusion: Robust mmWave-RGB Fusion for 3D Human Body Reconstruction in All Weather Conditions

3D human reconstruction from RGB images achieves decent results in good weather conditions but degrades dramatically in rough weather. Complementary, mmWave radars have been employed to reconstruct 3D human joints and meshes in rough weather. However, combining RGB and mmWave signals for robust all-weather 3D human reconstruction is still an open challenge, given the sparse nature of mmWave and the vulnerability of RGB images. In this paper, we present ImmFusion, the first mmWave-RGB fusion solution to reconstruct 3D human bodies in all weather conditions robustly. Specifically, our ImmFusion consists of image and point backbones for token feature extraction and a Transformer module for token fusion. The image and point backbones refine global and local features from original data, and the Fusion Transformer Module aims for effective information fusion of two modalities by dynamically selecting informative tokens. Extensive experiments on a large-scale dataset, mmBody, captured in various environments demonstrate that ImmFusion can efficiently utilize the information of two modalities to achieve a robust 3D human body reconstruction in all weather conditions. In addition, our method's accuracy is significantly superior to that of state-of-the-art Transformer-based LiDAR-camera fusion methods.

cs.CV

Combinatorial formulas for some generalized Ekeland-Hofer-Zehnder capacities of convex polytopes

Motivated by Pazit Haim-Kislev's combinatorial formula for the Ekeland-Hofer-Zehnder capacities of convex polytopes, we give corresponding formulas for $Ψ$-Ekeland-Hofer-Zehnder and coisotropic Ekeland-Hofer-Zehnder capacities of convex polytopes introduced by the second named author and others recently. Contrary to Pazit Haim-Kislev's subadditivity result for the Ekeland-Hofer-Zehnder capacities of convex domains, we show that the coisotropic Hofer-Zehnder capacities satisfy the subadditivity for suitable hyperplane cuts of two-dimensional convex domains in the reverse direction.

math.SG

MuVAM: A Multi-View Attention-based Model for Medical Visual Question Answering

Medical Visual Question Answering (VQA) is a multi-modal challenging task widely considered by research communities of the computer vision and natural language processing. Since most current medical VQA models focus on visual content, ignoring the importance of text, this paper proposes a multi-view attention-based model(MuVAM) for medical visual question answering which integrates the high-level semantics of medical images on the basis of text description. Firstly, different methods are utilized to extract the features of the image and the question for the two modalities of vision and text. Secondly, this paper proposes a multi-view attention mechanism that include Image-to-Question (I2Q) attention and Word-to-Text (W2T) attention. Multi-view attention can correlate the question with image and word in order to better analyze the question and get an accurate answer. Thirdly, a composite loss is presented to predict the answer accurately after multi-modal feature fusion and improve the similarity between visual and textual cross-modal features. It consists of classification loss and image-question complementary (IQC) loss. Finally, for data errors and missing labels in the VQA-RAD dataset, we collaborate with medical experts to correct and complete this dataset and then construct an enhanced dataset, VQA-RADPh. The experiments on these two datasets show that the effectiveness of MuVAM surpasses the state-of-the-art method.

cs.CV

Higher $P$-symmetric Ekeland-Hofer capacities

This paper is devoted to the construction of analogues of higher Ekeland-Hofer symplectic capacities for $P$-symmetric subsets in the standard symplectic space $(\mathbb{R}^{2n},ω_0)$, which is motivated by Long and Dong's study $P$-symmetric closed characteristics on $P$-symmetric convex bodies. We study the relationship between these capacities and other capacities, and give some computation examples. Moreover, we also define higher real symmetric Ekeland-Hofer capacities as a complement of Jin and the second named author's recent study of the real symmetric analogue about the first Ekeland-Hofer capacity.

math.SG

A note on symmetrical symplectic capacities

For a convex domain in the standard Euclidean symplectic space which is invariant under a linear anti-symplectic involution $τ$ we show that its Ekeland-Hofer-Zehnder capacity is equal to the $τ$-symmetrical symplectic capacity of it.

math.SG

The Poisson bracket invariant for open covers consisting of topological disks on surfaces

L. Buhovsky, A. Logunov and S. Tanny proved the (strong) Poisson bracket conjecture by Leonid Polterovich in dimension $2$. In this note, instead of open cover consisting of displaceable sets in their work, we consider open cover constituted of topological discs and give a necessary and sufficient condition that Poisson bracket invariants of these covers are positive.

math.SG

Quantization Noise Shaping for Information Maximizing ADCs

ADCs sit at the interface of the analog and digital worlds and fundamentally determine what information is available in the digital domain for processing. This paper shows that a configurable ADC can be designed for signals with non constant information as a function of frequency such that within a fixed power budget the ADC maximizes the information in the converted signal by frequency shaping the quantization noise. Quantization noise shaping can be realized via loop filter design for a single channel delta sigma ADC and extended to common time and frequency interleaved multi channel structures. Results are presented for example wireline and wireless style channels.

cs.IT