SearcharxivSearch

arXiv subjects

Jianpeng Ma

Publications and source records attributed to Jianpeng Ma.

At least 19 recordsLinked to original sources

Intern-S1: A Scientific Multimodal Foundation Model

In recent years, a plethora of open-source foundation models have emerged, achieving remarkable progress in some widely attended fields, with performance being quite close to that of closed-source models. However, in high-value but more challenging scientific professional fields, either the fields still rely on expert models, or the progress of general foundation models lags significantly compared to those in popular areas, far from sufficient for transforming scientific research and leaving substantial gap between open-source models and closed-source models in these scientific domains. To mitigate this gap and explore a step further toward Artificial General Intelligence (AGI), we introduce Intern-S1, a specialized generalist equipped with general understanding and reasoning capabilities with expertise to analyze multiple science modal data. Intern-S1 is a multimodal Mixture-of-Experts (MoE) model with 28 billion activated parameters and 241 billion total parameters, continually pre-trained on 5T tokens, including over 2.5T tokens from scientific domains. In the post-training stage, Intern-S1 undergoes offline and then online reinforcement learning (RL) in InternBootCamp, where we propose Mixture-of-Rewards (MoR) to synergize the RL training on more than 1000 tasks simultaneously. Through integrated innovations in algorithms, data, and training systems, Intern-S1 achieved top-tier performance in online RL training. On comprehensive evaluation benchmarks, Intern-S1 demonstrates competitive performance on general reasoning tasks among open-source models and significantly outperforms open-source models in scientific domains, surpassing closed-source state-of-the-art models in professional tasks, such as molecular synthesis planning, reaction condition prediction, predicting thermodynamic stabilities for crystals. Our models are available at https://huggingface.co/internlm/Intern-S1.

cs.LG

PP-Tac: Paper Picking Using Tactile Feedback in Dexterous Robotic Hands

Robots are increasingly envisioned as human companions, assisting with everyday tasks that often involve manipulating deformable objects. Although recent advances in robotic hardware and embodied AI have expanded their capabilities, current systems still struggle with handling thin, flat, and deformable objects such as paper and fabric. This limitation arises from the lack of suitable perception techniques for robust state estimation under diverse object appearances, as well as the absence of planning techniques for generating appropriate grasp motions. To bridge these gaps, this paper introduces PP-Tac, a robotic system for picking up paper-like objects. PP-Tac features a multi-fingered robotic hand with high-resolution omnidirectional tactile sensors \sensorname. This hardware configuration enables real-time slip detection and online frictional force control that mitigates such slips. Furthermore, grasp motion generation is achieved through a trajectory synthesis pipeline, which first constructs a dataset of finger's pinching motions. Based on this dataset, a diffusion-based policy is trained to control the hand-arm robotic system. Experiments demonstrate that PP-Tac can effectively grasp paper-like objects of varying material, thickness, and stiffness, achieving an overall success rate of 87.5\%. To our knowledge, this work is the first attempt to grasp paper-like deformable objects using a tactile dexterous hand. Our project webpage can be found at: https://peilin-666.github.io/projects/PP-Tac/

cs.RO

CAMEL2: Enhancing weakly supervised learning for histopathology images by incorporating the significance ratio

Histopathology image analysis plays a crucial role in cancer diagnosis. However, training a clinically applicable segmentation algorithm requires pathologists to engage in labour-intensive labelling. In contrast, weakly supervised learning methods, which only require coarse-grained labels at the image level, can significantly reduce the labeling efforts. Unfortunately, while these methods perform reasonably well in slide-level prediction, their ability to locate cancerous regions, which is essential for many clinical applications, remains unsatisfactory. Previously, we proposed CAMEL, which achieves comparable results to those of fully supervised baselines in pixel-level segmentation. However, CAMEL requires 1,280x1,280 image-level binary annotations for positive WSIs. Here, we present CAMEL2, by introducing a threshold of the cancerous ratio for positive bags, it allows us to better utilize the information, consequently enabling us to scale up the image-level setting from 1,280x1,280 to 5,120x5,120 while maintaining the accuracy. Our results with various datasets, demonstrate that CAMEL2, with the help of 5,120x5,120 image-level binary annotations, which are easy to annotate, achieves comparable performance to that of a fully supervised baseline in both instance- and slide-level classifications.

cs.CV

Deep Learning-based Time-varying Channel Estimation for RIS Assisted Communication

Reconfigurable intelligent surface (RIS) is considered as a revolutionary technology for future wireless communication networks. In this letter, we consider the acquisition of the time-varying cascaded channels, which is a challenging task due to the massive number of passive RIS elements and the small channel coherence time. To reduce the pilot overhead, a deep learning-based channel extrapolation is implemented over both antenna and time domains. We divide the neural network into two parts, i.e., the time-domain and the antenna-domain extrapolation networks, where the neural ordinary differential equations (ODE) are utilized. In the former, ODE accurately describes the dynamics of the RIS channels and improves the recurrent neural network's performance of time series reconstruction. In the latter, ODE is resorted to modify the relations among different data layers in a feedforward neural network. We cascade the two networks and jointly train them. Simulation results show that the proposed scheme can effectively extrapolate the cascaded RIS channels in high mobility scenario.

eess.SP

Deep Learning Based Antenna-time Domain Channel Extrapolation for Hybrid mmWave Massive MIMO

In a time-varying massive multiple-input multipleoutput (MIMO) system, the acquisition of the downlink channel state information at the base station (BS) is a very challenging task due to the prohibitively high overheads associated with downlink training and uplink feedback. In this paper, we consider the hybrid precoding structure at BS and examine the antennatime domain channel extrapolation. We design a latent ordinary differential equation (ODE)-based network under the variational auto-encoder (VAE) framework to learn the mapping function from the partial uplink channels to the full downlink ones at the BS side. Specifically, the gated recurrent unit is adopted for the encoder and the fully-connected neural network is used for the decoder. The end-to-end learning is utilized to optimize the network parameters. Simulation results show that the designed network can efficiently infer the full downlink channels from the partial uplink ones, which can significantly reduce the channel training overhead.

cs.IT

Deep Learning Based RIS Channel Extrapolation with Element-grouping

Reconfigurable intelligent surface (RIS) is considered as a revolutionary technology for future wireless communication networks. In this letter, we consider the acquisition of the cascaded channels, which is a challenging task due to the massive number of passive RIS elements. To reduce the pilot overhead, we adopt the element-grouping strategy, where each element in one group shares the same reflection coefficient and is assumed to have the same channel condition. We analyze the channel interference caused by the element-grouping strategy and further design two deep learning based networks. The first one aims to refine the partial channels by eliminating the interference, while the second one tries to extrapolate the full channels from the refined partial channels. We cascade the two networks and jointly train them. Simulation results show that the proposed scheme provides significant gain compared to the conventional element-grouping method without interference elimination.

eess.SP

Ordinary Differential Equation-based CNN for Channel Extrapolation over RIS-assisted Communication

The reconfigurable intelligent surface (RIS) is considered as a promising new technology for reconfiguring wireless communication environments. To acquire the channel information accurately and efficiently, we only turn on a fraction of all the RIS elements, formulate a sub-sampled RIS channel, and design a deep learning based scheme to extrapolate the full channel information from the partial one. Specifically, inspired by the ordinary differential equation (ODE), we set up connections between different data layers in a convolutional neural network (CNN) and improve its structure. Simulation results are provided to demonstrate that our proposed ODE-based CNN structure can achieve faster convergence speed and better solution than the cascaded CNN.

cs.IT

Deep Learning Optimized Sparse Antenna Activation for Reconfigurable Intelligent Surface Assisted Communication

To capture the communications gain of the massive radiating elements with low power cost, the conventional reconfigurable intelligent surface (RIS) usually works in passive mode. However, due to the cascaded channel structure and the lack of signal processing ability, it is difficult for RIS to obtain the individual channel state information and optimize the beamforming vector. In this paper, we add signal processing units for a few antennas at RIS to partially acquire the channels. To solve the crucial active antenna selection problem, we construct an active antenna selection network that utilizes the probabilistic sampling theory to select the optimal locations of these active antennas. With this active antenna selection network, we further design two deep learning (DL) based schemes, i.e., the channel extrapolation scheme and the beam searching scheme, to enable the RIS communication system. The former utilizes the selection network and a convolutional neural network to extrapolate the full channels from the partial channels received by the active RIS antennas, while the latter adopts a fully-connected neural network to achieve the direct mapping between the partial channels and the optimal beamforming vector with maximal transmission rate. Simulation results are provided to demonstrate the effectiveness of the designed DL-based schemes.

eess.SP

Deep Learning Based Antenna Selection for Channel Extrapolation in FDD Massive MIMO

In massive multiple-input multiple-output (MIMO) systems, the large number of antennas would bring a great challenge for the acquisition of the accurate channel state information, especially in the frequency division duplex mode. To overcome the bottleneck of the limited number of radio links in hybrid beamforming, we utilize the neural networks (NNs) to capture the inherent connection between the uplink and downlink channel data sets and extrapolate the downlink channels from a subset of the uplink channel state information. We study the antenna subset selection problem in order to achieve the best channel extrapolation and decrease the data size of NNs. The probabilistic sampling theory is utilized to approximate the discrete antenna selection as a continuous and differentiable function, which makes the back propagation of the deep learning feasible. Then, we design the proper off-line training strategy to optimize both the antenna selection pattern and the extrapolation NNs. Finally, numerical results are presented to verify the effectiveness of our proposed massive MIMO channel extrapolation algorithm.

eess.SP

Graph Neural Network based Channel Tracking for Massive MIMO Networks

In this paper, we resort to the graph neural network (GNN) and propose the new channel tracking method for the massive multiple-input multiple-output networks under the high mobility scenario. We first utilize a small number of pilots to achieve the initial channel estimation. Then, we represent the obtained channel data in the form of graphs and describe the channel spatial correlation by the weights along the edges of the graph. Furthermore, we introduce the computation steps of the main unit for the GNN and design a GNN-based channel tracking framework, which includes an encoder, a core network and a decoder. Simulation results corroborate that our proposed GNN-based scheme can achieve better performance than the works with feedforward neural network.

cs.IT

Uplink-aided High Mobility Downlink Channel Estimation over Massive MIMO-OTFS System

Although it is often used in the orthogonal frequency division multiplexing (OFDM) systems, application of massive multiple-input multiple-output (MIMO) over the orthogonal time frequency space (OTFS) modulation could suffer from enormous training overhead in high mobility scenarios. In this paper, we propose one uplink-aided high mobility downlink channel estimation scheme for the massive MIMO-OTFS networks. Specifically, we firstly formulate the time domain massive MIMO-OTFS signal model along the uplink and adopt the expectation maximization based variational Bayesian (EM-VB) framework to recover the uplink channel parameters including the angle, the delay, the Doppler frequency, and the channel gain for each physical scattering path. Correspondingly, with the help of the fast Bayesian inference, one low complex approach is constructed to overcome the bottleneck of the EM-VB. Then, we fully exploit the angle, delay and Doppler reciprocity between the uplink and the downlink and reconstruct the angles, the delays, and the Doppler frequencies for the downlink massive channels at the base station. Furthermore, we examine the downlink massive MIMO channel estimation over the delay-Doppler-angle domain. The channel dispersion of the OTFS over the delay-Doppler domain is carefully analyzed. Various numerical examples are presented to confirm the validity and robustness of the proposed scheme.

eess.SP

CAMEL: A Weakly Supervised Learning Framework for Histopathology Image Segmentation

Histopathology image analysis plays a critical role in cancer diagnosis and treatment. To automatically segment the cancerous regions, fully supervised segmentation algorithms require labor-intensive and time-consuming labeling at the pixel level. In this research, we propose CAMEL, a weakly supervised learning framework for histopathology image segmentation using only image-level labels. Using multiple instance learning (MIL)-based label enrichment, CAMEL splits the image into latticed instances and automatically generates instance-level labels. After label enrichment, the instance-level labels are further assigned to the corresponding pixels, producing the approximate pixel-level labels and making fully supervised training of segmentation models possible. CAMEL achieves comparable performance with the fully supervised approaches in both instance-level classification and pixel-level segmentation on CAMELYON16 and a colorectal adenoma dataset. Moreover, the generality of the automatic labeling methodology may benefit future weakly supervised learning studies for histopathology image analysis.

eess.IV

Time-Varying Downlink Channel Tracking for Quantized Massive MIMO Networks

This paper proposes a Bayesian downlink channel estimation algorithm for time-varying massive MIMO networks. In particular, the quantization effects at the receiver are considered. In order to fully exploit the sparsity and time correlations of channels, we formulate the time-varying massive MIMO channel as the simultaneously sparse signal model. Then, we propose a sparse Bayesian learning (SBL) framework to learn the model parameters of the sparse virtual channel. To reduce complexity, we employ the expectation maximization (EM) algorithm to achieve the approximated solution. Specifically, the factor graph and the general approximate message passing (GAMP) algorithms are used to compute the desired posterior statistics in the expectation step, so that high-dimensional integrals over the marginal distributions can be avoided. The non-zero supporting vector of a virtual channel is then obtained from channel statistics by a k-means clustering algorithm. After that, the reduced dimensional GAMP based scheme is applied to make the full use of the channel temporal correlation so as to enhance the virtual channel tracking accuracy. Finally, we demonstrate the efficacy of the proposed schemes through simulations.

cs.IT

Interference-Alignment and Soft-Space-Reuse Based Cooperative Transmission for Multi-cell Massive MIMO Networks

As a revolutionary wireless transmission strategy, interference alignment (IA) can improve the capacity of the cell-edge users. However, the acquisition of the global channel state information (CSI) for IA leads to unacceptable overhead in the massive MIMO systems. To tackle this problem, in this paper, we propose an IA and soft-space-reuse (IA-SSR) based cooperative transmission scheme under the two-stage precoding framework. Specifically, the cell-center and the cell-edge users are separately treated to fully exploit the spatial degrees of freedoms (DoF). Then, the optimal power allocation policy is developed to maximize the sum-capacity of the network. Next, a low-cost channel estimator is designed for the proposed IA-SSR framework. Some practical issues in IA-SSR implementation are also discussed. Finally, plenty of numerical results are presented to show the efficiency of the proposed algorithm.

cs.IT

Base Station Selection for Massive MIMO Networks with Two-stage Precoding

The two-stage precoding has been proposed to reduce the overhead of both the channel training and the channel state information (CSI) feedback for the massive multiple-input multiple-output (MIMO) system. But the overlap of the angle-spreading-ranges (ASR) for different user clusters may seriously degrade the performance of the two-stage precoding. In this letter, we propose one ASR overlap mitigating scheme through the base station (BS) selection. Firstly, the BS selection is formulated as a sum signal-to-interference-plus-noise ratio (SINR) maximization problem. Then, the problem is solved by a low-complex algorithm through maximizing signal-to-leakage-plus-noise ratio (SLNR). In addition, we propose one low-overhead algorithm with the lower bound on the average SLNR as the objective function. Finally, we demonstrate the efficacy of the proposed schemes through the numerical simulations.

cs.IT

OPUS-Beta: A Statistical Potential for Beta-Sheet Contact Pattern in Proteins

Developing an accurate scoring function is essential for successfully predicting protein structures. In this study, we developed a statistical potential function, called OPUS-Beta, for energetically evaluating beta-sheet contact pattern (the entire residue-residue beta-contacts of a protein) independent of the atomic coordinate information. The OPUS-Beta potential contains five terms, i.e., a self-packing term, a pairwise inter-strand packing term, a pairwise intra-strand packing term, a lattice term and a hydrogen-bonding term. The results show that, in recognizing the native beta-contact pattern from decoys, OPUS-Beta potential outperforms the existing methods in literature, especially in combination with a method using 2D-recursive neural networks (about 5% and 23% improvements in top-1 and top-5 selections). We expect OPUS-Beta potential to be useful in beta-sheet modeling for proteins.

q-bio.BM

Estimating statistical distributions using an integral identity

We present an identity for an unbiased estimate of a general statistical distribution. The identity computes the distribution density from dividing a histogram sum over a local window by a correction factor from a mean-force integral, and the mean force can be evaluated as a configuration average. We show that the optimal window size is roughly the inverse of the local mean-force fluctuation. The new identity offers a more robust and precise estimate than a previous one by Adib and Jarzynski [J. Chem. Phys. 122, 014114, (2005)]. It also allows a straightforward generalization to an arbitrary ensemble and a joint distribution of multiple variables. Particularly we derive a mean-force enhanced version of the weighted histogram analysis method (WHAM). The method can be used to improve distributions computed from molecular simulations. We illustrate the use in computing a potential energy distribution, a volume distribution in a constant-pressure ensemble, a radial distribution function and a joint distribution of amino acid backbone dihedral angles.

physics.comp-ph

Enhanced sampling and applications in protein folding in explicit solvent

We report a single-copy tempering method for simulating large complex systems. In a generalized ensemble, the method uses runtime estimate of the thermal average energy computed from a novel integral identity to guide a continuous temperature-space random walk. We first validated the method in a two-dimensional Ising model and a Lennard-Jones liquid system. It was then applied to folding of three small proteins, trpzip2, trp-cage, and villin headpiece in explicit solvent. Within 0.5~1 microsecond, all three systems were folded into atomic accuracy: the alpha carbon root mean square deviations of the best folded conformations from the native states were 0.2 A, 0.4 A, and 0.4 A, for trpzip2, trp-cage, and villin headpiece, respectively.

physics.bio-ph