SearcharxivSearch

arXiv subjects

Ming Zheng

Publications and source records attributed to Ming Zheng.

14 recordsLinked to original sources

Insight Miner: A Time Series Analysis Dataset for Cross-Domain Alignment with Natural Language

Time-series data is critical across many scientific and industrial domains, including environmental analysis, agriculture, transportation, and finance. However, mining insights from this data typically requires deep domain expertise, a process that is both time-consuming and labor-intensive. In this paper, we propose \textbf{Insight Miner}, a large-scale multimodal model (LMM) designed to generate high-quality, comprehensive time-series descriptions enriched with domain-specific knowledge. To facilitate this, we introduce \textbf{TS-Insights}\footnote{Available at \href{https://huggingface.co/datasets/zhykoties/time-series-language-alignment}{https://huggingface.co/datasets/zhykoties/time-series-language-alignment}.}, the first general-domain dataset for time series and language alignment. TS-Insights contains 100k time-series windows sampled from 20 forecasting datasets. We construct this dataset using a novel \textbf{agentic workflow}, where we use statistical tools to extract features from raw time series before synthesizing them into coherent trend descriptions with GPT-4. Following instruction tuning on TS-Insights, Insight Miner outperforms state-of-the-art multimodal models, such as LLaVA \citep{liu2023llava} and GPT-4, in generating time-series descriptions and insights. Our findings suggest a promising direction for leveraging LMMs in time series analysis, and serve as a foundational step toward enabling LLMs to interpret time series as a native input modality.

cs.LG

Deep non-parametric logistic model with case-control data and external summary information

The case-control sampling design serves as a pivotal strategy in mitigating the imbalanced structure observed in binary data. We consider the estimation of a non-parametric logistic model with the case-control data supplemented by external summary information. The incorporation of external summary information ensures the identifiability of the model. We propose a two-step estimation procedure. In the first step, the external information is utilized to estimate the marginal case proportion. In the second step, the estimated proportion is used to construct a weighted objective function for parameter training. A deep neural network architecture is employed for functional approximation. We further derive the non-asymptotic error bound of the proposed estimator. Following this the convergence rate is obtained and is shown to reach the optimal speed of the non-parametric regression estimation. Simulation studies are conducted to evaluate the theoretical findings of the proposed method. A real data example is analyzed for illustration.

stat.ML

Statistical inference for case-control logistic regression via integrating external summary data

Case-control sampling is a commonly used retrospective sampling design to alleviate imbalanced structure of binary data. When fitting the logistic regression model with case-control data, although the slope parameter of the model can be consistently estimated, the intercept parameter is not identifiable, and the marginal case proportion is not estimatable, either. We consider the situations in which besides the case-control data from the main study, called internal study, there also exists summary-level information from related external studies. An empirical likelihood based approach is proposed to make inference for the logistic model by incorporating the internal case-control data and external information. We show that the intercept parameter is identifiable with the help of external information, and then all the regression parameters as well as the marginal case proportion can be estimated consistently. The proposed method also accounts for the possible variability in external studies. The resultant estimators are shown to be asymptotically normally distributed. The asymptotic variance-covariance matrix can be consistently estimated by the case-control data. The optimal way to utilized external information is discussed. Simulation studies are conducted to verify the theoretical findings. A real data set is analyzed for illustration.

stat.ME

SEMRes-DDPM: Residual Network Based Diffusion Modelling Applied to Imbalanced Data

In the field of data mining and machine learning, commonly used classification models cannot effectively learn in unbalanced data. In order to balance the data distribution before model training, oversampling methods are often used to generate data for a small number of classes to solve the problem of classifying unbalanced data. Most of the classical oversampling methods are based on the SMOTE technique, which only focuses on the local information of the data, and therefore the generated data may have the problem of not being realistic enough. In the current oversampling methods based on generative networks, the methods based on GANs can capture the true distribution of data, but there is the problem of pattern collapse and training instability in training; in the oversampling methods based on denoising diffusion probability models, the neural network of the inverse diffusion process using the U-Net is not applicable to tabular data, and although the MLP can be used to replace the U-Net, the problem exists due to the simplicity of the structure and the poor effect of removing noise. problem of poor noise removal. In order to overcome the above problems, we propose a novel oversampling method SEMRes-DDPM.In the SEMRes-DDPM backward diffusion process, a new neural network structure SEMST-ResNet is used, which is suitable for tabular data and has good noise removal effect, and it can generate tabular data with higher quality. Experiments show that the SEMResNet network removes noise better than MLP; SEMRes-DDPM generates data distributions that are closer to the real data distributions than TabDDPM with CWGAN-GP; on 20 real unbalanced tabular datasets with 9 classification models, SEMRes-DDPM improves the quality of the generated tabular data in terms of three evaluation metrics (F1, G-mean, AUC) with better classification performance than other SOTA oversampling methods.

cs.LG

Deep Learning Assisted Raman Spectroscopy for Rapid Identification of 2D Materials

Two-dimensional (2D) materials have attracted extensive attention due to their unique characteristics and application potentials. Raman spectroscopy, as a rapid and non-destructive probe, exhibits distinct features and holds notable advantages in the structural characterization of 2D materials. However, traditional data analysis of Raman spectra relies on manual interpretation and feature extraction, which are both time-consuming and subjective. In this work, we employ deep learning techniques, including classificatory and generative deep learning, to assist the analysis of Raman spectra of typical 2D materials. For the limited and unevenly distributed Raman spectral data, we propose a data augmentation approach based on Denoising Diffusion Probabilistic Models (DDPM) to augment the training dataset and construct a four-layer Convolutional Neural Network (CNN) for 2D material classification. Experimental results illustrate the effectiveness of DDPM in addressing data limitations and significantly improved classification model performance. The proposed DDPM-CNN method shows high reliability, with 100%classification accuracy. Our work demonstrates the practicality of deep learning-assisted Raman spectroscopy for high-precision recognition and classification of 2D materials, offering a promising avenue for rapid and automated spectral analysis.

physics.app-ph

Neural Frailty Machine: Beyond proportional hazard assumption in neural survival regressions

We present neural frailty machine (NFM), a powerful and flexible neural modeling framework for survival regressions. The NFM framework utilizes the classical idea of multiplicative frailty in survival analysis to capture unobserved heterogeneity among individuals, at the same time being able to leverage the strong approximation power of neural architectures for handling nonlinear covariate dependence. Two concrete models are derived under the framework that extends neural proportional hazard models and nonparametric hazard regression models. Both models allow efficient training under the likelihood objective. Theoretically, for both proposed models, we establish statistical guarantees of neural function approximation with respect to nonparametric components via characterizing their rate of convergence. Empirically, we provide synthetic experiments that verify our theoretical statements. We also conduct experimental evaluations over $6$ benchmark datasets of different scales, showing that the proposed NFM models outperform state-of-the-art survival models in terms of predictive performance. Our code is publicly availabel at https://github.com/Rorschach1989/nfm

cs.LG

Electronic transport property of PbS nanowire devices

Lead sulfide is an important photosensitive material, and its photoelectric properties have received widespread attention. We completed the preparation of PbS nanowires and single PbS nanowire devices, then we conducted electrical performance tests on the PbS nanowire devices. We found that a single lead sulfide nanowire device exhibits memristive properties that depend on the applied voltage and the power density of light. From this we studied the electrical transport properties of single lead sulfide nanodevices.

physics.app-ph

Applications of Raman Spectroscopy in Clinical Medicine

Raman spectroscopy provides spectral information related to the specific molecular structures of substances and has been well established as a powerful tool for studying biological tissues and diagnosing diseases. This article reviews recent advances in Raman spectroscopy and its applications in diagnosing various critical diseases, including cancers, infections, and neurodegenerative diseases, and in predicting surgical outcomes. These advances are explored through discussion of state-of-the-art forms of Raman spectroscopy, such as surface-enhanced Raman spectroscopy, resonance Raman spectroscopy, and tip-enhanced Raman spectroscopy employed in biomedical sciences. We discuss biomedical applications, including various aspects and methods of ex vivo and in vivo medical diagnosis, sample collection, data processing, and achievements in realizing the correlation between Raman spectra and biochemical information in certain diseases. Finally, we present the limitations of the current study and provide perspectives for future research.

physics.app-ph

Variable selection in doubly truncated regression

Doubly truncated data arise in many areas such as astronomy, econometrics, and medical studies. For the regression analysis with doubly truncated response variables, the existence of double truncation may bring bias for estimation as well as affect variable selection. We propose a simultaneous estimation and variable selection procedure for the doubly truncated regression, allowing a diverging number of regression parameters. To remove the bias introduced by the double truncation, a Mann-Whitney-type loss function is used. The adaptive LASSO penalty is then added into the loss function to achieve simultaneous estimation and variable selection. An iterative algorithm is designed to optimize the resulting objective function. We establish the consistency and the asymptotic normality of the proposed estimator. The oracle property of the proposed selection procedure is also obtained. Some simulation studies are conducted to show the finite sample performance of the proposed approach. We also apply the method to analyze a real astronomical data.

stat.ME

Band Structure Dependent Electronic Localization in Macroscopic Films of Single-Chirality Single-Wall Carbon Nanotubes

Significant understanding has been achieved over the last few decades regarding chirality-dependent properties of single-wall carbon nanotubes (SWCNTs), primarily through single-tube studies. However, macroscopic manifestations of chirality dependence have been limited, especially in electronic transport, despite the fact that such distinct behaviors are needed for many applications of SWCNT-based devices. In addition, developing reliable transport theory is challenging since a description of localization phenomena in an assembly of nanoobjects requires precise knowledge of disorder on multiple spatial scales, particularly if the ensemble is heterogeneous. Here, we report an observation of pronounced chirality-dependent electronic localization in temperature and magnetic field dependent conductivity measurements on macroscopic films of single-chirality SWCNTs. The samples included large-gap semiconducting (6,5) and (10,3) films, narrow-gap semiconducting (7,4) and (8,5) films, and armchair metallic (6,6) films. Experimental data and theoretical calculations revealed Mott variable-range-hopping dominated transport in all samples, while localization lengths fall into three distinct categories depending on their band gaps. Armchair films have the largest localization length. Our detailed analyses on electronic transport properties of single-chirality SWCNT films provide significant new insight into electronic transport in ensembles of nanoobjects, offering foundations for designing and deploying macroscopic SWCNT solid-state devices.

cond-mat.mes-hall

Artificial neural networks condensation: A strategy to facilitate adaption of machine learning in medical settings by reducing computational burden

Machine Learning (ML) applications on healthcare can have a great impact on people's lives helping deliver better and timely treatment to those in need. At the same time, medical data is usually big and sparse requiring important computational resources. Although it might not be a problem for wide-adoption of ML tools in developed nations, availability of computational resource can very well be limited in third-world nations. This can prevent the less favored people from benefiting of the advancement in ML applications for healthcare. In this project we explored methods to increase computational efficiency of ML algorithms, in particular Artificial Neural Nets (NN), while not compromising the accuracy of the predicted results. We used in-hospital mortality prediction as our case analysis based on the MIMIC III publicly available dataset. We explored three methods on two different NN architectures. We reduced the size of recurrent neural net (RNN) and dense neural net (DNN) by applying pruning of "unused" neurons. Additionally, we modified the RNN structure by adding a hidden-layer to the LSTM cell allowing to use less recurrent layers for the model. Finally, we implemented quantization on DNN forcing the weights to be 8-bits instead of 32-bits. We found that all our methods increased computational efficiency without compromising accuracy and some of them even achieved higher accuracy than the pre-condensed baseline models.

cs.LG

Regression analysis of doubly truncated data

Doubly truncated data are found in astronomy, econometrics and survival analysis literature. They arise when each observation is confined to an interval, i.e., only those which fall within their respective intervals are observed along with the intervals. Unlike the more widely studied one-sided truncation that can be handled effectively by the counting process-based approach, doubly truncated data are much more difficult to handle. In their analysis of an astronomical data set, Efron and Petrosian (1999) proposed some nonparametric methods, including a generalization of Kendall's tau test, for doubly truncated data. Motivated by their approach, as well as by the work of Bhattacharya et al. (1983) for right truncated data, we proposed a general method for estimating the regression parameter when the dependent variable is subject to the double truncation. It extends the Mann-Whitney-type rank estimator and can be computed easily by existing software packages. We show that the resulting estimator is consistent and asymptotically normal. A resampling scheme is proposed with large sample justification for approximating the limiting distribution. The quasar data in Efron and Petrosian (1999) are re-analyzed by the new method. Simulation results show that the proposed method works well. Extension to weighted rank estimation are also given.

stat.ME

Two-color spectroscopy of UV excited ssDNA complex with a single-wall nanotube probe: Fast nucleobase autoionization mechanism

DNA autoionization is a fundamental process wherein UV-photoexcited nucleobases dissipate energy by charge transfer to the environment without undergoing chemical damage. Here, single-wall carbon nanotubes (SWNT) are explored as a photoluminescent reporter for studying the mechanism and rates of DNA autoionization. Two-color photoluminescence spectroscopy allows separate photoexcitation of the DNA and the SWNTs in the UV and visible range, respectively. A strong SWNT photoluminescence quenching is observed when the UV pump is resonant with the DNA absorption, consistent with charge transfer from the excited states of the DNA to the SWNT. Semiempirical calculations of the DNA-SWNT electronic structure, combined with a Green's function theory for charge transfer, show a 20 fs autoionization rate, dominated by the hole transfer. Rate-equation analysis of the spectroscopy data confirms that the quenching rate is limited by the thermalization of the free charge carriers transferred to the nanotube reservoir. The developed approach has a great potential for monitoring DNA excitation, autoionization, and chemical damage both {\it in vivo} and {\it in vitro}.

cond-mat.mes-hall

Optical Characterizations and Electronic Devices of Nearly Pure (10,5) Single-Walled Carbon Nanotubes

It remains an elusive goal to achieve high performance single-walled carbon nanotube (SWNTs) field effect transistors (FETs) comprised of only single chirality SWNTs. Many separation mechanisms have been devised and various degrees of separation demonstrated, yet it is still difficult to reach the goal of total fractionation of a given nanotube mixture into its single chirality components. Chromatography has been reported to separate small SWNTs (diameter less than 0.9nm) according to their diameter, chirality and length. The separation efficiency decreased with increasing tube diameter by using ssDNA sequence d(GT)n (n=10-45). Here we report our result on the separation of single chirality (10,5) SWNTs (diameter = 1.03nm) from HiPco tubes with ion exchange chromatography. The separation efficiency was improved by using a new DNA sequence (TTTA)3T which can recognize SWNTs with the specific chirality (10,5). The chirality of the separated tubes was examined by optical absorption, Raman, photoluminescence excitation/emission and electrical transport measurement. All spectroscopic methods gave single peak of (10,5) tubes. The purity was 99% according to the electrical measurement. The FETs comprised of separated SWNTs in parallel gave Ion/Ioff ratio up to 106 owning to the single chirality enriched (10,5) tubes. This is the first time that SWNT FETs with single chirality SWNTs were achieved. The chromatography method has the potential to separate even larger diameter semiconducting SWNTs from other starting material for further improving the performance of the SWNT FETs.

cond-mat.mtrl-sci