Searcharxiv⌕ Search

arXiv subjects

Ming Ma

Publications and source records attributed to Ming Ma.

At least 37 records · Page 2Linked to original sources

Fast Burst-Sparsity Learning Approach for Massive MIMO-OTFS Channel Estimation

Accurate channel estimation in orthogonal time frequency space (OTFS) systems with massive multiple-input multiple-output (MIMO) configurations is challenging due to high-dimensional sparse representation (SR). Existing methods often face performance degradation and/or high computational complexity. To address these issues and exploit intricate channel sparsity structure, this letter first leverages a novel hybrid burst-sparsity prior to capture the burst/common sparse structure in the angle/delay domain, and then utilizes an independent variational Bayesian inference (VBI) factorization technique to efficiently solve the high-dimensional SR problem. Additionally, an angle/Doppler refinement approach is incorporated into the proposed method to automatically mitigate off-grid mismatches.

eess.SP↗

Evolution of Interfacial Hydration Structure Induced by Ion Condensation and Correlation Effects

Interfacial hydration structures are crucial in wide-ranging applications, including battery, colloid, lubrication etc. Multivalent ions like Mg2+ and La3+ show irreplaceable roles in these applications, which are hypothesized due to their unique interfacial hydration structures. However, this hypothesis lacks experimental supports. Here, using three-dimensional atomic force microscopy (3D-AFM), we provide the first observation for their interfacial hydration structures with molecular resolution. We observed the evolution of layered hydration structures at La(NO3)3 solution-mica interfaces with concentration. As concentration increases from 25 mM to 2 M, the layer number varies from 2 to 1 and back to 2, and the interlayer thickness rises from 0.25 to 0.34 nm, with hydration force increasing from 0.27+-0.07 to 1.04+-0.24 nN. Theory and molecular simulation reveal that multivalence induces concentration-dependent ion condensation and correlation effects, resulting in compositional and structural evolution within interfacial hydration structures. Additional experiments with MgCl2-mica, La(NO3)3-graphite and Al(NO3)3-mica interfaces together with literature comparison confirm the universality of this mechanism for both multivalent and monovalent ions. New factors affecting interfacial hydration structures are revealed, including concentration and solvent dielectric constant. This insight provides guidance for designing interfacial hydration structures to optimize solid-liquid-interphase for battery life extension, modulate colloid stability and develop efficient lubricants.

physics.chem-ph↗

MoGU: A Framework for Enhancing Safety of Open-Sourced LLMs While Preserving Their Usability

Large Language Models (LLMs) are increasingly deployed in various applications. As their usage grows, concerns regarding their safety are rising, especially in maintaining harmless responses when faced with malicious instructions. Many defense strategies have been developed to enhance the safety of LLMs. However, our research finds that existing defense strategies lead LLMs to predominantly adopt a rejection-oriented stance, thereby diminishing the usability of their responses to benign instructions. To solve this problem, we introduce the MoGU framework, designed to enhance LLMs' safety while preserving their usability. Our MoGU framework transforms the base LLM into two variants: the usable LLM and the safe LLM, and further employs dynamic routing to balance their contribution. When encountering malicious instructions, the router will assign a higher weight to the safe LLM to ensure that responses are harmless. Conversely, for benign instructions, the router prioritizes the usable LLM, facilitating usable and helpful responses. On various open-sourced LLMs, we compare multiple defense strategies to verify the superiority of our MoGU framework. Besides, our analysis provides key insights into the effectiveness of MoGU and verifies that our designed routing mechanism can effectively balance the contribution of each variant by assigning weights. Our work released the safer Llama2, Vicuna, Falcon, Dolphin, and Baichuan2.

cs.CL↗

Don't Ignore Dual Logic Ability of LLMs while Privatizing: A Data-Intensive Analysis in Medical Domain

Extensive studies have been devoted to privatizing general-domain Large Language Models (LLMs) as Domain-Specific LLMs via feeding specific-domain data. However, these privatization efforts often ignored a critical aspect: Dual Logic Ability, which is a core reasoning ability for LLMs. The dual logic ability of LLMs ensures that they can maintain a consistent stance when confronted with both positive and negative statements about the same fact. Our study focuses on how the dual logic ability of LLMs is affected during the privatization process in the medical domain. We conduct several experiments to analyze the dual logic ability of LLMs by examining the consistency of the stance in responses to paired questions about the same fact. In our experiments, interestingly, we observed a significant decrease in the dual logic ability of existing LLMs after privatization. Besides, our results indicate that incorporating general domain dual logic data into LLMs not only enhances LLMs' dual logic ability but also further improves their accuracy. These findings underscore the importance of prioritizing LLMs' dual logic ability during the privatization process. Our study establishes a benchmark for future research aimed at exploring LLMs' dual logic ability during the privatization process and offers valuable guidance for privatization efforts in real-world applications.

cs.CL↗

Analyzing the Inherent Response Tendency of LLMs: Real-World Instructions-Driven Jailbreak

Extensive work has been devoted to improving the safety mechanism of Large Language Models (LLMs). However, LLMs still tend to generate harmful responses when faced with malicious instructions, a phenomenon referred to as "Jailbreak Attack". In our research, we introduce a novel automatic jailbreak method RADIAL, which bypasses the security mechanism by amplifying the potential of LLMs to generate affirmation responses. The jailbreak idea of our method is "Inherent Response Tendency Analysis" which identifies real-world instructions that can inherently induce LLMs to generate affirmation responses and the corresponding jailbreak strategy is "Real-World Instructions-Driven Jailbreak" which involves strategically splicing real-world instructions identified through the above analysis around the malicious instruction. Our method achieves excellent attack performance on English malicious instructions with five open-source advanced LLMs while maintaining robust attack performance in executing cross-language attacks against Chinese malicious instructions. We conduct experiments to verify the effectiveness of our jailbreak idea and the rationality of our jailbreak strategy design. Notably, our method designed a semantically coherent attack prompt, highlighting the potential risks of LLMs. Our study provides detailed insights into jailbreak attacks, establishing a foundation for the development of safer LLMs.

cs.CL↗

Robustness-enhanced Uplift Modeling with Adversarial Feature Desensitization

Uplift modeling has shown very promising results in online marketing. However, most existing works are prone to the robustness challenge in some practical applications. In this paper, we first present a possible explanation for the above phenomenon. We verify that there is a feature sensitivity problem in online marketing using different real-world datasets, where the perturbation of some key features will seriously affect the performance of the uplift model and even cause the opposite trend. To solve the above problem, we propose a novel robustness-enhanced uplift modeling framework with adversarial feature desensitization (RUAD). Specifically, our RUAD can more effectively alleviate the feature sensitivity of the uplift model through two customized modules, including a feature selection module with joint multi-label modeling to identify a key subset from the input features and an adversarial feature desensitization module using adversarial training and soft interpolation operations to enhance the robustness of the model against this selected subset of features. Finally, we conduct extensive experiments on a public dataset and a real product dataset to verify the effectiveness of our RUAD in online marketing. In addition, we also demonstrate the robustness of our RUAD to the feature sensitivity, as well as the compatibility with different uplift models.

cs.LG↗

Kinetic Friction of Structurally Superlubric 2D Material Interfaces

The ultra-low kinetic friction F_k of 2D structurally superlubric interfaces, connected with the fast motion of the incommensurate moiré pattern, is often invoked for its linear increase with velocity v_0 and area A, but never seriously addressed and calculated so far. Here we do that, exemplifying with a twisted graphene layer sliding on top of bulk graphite -- a demonstration case that could easily be generalized to other systems. Neglecting quantum effects and assuming a classical Langevin dynamics, we derive friction expressions valid in two temperature regimes. At low temperatures the nonzero sliding friction velocity derivative dF_k/dv_0 is shown by Adelman-Doll-Kantorovich type approximations to be equivalent to that of a bilayer whose substrate is affected by an analytically derived effective damping parameter, replacing the semi-infinite substrate. At high temperatures, friction grows proportional to temperature as analytically required by fluctuation-dissipation. The theory is validated by non-equilibrium molecular dynamics simulations with different contact areas, velocities, twist angles and temperatures. Using 6^{\circ}-twisted graphene on Bernal graphite as a prototype we find a shear stress of measurable magnitude, from 25 kPa at low temperature to 260 kPa at room temperature, yet only at high sliding velocities such as 100 m/s. However, it will linearly drop many orders of magnitude below measurable values at common experimental velocities such as 1 μm/s, a factor 10^{-8} lower. The low but not ultra-low "engineering superlubric" friction measured in existing experiments should therefore be attributed to defects and/or edges, whose contribution surpasses by far the negligible moiré contribution.

cond-mat.mtrl-sci↗

A Unified Formula of the Optimal Portfolio for Piecewise Hyperbolic Absolute Risk Aversion Utilities

We propose a general family of piecewise hyperbolic absolute risk aversion (PHARA) utilities, including many classic and non-standard utilities as examples. A typical application is the composition of a HARA preference and a piecewise linear payoff in asset allocation. We derive a unified closed-form formula of the optimal portfolio, which is a four-term division. The formula has clear economic meanings, reflecting the behavior of risk aversion, risk seeking, loss aversion and first-order risk aversion. We conduct a general asymptotic analysis to the optimal portfolio, which directly serves as an analytical tool for financial analysis. We compare this PHARA portfolio with those of other utility families both analytically and numerically. One main finding is that risk-taking behaviors are greatly increased by non-concavity and reduced by non-differentiability of the PHARA utility. Finally, we use financial data to test the performance of the PHARA portfolio in the market.

q-fin.MF↗

CWP: Instance complexity weighted channel-wise soft masks for network pruning

Existing differentiable channel pruning methods often attach scaling factors or masks behind channels to prune filters with less importance, and implicitly assume uniform contribution of input samples to filter importance. Specifically, the effects of instance complexity on pruning performance are not yet fully investigated in static network pruning. In this paper, we propose a simple yet effective differentiable network pruning method CWP based on instance complexity weighted filter importance scores. We define instance complexity related weight for each instance by giving higher weights to hard instances, and measure the weighted sum of instance-specific soft masks to model non-uniform contribution of different inputs, which encourages hard instances to dominate the pruning process and the model performance to be well preserved. In addition, we introduce a regularizer to maximize polarization of the masks, such that a sweet spot can be easily found to identify the filters to be pruned. Performance evaluations on various network architectures and datasets demonstrate CWP has advantages over the state-of-the-arts in pruning large networks. For instance, CWP improves the accuracy of ResNet56 on CIFAR-10 dataset by 0.32% aftering removing 64.11% FLOPs, and prunes 87.75% FLOPs of ResNet50 on ImageNet dataset with only 0.93% Top-1 accuracy loss.

cs.LG↗

Accurate estimation of dynamical quantities for nonequilibrium nanoscale system

Fluctuations of dynamical quantities are fundamental and inevitable. For the booming research in nanotechnology, huge relative fluctuation comes with the reduction of system size, leading to large uncertainty for the estimates of dynamical quantities. Thus, increasing statistical efficiency, i.e., reducing the number of samples required to achieve a given accuracy, is of great significance for accurate estimation. Here we propose a theory as a fundamental solution for such problem by constructing auxiliary path for each real path. The states on auxiliary paths constitute canonical ensemble and share the same macroscopic properties with the initial states of the real path. By implementing the theory in molecular dynamics simulations, we obtain a nanoscale Couette flow field with an accuracy of 0.2 μm/s with relative standard error < 0.1. The required number of samples is reduced by 12 orders compared to conventional method. The predicted thermolubric behavior of water sliding on a self-assembled surface is directly validated by experiment under the same velocity. As the theory only assumes the system is initially in thermal equilibrium then driven from that equilibrium by an external perturbation, we believe it could serve as a general approach for extracting the accurate estimate of dynamical quantities from large fluctuations to provide insights on atomic level under experimental conditions, and benefit the studies on mass transport across (biological) nanochannels and fluid film lubrication of nanometer thickness.

stat.CO↗

Translucency and negative temperature-dependence for the slip length of water on graphene

Carbonous materials, such as graphene and carbon nanotube, have attracted tremendous attention in the fields of nanofluidics due to the slip at the interface between solid and liquid. The dependence of slip length for water on the types of supporting substrates and thickness of carbonous layer, which is critical for applications such as sustainable cooling of electronic devices, remains unknown. In this paper, using colloidal probe atomic force microscope, we measured the slip length of water on graphene ls supported by hydrophilic and hydrophobic substrates, i.e., SiO2 and octadecyltrimethoxysilane (OTS). The ls on single-layer graphene supported by SiO2 is found to be 1.6~1.9 nm, and by OTS is 8.5~0.9 nm. With the thickness of few-layer graphene increases to 3~4 layers, both ls gradually converge to the value of graphite (4.3~3.5 nm). Such thickness dependence is termed slip length translucency. Further, ls is found to decrease by about 70% with the temperature increases from 300 K to 350 K for 2-layer graphene supported by SiO2. These observations are explained by analysis based on Green-Kubo relation and McLachlan theory. Our results provide the first set of reference values for the slip length of water on supported few-layer graphene. They can not only serve as a direct experimental reference for solid-liquid interaction, but also provide guideline for the design of nanofluidics-based devices, for example the thermo-mechanical nanofluidic devices.

physics.flu-dyn↗

a novel attention-based network for fast salient object detection

In the current salient object detection network, the most popular method is using U-shape structure. However, the massive number of parameters leads to more consumption of computing and storage resources which are not feasible to deploy on the limited memory device. Some others shallow layer network will not maintain the same accuracy compared with U-shape structure and the deep network structure with more parameters will not converge to a global minimum loss with great speed. To overcome all of these disadvantages, we proposed a new deep convolution network architecture with three contributions: (1) using smaller convolution neural networks (CNNs) to compress the model in our improved salient object features compression and reinforcement extraction module (ISFCREM) to reduce parameters of the model. (2) introducing channel attention mechanism in ISFCREM to weigh different channels for improving the ability of feature representation. (3) applying a new optimizer to accumulate the long-term gradient information during training to adaptively tune the learning rate. The results demonstrate that the proposed method can compress the model to 1/3 of the original size nearly without losing the accuracy and converging faster and more smoothly on six widely used datasets of salient object detection compared with the others models. Our code is published in https://gitee.com/binzhangbinzhangbin/code-a-novel-attention-based-network-for-fast-salient-object-detection.git

cs.CV↗

Multimedia Edge Computing

In this paper, we investigate the recent studies on multimedia edge computing, from sensing not only traditional visual/audio data but also individuals' geographical preference and mobility behaviors, to performing distributed machine learning over such data using the joint edge and cloud infrastructure and using evolutional strategies like reinforcement learning and online learning at edge devices to optimize the quality of experience for multimedia services at the last mile proactively. We provide both a retrospective view of recent rapid migration (resp. merge) of cloud multimedia to (resp. and) edge-aware multimedia and insights on the fundamental guidelines for designing multimedia edge computing strategies that target satisfying the changing demand of quality of experience. By showing the recent research studies and industrial solutions, we also provide future directions towards high-quality multimedia services over edge computing.

cs.MM↗

The Normal Map Based on Area-Preserving Parameterization

In this paper, we present an approach to enhance and improve the current normal map rendering technique. Our algorithm is based on semi-discrete Optimal Mass Transportation (OMT) theory and has a solid theoretical base. The key difference from previous normal map method is that we preserve the local area when we unwrap a disk-like 3D surface onto 2D plane. Compared to the currently used techniques which is based on conformal parameterization, our method does not need to cut a surface into many small pieces to avoid the large area distortion. The following charts packing step is also unnecessary in our framework. Our method is practical and makes the normal map technique more robust and efficient.

cs.GR↗

Load-velocity-temperature relationship in frictional response of microscopic contacts

Frictional properties of interfaces with dynamic chemical bonds have been the subject of intensive experimental investigation and modeling, as it provides important insights into the molecular origin of the empirical rate and state laws, which have been highly successful in describing friction from nano to geophysical scales. Using previously developed theoretical approaches requires time-consuming simulations that are impractical for many realistic tribological systems. To solve this problem and set a framework for understanding microscopic mechanisms of friction at interfaces including multiple microscopic contacts, we developed an analytical approach for description of friction mediated by dynamical formation and rupture of microscopic interfacial contacts, which allows to calculate frictional properties on the time and length scales that are relevant to tribological experimental conditions. The model accounts for the presence of various types of contacts at the frictional interface and predicts novel dependencies of friction on sliding velocity, temperature, and normal load, which are amenable to experimental observations. Our model predicts the velocity-temperature scaling, which relies on the interplay between the effects of shear and temperature on the rupture of interfacial contacts. The proposed scaling can be used to extrapolate the simulation results to a range of very low sliding velocities used in nanoscale friction experiments, which is still unreachable by simulations. For interfaces including two types of interfacial contacts with distinct properties, our model predicts novel double-peaked dependencies of friction on temperature and velocity. Our work provides a promising avenue for the interpretation of the experimental data on friction at interfaces including microscopic contacts and opens new pathways for the rational control of the frictional response.

cond-mat.mes-hall↗

Atlas Based Segmentations via Semi-Supervised Diffeomorphic Registrations

Purpose: Segmentation of organs-at-risk (OARs) is a bottleneck in current radiation oncology pipelines and is often time consuming and labor intensive. In this paper, we propose an atlas-based semi-supervised registration algorithm to generate accurate segmentations of OARs for which there are ground truth contours and rough segmentations of all other OARs in the atlas. To the best of our knowledge, this is the first study to use learning-based registration methods for the segmentation of head and neck patients and demonstrate its utility in clinical applications. Methods: Our algorithm cascades rigid and deformable deformation blocks, and takes on an atlas image (M), set of atlas-space segmentations (S_A), and a patient image (F) as inputs, while outputting patient-space segmentations of all OARs defined on the atlas. We train our model on 475 CT images taken from public archives and Stanford RadOnc Clinic (SROC), validate on 5 CT images from SROC, and test our model on 20 CT images from SROC. Results: Our method outperforms current state of the art learning-based registration algorithms and achieves an overall dice score of 0.789 on our test set. Moreover, our method yields a performance comparable to manual segmentation and supervised segmentation, while solving a much more complex registration problem. Whereas supervised segmentation methods only automate the segmentation process for a select few number of OARs, we demonstrate that our methods can achieve similar performance for OARs of interest, while also providing segmentations for every other OAR on the provided atlas. Conclusions: Our proposed algorithm has significant clinical applications and could help reduce the bottleneck for segmentation of head and neck OARs. Further, our results demonstrate that semi-supervised diffeomorphic registration can be accurately applied to both registration and segmentation problems.

cs.CV↗

Robust consumption-investment problem Under CRRA and CARA utilities with time-varying confidence sets

We consider a robust consumption-investment problem under CRRA and CARA utilities. The time-varying confidence sets are specified by $Θ$, a correspondence from $[0,T]$ to the space of Lévy triplets, and describe priori information about drift, volatility and jump. Under each possible measure, the log-price processes of stocks are semimartingales and the triplet of their differential characteristics is a measurable selector from the correspondence $Θ$ almost surely. By proposing and studying the global kernel, an optimal policy and a worst-case measure are generated from a saddle point of the global kernel, and they also constitute a saddle point of the objective function.

math.OC↗

Brain Morphometry Analysis with Surface Foliation Theory

Brain morphometry study plays a fundamental role in neuroimaging research. In this work, we propose a novel method for brain surface morphometry analysis based on surface foliation theory. Given brain cortical surfaces with automatically extracted landmark curves, we first construct finite foliations on surfaces. A set of admissible curves and a height parameter for each loop are provided by users. The admissible curves cut the surface into a set of pairs of pants. A pants decomposition graph is then constructed. Strebel differential is obtained by computing a unique harmonic map from surface to pants decomposition graph. The critical trajectories of Strebel differential decompose the surface into topological cylinders. After conformally mapping those topological cylinders to standard cylinders, parameters of standard cylinders (height, circumference) are intrinsic geometric features of the original cortical surfaces and thus can be used for morphometry analysis purpose. In this work, we propose a set of novel surface features rooted in surface foliation theory. To the best of our knowledge, this is the first work to make use of surface foliation theory for brain morphometry analysis. The features we computed are intrinsic and informative. The proposed method is rigorous, geometric, and automatic. Experimental results on classifying brain cortical surfaces between patients with Alzheimer's disease and healthy control subjects demonstrate the efficiency and efficacy of our method.

cs.CG↗