SearcharxivSearch

arXiv subjects

Zhong Jin

Publications and source records attributed to Zhong Jin.

12 recordsLinked to original sources

Revisiting Parameter Redundancy in Vision-Language-Action Models: Insights from VLM-to-VLA Adaptation

Vision-Language-Action (VLA) models have made significant strides in embodied intelligence by integrating the powerful representations of pre-trained Vision-Language Models (VLMs). However, the massive parameter scale of VLAs imposes a heavy computational burden, and these models exhibit extreme sensitivity to parameter pruning. Current paradigms often treat the resulting performance degradation as inevitable, relying on fine-tuning or low-rank corrections to recover efficacy. We challenge this convention by questioning whether the removed parameters are truly redundant if VLA pruning necessitates performance recovery to be effective, or if this paradigm masks the indiscriminate pruning of critical parameters. We revisit parameter redundancy through the lens of VLM-to-VLA adaptation, first quantifying the spatial distribution of parameter divergence during adaptation to reveal structured patterns across different modules. Subsequently, we introduce controlled pruning as a diagnostic probe: by comparing the direct impact of removing different parameter subsets on VLA performance without any fine-tuning, we establish a causal link between adaptation-induced divergence signals and functional contributions. Based on the discovered modular heterogeneities, we design a multi-module joint pruning scheme. Evaluations on the LIBERO benchmark demonstrate that our approach reduces the parameters of OpenVLA and $π_{0.5}$ by 12\%--30\% while maintaining approximately 90\% of the original performance without any post-pruning recovery. In contrast, existing parameter pruning criteria result in total performance collapse when evaluated under the same recovery-free constraints. Our study reveals the parameter evolution mechanism in VLA adaptation and provides a new path for deploying efficient, robust robotic policies in resource-constrained environments.

cs.RO

TopoMAS: Large Language Model Driven Topological Materials Multiagent System

Topological materials occupy a frontier in condensed-matter physics thanks to their remarkable electronic and quantum properties, yet their cross-scale design remains bottlenecked by inefficient discovery workflows. Here, we introduce TopoMAS (Topological materials Multi-Agent System), an interactive human-AI framework that seamlessly orchestrates the entire materials-discovery pipeline: from user-defined queries and multi-source data retrieval, through theoretical inference and crystal-structure generation, to first-principles validation. Crucially, TopoMAS closes the loop by autonomously integrating computational outcomes into a dynamic knowledge graph, enabling continuous knowledge refinement. In collaboration with human experts, it has already guided the identification of novel topological phases SrSbO3, confirmed by first-principles calculations. Comprehensive benchmarks demonstrate robust adaptability across base Large Language Model, with the lightweight Qwen2.5-72B model achieving 94.55% accuracy while consuming only 74.3-78.4% of tokens required by Qwen3-235B and 83.0% of DeepSeek-V3's usage--delivering responses twice as fast as Qwen3-235B. This efficiency establishes TopoMAS as an accelerator for computation-driven discovery pipelines. By harmonizing rational agent orchestration with a self-evolving knowledge graph, our framework not only delivers immediate advances in topological materials but also establishes a transferable, extensible paradigm for materials-science domain.

cond-mat.mtrl-sci

Enhancing Large Language Models with Domain-Specific Knowledge: The Case in Topological Materials

Large language models (LLMs), such as ChatGPT, have demonstrated impressive performance in the text generation task, showing the ability to understand and respond to complex instructions. However, the performance of naive LLMs in speciffc domains is limited due to the scarcity of domain-speciffc corpora and specialized training. Moreover, training a specialized large-scale model necessitates signiffcant hardware resources, which restricts researchers from leveraging such models to drive advances. Hence, it is crucial to further improve and optimize LLMs to meet speciffc domain demands and enhance their scalability. Based on the condensed matter data center, we establish a material knowledge graph (MaterialsKG) and integrate it with literature. Using large language models and prompt learning, we develop a specialized dialogue system for topological materials called TopoChat. Compared to naive LLMs, TopoChat exhibits superior performance in structural and property querying, material recommendation, and complex relational reasoning. This system enables efffcient and precise retrieval of information and facilitates knowledge interaction, thereby encouraging the advancement on the ffeld of condensed matter materials.

cs.CL

Transformed Low-Rank Parameterization Can Help Robust Generalization for Tensor Neural Networks

Achieving efficient and robust multi-channel data learning is a challenging task in data science. By exploiting low-rankness in the transformed domain, i.e., transformed low-rankness, tensor Singular Value Decomposition (t-SVD) has achieved extensive success in multi-channel data representation and has recently been extended to function representation such as Neural Networks with t-product layers (t-NNs). However, it still remains unclear how t-SVD theoretically affects the learning behavior of t-NNs. This paper is the first to answer this question by deriving the upper bounds of the generalization error of both standard and adversarially trained t-NNs. It reveals that the t-NNs compressed by exact transformed low-rank parameterization can achieve a sharper adversarial generalization bound. In practice, although t-NNs rarely have exactly transformed low-rank weights, our analysis further shows that by adversarial training with gradient flow (GF), the over-parameterized t-NNs with ReLU activations are trained with implicit regularization towards transformed low-rank parameterization under certain conditions. We also establish adversarial generalization bounds for t-NNs with approximately transformed low-rank weights. Our analysis indicates that the transformed low-rank parameterization can promisingly enhance robust generalization for t-NNs.

cs.LG

Continuous Conditional Random Field Convolution for Point Cloud Segmentation

Point cloud segmentation is the foundation of 3D environmental perception for modern intelligent systems. To solve this problem and image segmentation, conditional random fields (CRFs) are usually formulated as discrete models in label space to encourage label consistency, which is actually a kind of postprocessing. In this paper, we reconsider the CRF in feature space for point cloud segmentation because it can capture the structure of features well to improve the representation ability of features rather than simply smoothing. Therefore, we first model the point cloud features with a continuous quadratic energy model and formulate its solution process as a message-passing graph convolution, by which it can be easily integrated into a deep network. We theoretically demonstrate that the message passing in the graph convolution is equivalent to the mean-field approximation of a continuous CRF model. Furthermore, we build an encoder-decoder network based on the proposed continuous CRF graph convolution (CRFConv), in which the CRFConv embedded in the decoding layers can restore the details of high-level features that were lost in the encoding stage to enhance the location ability of the network, thereby benefiting segmentation. Analogous to the CRFConv, we show that the classical discrete CRF can also work collaboratively with the proposed network via another graph convolution to further improve the segmentation results. Experiments on various point cloud benchmarks demonstrate the effectiveness and robustness of the proposed method. Compared with the state-of-the-art methods, the proposed method can also achieve competitive segmentation performance.

cs.CV

WakaVT: A Sequential Variational Transformer for Waka Generation

Poetry generation has long been a challenge for artificial intelligence. In the scope of Japanese poetry generation, many researchers have paid attention to Haiku generation, but few have focused on Waka generation. To further explore the creative potential of natural language generation systems in Japanese poetry creation, we propose a novel Waka generation model, WakaVT, which automatically produces Waka poems given user-specified keywords. Firstly, an additive mask-based approach is presented to satisfy the form constraint. Secondly, the structures of Transformer and variational autoencoder are integrated to enhance the quality of generated content. Specifically, to obtain novelty and diversity, WakaVT employs a sequence of latent variables, which effectively captures word-level variability in Waka data. To improve linguistic quality in terms of fluency, coherence, and meaningfulness, we further propose the fused multilevel self-attention mechanism, which properly models the hierarchical linguistic structure of Waka. To the best of our knowledge, we are the first to investigate Waka generation with models based on Transformer and/or variational autoencoder. Both objective and subjective evaluation results demonstrate that our model outperforms baselines significantly.

cs.CL

A Forecasting System of Computational Time of DFT/TDDFT Calculations under the Multiverse ansatz via Machine Learning and Cheminformatics

A top-level designed forecasting system for predicting computational times of density-functional theory (DFT)/time-dependent density-functional theory (TDDFT) calculations is presented. The computational time is assumed as the intrinsic property for the molecule. Basing on this assumption, the forecasting system is established using the "reinforced concrete", which combines the cheminformatics, several machine-learning (ML) models, and the framework of many-world interpretation (MWI) in multiverse ansatz. Herein, the cheminformatics is used to recognize the topological structure of molecules, the ML/AI models are used to build the relationships between topology and computational cost, and the MWI framework is used to hold various combinations of DFT functionals and basis sets in DFT/TDDFT calculations. Calculated results of molecules from DrugBank dataset show that 1) it can give quantitative predictions of computational costs, typical mean relative errors can be less than 0.2 for DFT/TDDFT calculations with derivations of 25% using the exactly pre-trained ML models, 2) it can also be employed to various combinations of DFT functional and basis set cases without exactly pre-trained ML models, while only slightly enlarge predicting errors.

physics.comp-ph

Canonical dual method for mixed integer fourth-order polynomial minimization problems with fixed cost terms

We study a canonical duality method to solve a mixed-integer nonconvex fourth-order polynomial minimization problem with fixed cost terms. This constrained nonconvex problem can be transformed into a continuous concave maximization dual problem without duality gap. The global optimality conditions are proposed and the existence and uniqueness criteria are discussed. Application to a decoupled mixed-integer problem is illustrated and analytic solution for a global minimum is obtained under some suitable conditions. Several examples are given to show the method is effective.

math.OC

On modeling and global solutions for d.c. optimization problems by canonical duality theory

This paper presents a canonical d.c. (difference of canonical and convex functions) programming problem, which can be used to model general global optimization problems in complex systems. It shows that by using the canonical duality theory, a large class of nonconvex minimization problems can be equivalently converted to a unified concave maximization problem over a convex domain, which can be solved easily under certain conditions. Additionally, a detailed proof for triality theory is provided, which can be used to identify local extremal solutions. Applications are illustrated and open problems are presented.

math.OC

Chemical reactivity imprint lithography on graphene: Controlling the substrate influence on electron transfer reactions

The chemical functionalization of graphene enables control over electronic properties and sensor recognition sites. However, its study is confounded by an unusually strong influence of the underlying substrate. In this paper, we show a stark difference in the rate of electron transfer chemistry with aryl diazonium salts on monolayer graphene supported on a broad range of substrates. Reactions proceed rapidly when graphene is on SiO_2 and Al_2O_3 (sapphire), but negligibly on alkyl-terminated and hexagonal boron nitride (hBN) surfaces. The effect is contrary to expectations based on doping levels and can instead be described using a reactivity model accounting for substrate-induced electron-hole puddles in graphene. Raman spectroscopic mapping is used to characterize the effect of the substrates on graphene. Reactivity imprint lithography (RIL) is demonstrated as a technique for spatially patterning chemical groups on graphene by patterning the underlying substrate, and is applied to the covalent tethering of proteins on graphene.

cond-mat.mtrl-sci

Terahertz and Infrared Spectroscopy of Gated Large-Area Graphene

We have fabricated a centimeter-size single-layer graphene device, with a gate electrode, which can modulate the transmission of terahertz and infrared waves. Using time-domain terahertz spectroscopy and Fourier-transform infrared spectroscopy in a wide frequency range (10-10000 cm^{-1}), we measured the dynamic conductivity change induced by electrical gating and thermal annealing. Both methods were able to effectively tune the Fermi energy, E_F, which in turn modified the Drude-like intraband absorption in the terahertz as well as the '2E_F onset' for interband absorption in the mid-infrared. These results not only provide fundamental insight into the electromagnetic response of Dirac fermions in graphene but also demonstrate the key functionalities of large-area graphene devices that are desired for components in terahertz and infrared optoelectronics.

cond-mat.mes-hall

Resistive switching in nanogap systems on SiO2 substrates

Voltage-controlled resistive switching is demonstrated in various gap systems on SiO2 substrate. The nanosized gaps are made by different means using different materials including metal, semiconductor, and metallic nonmetal. The switching site is further reduced by using multi-walled carbon nanotubes and single-walled carbon nanotubes. The switching in all the gap systems shares the same characteristics. This independence of switching on the material compositions of the electrodes, accompanied by observable damage to the SiO2 substrate at the gap region, bespeaks the intrinsic switching from post-breakdown SiO2. It calls for caution when studying resistive switching in nanosystems on oxide substrates, since oxide breakdown extrinsic to the nanosystem can mimic resistive switching. Meanwhile, the high ON/OFF ratio (10E5), fast switching time (2 us, test limit), durable cycles demonstrated show promising memory properties. The intermediate states observed reveal the filamentary conduction nature.

cond-mat.mes-hall