SearcharxivSearch

arXiv subjects

Ziyang Yu

Publications and source records attributed to Ziyang Yu.

27 records · Page 2Linked to original sources

Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models

The burgeoning field of Large Language Models (LLMs), exemplified by sophisticated models like OpenAI's ChatGPT, represents a significant advancement in artificial intelligence. These models, however, bring forth substantial challenges in the high consumption of computational, memory, energy, and financial resources, especially in environments with limited resource capabilities. This survey aims to systematically address these challenges by reviewing a broad spectrum of techniques designed to enhance the resource efficiency of LLMs. We categorize methods based on their optimization focus: computational, memory, energy, financial, and network resources and their applicability across various stages of an LLM's lifecycle, including architecture design, pretraining, finetuning, and system design. Additionally, the survey introduces a nuanced categorization of resource efficiency techniques by their specific resource types, which uncovers the intricate relationships and mappings between various resources and corresponding optimization techniques. A standardized set of evaluation metrics and datasets is also presented to facilitate consistent and fair comparisons across different models and techniques. By offering a comprehensive overview of the current sota and identifying open research avenues, this survey serves as a foundational reference for researchers and practitioners, aiding them in developing more sustainable and efficient LLMs in a rapidly evolving landscape.

cs.LG

Force-Guided Bridge Matching for Full-Atom Time-Coarsened Dynamics of Peptides

Molecular Dynamics (MD) is crucial in various fields such as materials science, chemistry, and pharmacology to name a few. Conventional MD software struggles with the balance between time cost and prediction accuracy, which restricts its wider application. Recently, data-driven approaches based on deep generative models have been devised for time-coarsened dynamics, which aim at learning dynamics of diverse molecular systems over a long timestep, enjoying both universality and efficiency. Nevertheless, most current methods are designed solely to learn from the data distribution regardless of the underlying Boltzmann distribution, and the physics priors such as energies and forces are constantly overlooked. In this work, we propose a conditional generative model called Force-guided Bridge Matching (FBM), which learns full-atom time-coarsened dynamics and targets the Boltzmann-constrained distribution. With the guidance of our delicately-designed intermediate force field, FBM leverages favourable physics priors into the generation process, giving rise to enhanced simulations. Experiments on two datasets consisting of peptides verify our superiority in terms of comprehensive metrics and demonstrate transferability to unseen systems.

physics.chem-ph

Rigid Protein-Protein Docking via Equivariant Elliptic-Paraboloid Interface Prediction

The study of rigid protein-protein docking plays an essential role in a variety of tasks such as drug design and protein engineering. Recently, several learning-based methods have been proposed for the task, exhibiting much faster docking speed than those computational methods. In this paper, we propose a novel learning-based method called ElliDock, which predicts an elliptic paraboloid to represent the protein-protein docking interface. To be specific, our model estimates elliptic paraboloid interfaces for the two input proteins respectively, and obtains the roto-translation transformation for docking by making two interfaces coincide. By its design, ElliDock is independently equivariant with respect to arbitrary rotations/translations of the proteins, which is an indispensable property to ensure the generalization of the docking process. Experimental evaluations show that ElliDock achieves the fastest inference time among all compared methods and is strongly competitive with current state-of-the-art learning-based models such as DiffDock-PP and Multimer particularly for antibody-antigen docking.

cs.LG

Staleness-Alleviated Distributed GNN Training via Online Dynamic-Embedding Prediction

Despite the recent success of Graph Neural Networks (GNNs), it remains challenging to train GNNs on large-scale graphs due to neighbor explosions. As a remedy, distributed computing becomes a promising solution by leveraging abundant computing resources (e.g., GPU). However, the node dependency of graph data increases the difficulty of achieving high concurrency in distributed GNN training, which suffers from the massive communication overhead. To address it, Historical value approximation is deemed a promising class of distributed training techniques. It utilizes an offline memory to cache historical information (e.g., node embedding) as an affordable approximation of the exact value and achieves high concurrency. However, such benefits come at the cost of involving dated training information, leading to staleness, imprecision, and convergence issues. To overcome these challenges, this paper proposes SAT (Staleness-Alleviated Training), a novel and scalable distributed GNN training framework that reduces the embedding staleness adaptively. The key idea of SAT is to model the GNN's embedding evolution as a temporal graph and build a model upon it to predict future embedding, which effectively alleviates the staleness of the cached historical embedding. We propose an online algorithm to train the embedding predictor and the distributed GNN alternatively and further provide a convergence analysis. Empirically, we demonstrate that SAT can effectively reduce embedding staleness and thus achieve better performance and convergence speed on multiple large-scale graph datasets.

cs.LG

Catalogue of topological electrons and phonons in all allotropes of carbon

Carbon, as one of the most common element in the earth, constructs hundreds of allotropic phases to present rich physical nature. In this work, by combining the ab inito calculations and symmetry analyses method, we systematically study a large number of allotropes of carbon (703), and discovered 315 ideal topological phononic materials and 32 topological electronic materials. The ideal topological phononic nature includes single, charge-two, three, four Weyl honons, the Dirac or Weyl nodal lines phonons, and nodal surfaces phonons. And the topological electron nature ncludes topological insulator, (Type-II) Dirac points, triple nodal points, the Dirac (Weyl) nodal lines, quadratic nodal lines and so on. For convenience, we take the uni in SG 178 and pbg in SG 230 as the examples to describe the topological features in the main. We find that it is the coexistence of single pair Weyl phonons and one-nodal surfaces phonons in the uni in SG 178, which can form the single surface arc in the (100) surface BZ and isolated double-helix surface states (IDHSSs)in the (110) surface BZ. In topological semimetal pbg in SG 230, we find that the perfect triple degenerate nodal point can be found in the near Fermi level, and it can form the clear surface states in the (001) and (110) surface BZ. Our work not only greatly expands the topological features in all allotropes of carbon, but also provide many ideal platforms to study the topological electrons and phonons.

cond-mat.mtrl-sci

DevelSet: Deep Neural Level Set for Instant Mask Optimization

With the feature size continuously shrinking in advanced technology nodes, mask optimization is increasingly crucial in the conventional design flow, accompanied by an explosive growth in prohibitive computational overhead in optical proximity correction (OPC) methods. Recently, inverse lithography technique (ILT) has drawn significant attention and is becoming prevalent in emerging OPC solutions. However, ILT methods are either time-consuming or in weak performance of mask printability and manufacturability. In this paper, we present DevelSet, a GPU and deep neural network (DNN) accelerated level set OPC framework for metal layer. We first improve the conventional level set-based ILT algorithm by introducing the curvature term to reduce mask complexity and applying GPU acceleration to overcome computational bottlenecks. To further enhance printability and fast iterative convergence, we propose a novel deep neural network delicately designed with level set intrinsic principles to facilitate the joint optimization of DNN and GPU accelerated level set optimizer. Experimental results show that DevelSet framework surpasses the state-of-the-art methods in printability and boost the runtime performance achieving instant level (around 1 second).

cs.CV

AdaOPC: A Self-Adaptive Mask Optimization Framework For Real Design Patterns

Optical proximity correction (OPC) is a widely-used resolution enhancement technique (RET) for printability optimization. Recently, rigorous numerical optimization and fast machine learning are the research focus of OPC in both academia and industry, each of which complements the other in terms of robustness or efficiency. We inspect the pattern distribution on a design layer and find that different sub-regions have different pattern complexity. Besides, we also find that many patterns repetitively appear in the design layout, and these patterns may possibly share optimized masks. We exploit these properties and propose a self-adaptive OPC framework to improve efficiency. Firstly we choose different OPC solvers adaptively for patterns of different complexity from an extensible solver pool to reach a speed/accuracy co-optimization. Apart from that, we prove the feasibility of reusing optimized masks for repeated patterns and hence, build a graph-based dynamic pattern library reusing stored masks to further speed up the OPC flow. Experimental results show that our framework achieves substantial improvement in both performance and efficiency.

cs.CV

CAD Tool Design Space Exploration via Bayesian Optimization

The design complexity is increasing as the technology node keeps scaling down. As a result, the electronic design automation (EDA) tools also become more and more complex. There are lots of parameters involved in EDA tools, which results in a huge design space. What's worse, the runtime cost of the EDA flow also goes up as the complexity increases, thus exhaustive exploration is prohibitive for modern designs. Therefore, an efficient design space exploration methodology is of great importance in advanced designs. In this paper we target at an automatic flow for reducing manual tuning efforts to achieve high quality circuits synthesis outcomes. It is based on Bayesian optimization which is a promising technique for optimizing black-box functions that are expensive to evaluate. Gaussian process regression is leveraged as the surrogate model in Bayesian optimization framework. In this work, we use 64-bit prefix adder design as a case study. We demonstrate that the Bayesian optimization is efficient and effective for performing design space exploration on EDA tool parameters, which has great potential for accelerating the design flow in advanced technology nodes.

cs.OH

Voltage-controlled skyrmion-based artificial synapse in a synthetic antiferromagnet

Spintronics exhibits significant potential in neuromorphic computing system with high speed, high integration density, and low dissipation. In this letter, we propose an ultralow-dissipation spintronic memristor composed of a synthetic antiferromagnet (SAF) and a piezoelectric substrate. Skyrmions/skyrmion bubbles can be generated in the upper layer of SAF with weak anisotropy energy (Ea). With a weak electric field on the heterostructure, the interlayer antiferromagnetic coupling can be manipulated, giving rise to a continuous transition between a large skyrmion bubble and a small skyrmion. This thus induces the variation of the resistance of a magnetic tunneling junction. The synapse based on this principle may manipulate the weight in a wide range at a cost of a very low energy consumption of 0.3 fJ. These results pave a way to ultralow power neuromorphic computing applications.

physics.comp-ph