Searcharxiv⌕ Search

arXiv subjects

Xin Lin

Publications and source records attributed to Xin Lin.

At least 73 records · Page 4Linked to original sources

Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs

In this technical report, we tackle the challenges of training large-scale Mixture of Experts (MoE) models, focusing on overcoming cost inefficiency and resource limitations prevalent in such systems. To address these issues, we present two differently sized MoE large language models (LLMs), namely Ling-Lite and Ling-Plus (referred to as "Bailing" in Chinese, spelled Bǎilíng in Pinyin). Ling-Lite contains 16.8 billion parameters with 2.75 billion activated parameters, while Ling-Plus boasts 290 billion parameters with 28.8 billion activated parameters. Both models exhibit comparable performance to leading industry benchmarks. This report offers actionable insights to improve the efficiency and accessibility of AI development in resource-constrained settings, promoting more scalable and sustainable technologies. Specifically, to reduce training costs for large-scale MoE models, we propose innovative methods for (1) optimization of model architecture and training processes, (2) refinement of training anomaly handling, and (3) enhancement of model evaluation efficiency. Additionally, leveraging high-quality data generated from knowledge graphs, our models demonstrate superior capabilities in tool use compared to other models. Ultimately, our experimental findings demonstrate that a 300B MoE LLM can be effectively trained on lower-performance devices while achieving comparable performance to models of a similar scale, including dense and MoE models. Compared to high-performance devices, utilizing a lower-specification hardware system during the pre-training phase demonstrates significant cost savings, reducing computing costs by approximately 20%. The models can be accessed at https://huggingface.co/inclusionAI.

cs.LG↗

SGC-Net: Stratified Granular Comparison Network for Open-Vocabulary HOI Detection

Recent open-vocabulary human-object interaction (OV-HOI) detection methods primarily rely on large language model (LLM) for generating auxiliary descriptions and leverage knowledge distilled from CLIP to detect unseen interaction categories. Despite their effectiveness, these methods face two challenges: (1) feature granularity deficiency, due to reliance on last layer visual features for text alignment, leading to the neglect of crucial object-level details from intermediate layers; (2) semantic similarity confusion, resulting from CLIP's inherent biases toward certain classes, while LLM-generated descriptions based solely on labels fail to adequately capture inter-class similarities. To address these challenges, we propose a stratified granular comparison network. First, we introduce a granularity sensing alignment module that aggregates global semantic features with local details, refining interaction representations and ensuring robust alignment between intermediate visual features and text embeddings. Second, we develop a hierarchical group comparison module that recursively compares and groups classes using LLMs, generating fine-grained and discriminative descriptions for each interaction category. Experimental results on two widely-used benchmark datasets, SWIG-HOI and HICO-DET, demonstrate that our method achieves state-of-the-art results in OV-HOI detection. Codes will be released on https://github.com/Phil0212/SGC-Net.

cs.CV↗

RobuRCDet: Enhancing Robustness of Radar-Camera Fusion in Bird's Eye View for 3D Object Detection

While recent low-cost radar-camera approaches have shown promising results in multi-modal 3D object detection, both sensors face challenges from environmental and intrinsic disturbances. Poor lighting or adverse weather conditions degrade camera performance, while radar suffers from noise and positional ambiguity. Achieving robust radar-camera 3D object detection requires consistent performance across varying conditions, a topic that has not yet been fully explored. In this work, we first conduct a systematic analysis of robustness in radar-camera detection on five kinds of noises and propose RobuRCDet, a robust object detection model in BEV. Specifically, we design a 3D Gaussian Expansion (3DGE) module to mitigate inaccuracies in radar points, including position, Radar Cross-Section (RCS), and velocity. The 3DGE uses RCS and velocity priors to generate a deformable kernel map and variance for kernel size adjustment and value distribution. Additionally, we introduce a weather-adaptive fusion module, which adaptively fuses radar and camera features based on camera signal confidence. Extensive experiments on the popular benchmark, nuScenes, show that our model achieves competitive results in regular and noisy conditions.

cs.CV↗

HawkEye: Statically and Accurately Profiling the Communication Cost of Models in Multi-party Learning

Multi-party computation (MPC) based machine learning, referred to as multi-party learning (MPL), has become an important technology for utilizing data from multiple parties with privacy preservation. In recent years, in order to apply MPL in more practical scenarios, various MPC-friendly models have been proposedto reduce the extraordinary communication overhead of MPL. Within the optimization of MPC-friendly models, a critical element to tackle the challenge is profiling the communication cost of models. However, the current solutions mainly depend on manually establishing the profiles to identify communication bottlenecks of models, often involving burdensome human efforts in a monotonous procedure. In this paper, we propose HawkEye, a static model communication cost profiling framework, which enables model designers to get the accurate communication cost of models in MPL frameworks without dynamically running the secure model training or inference processes on a specific MPL framework. Firstly, to profile the communication cost of models with complex structures, we propose a static communication cost profiling method based on a prefix structure that records the function calling chain during the static analysis. Secondly, HawkEye employs an automatic differentiation library to assist model designers in profiling the communication cost of models in PyTorch. Finally, we compare the static profiling results of HawkEye against the profiling results obtained through dynamically running secure model training and inference processes on five popular MPL frameworks, CryptFlow2, CrypTen, Delphi, Cheetah, and SecretFlow-SEMI2K. The experimental results show that HawkEye can accurately profile the model communication cost without dynamic profiling.

cs.CR↗

UMC: Unified Resilient Controller for Legged Robots with Joint Malfunctions

Adaptation to unpredictable damages is crucial for autonomous legged robots, yet existing methods based on multi-policy or meta-learning frameworks face challenges like limited generalization and complex maintenance. To address this issue, we first analyze and summarize eight types of damage scenarios, including sensor failures and joint malfunctions. Then, we propose a novel, model-free, two-stage training framework, Unified Malfunction Controller (UMC), incorporating a masking mechanism to enhance damage resilience. Specifically, the model is initially trained with normal environments to ensure robust performance under standard conditions. In the second stage, we use masks to prevent the legged robot from relying on malfunctioning limbs, enabling adaptive gait and movement adjustments upon malfunction. Experimental results demonstrate that our approach improves the task completion capability by an average of 36% for the transformer and 39% for the MLP across three locomotion tasks. The source code and trained models will be made available to the public.

cs.RO↗

Linear Enhancement of Spin-Orbit Torques and Absence of Bulk Rashba-Type Spin Splitting in Perpendicularly Magnetized [Pt/Co/W]n Superlattices

The development of magnetic heterostructures with strong spin-orbit torques (SOTs), low impedance, strong perpendicular magnetic anisotropy (PMA), and good integration compatibility at the same time is central for high-performance spintronic memory and computing applications. Here, we report the development of the symmetry-broken spin-orbit superlattice [Pt/Co/W]n that can be sputtered-deposited on commercial oxidized silicon substrates and have giant SOTs, strong uniaxial PMA of 9.2 Merg/cm3. The dampinglike and fieldlike SOTs of the [Pt/Co/W]n superlattices exhibit a linear increase with the repeat number n and reach the giant values of 225% and -33% (two orders of magnitude greater than that in clean-limit Pt) at n = 12, respectively. The dampinglike SOT is also of the opposite sign and much greater in magnitude than the fieldlike SOT, regardless of the number of n. These results clarify that the spin current that generates SOTs in the [Pt/Co/W]n superlattices arises predominantly from the spin Hall effect rather than bulk Rashba-type spin splitting, providing a unified understanding of the SOTs in the superlattices. We also demonstrate deterministic switching in thicker-than-50-nm PMA [Pt/Co/W]12 superlattices at a low current density. This work establishes the [Pt/Co/W]n superlattice as a compelling material candidate for ultra-fast, low-power, long-retention nonvolatile spintronic memory and computing technologies.

cond-mat.mtrl-sci↗

Dual-Representation Interaction Driven Image Quality Assessment with Restoration Assistance

No-Reference Image Quality Assessment for distorted images has always been a challenging problem due to image content variance and distortion diversity. Previous IQA models mostly encode explicit single-quality features of synthetic images to obtain quality-aware representations for quality score prediction. However, performance decreases when facing real-world distortion and restored images from restoration models. The reason is that they do not consider the degradation factors of the low-quality images adequately. To address this issue, we first introduce the DRI method to obtain degradation vectors and quality vectors of images, which separately model the degradation and quality information of low-quality images. After that, we add the restoration network to provide the MOS score predictor with degradation information. Then, we design the Representation-based Semantic Loss (RS Loss) to assist in enhancing effective interaction between representations. Extensive experimental results demonstrate that the proposed method performs favorably against existing state-of-the-art models on both synthetic and real-world datasets.

eess.IV↗

Restore Anything with Masks: Leveraging Mask Image Modeling for Blind All-in-One Image Restoration

All-in-one image restoration aims to handle multiple degradation types using one model. This paper proposes a simple pipeline for all-in-one blind image restoration to Restore Anything with Masks (RAM). We focus on the image content by utilizing Mask Image Modeling to extract intrinsic image information rather than distinguishing degradation types like other methods. Our pipeline consists of two stages: masked image pre-training and fine-tuning with mask attribute conductance. We design a straightforward masking pre-training approach specifically tailored for all-in-one image restoration. This approach enhances networks to prioritize the extraction of image content priors from various degradations, resulting in a more balanced performance across different restoration tasks and achieving stronger overall results. To bridge the gap of input integrity while preserving learned image priors as much as possible, we selectively fine-tuned a small portion of the layers. Specifically, the importance of each layer is ranked by the proposed Mask Attribute Conductance (MAC), and the layers with higher contributions are selected for finetuning. Extensive experiments demonstrate that our method achieves state-of-the-art performance. Our code and model will be released at \href{https://github.com/Dragonisss/RAM}{https://github.com/Dragonisss/RAM}.

cs.CV↗

Hidden Turbulence in van Gogh's \textbf{\textit{The Starry Night}}

Turbulent skies have often inspired artists, particularly in the iconic swirls of Vincent van Gogh's \textbf{\textit{The Starry Night}}. For an extended period, debate has raged over whether the flow pattern in this masterpiece adheres to Kolmogorov's theory of turbulence. In contrast to previous studies that examined only part of this painting, {\textit{all and only the}} whirls/eddies in the painting are taken into account in this work, following the Richardson-Kolmogorov's cascade picture of turbulence. Consequently, the luminance's Fourier power spectrum spontaneously exhibits a characteristic $-5/3$ Kolmogorov-like power-law. This result suggests that van Gogh had a very careful observation of real flows, so that not only the sizes of whirls/eddies in \textbf{\textit{The Starry Night}} but also their relative distances and intensity follow the physical law that governs turbulent flows. Moreover, a "$-1$"-like power-law persists in the spectrum below the scales of the smallest whirls, hinting at Batchelor-type scalar turbulence with a high Schmidt number. Our study thus unveils the hidden turbulence captured within \textbf{\textit{The Starry Night}}.

physics.flu-dyn↗

Multi-task Image Restoration Guided By Robust DINO Features

Multi-task image restoration has gained significant interest due to its inherent versatility and efficiency compared to its single-task counterpart. However, performance decline is observed with an increase in the number of tasks, primarily attributed to the restoration model's challenge in handling different tasks with distinct natures at the same time. Thus, a perspective emerged aiming to explore the degradation-insensitive semantic commonalities among different degradation tasks. In this paper, we observe that the features of DINOv2 can effectively model semantic information and are independent of degradation factors. Motivated by this observation, we propose \mbox{\textbf{DINO-IR}}, a multi-task image restoration approach leveraging robust features extracted from DINOv2 to solve multi-task image restoration simultaneously. We first propose a pixel-semantic fusion (PSF) module to dynamically fuse DINOV2's shallow features containing pixel-level information and deep features containing degradation-independent semantic information. To guide the restoration model with the features of DINOv2, we develop a DINO-Restore adaption and fusion module to adjust the channel of fused features from PSF and then integrate them with the features from the restoration model. By formulating these modules into a unified deep model, we propose a DINO perception contrastive loss to constrain the model training. Extensive experimental results demonstrate that our DINO-IR performs favorably against existing multi-task image restoration approaches in various tasks by a large margin. The source codes and trained models will be made available.

cs.CV↗

Efficient generation of out-of-plane polarized spin current in polycrystalline heavy metal devices with broken electric symmetries

Spin currents of perpendicularly polarized spins (z spins) by an in-plane charge current have received blooming interest for the potential in energy-efficient spin-orbit torque switching of perpendicular magnetization in the absence of a magnetic field. However, generation of z spins is limited mainly to magnetically or crystallographically low-symmetry single crystals (such as non-colinear antiferromagnets) that are hardly compatible with the integration to semiconductor circuits. Here, we report efficient generation of z spins in sputter-deposited polycrystalline heavy metal devices via a new mechanism of broken electric symmetries in both the transverse and perpendicular directions. Both the dampinglike and fieldlike spin-orbit torques of z spins can be tuned significantly by varying the degree of the electric asymmetries via the length, width, and thickness of devices as well as by varying the type of the heavy metals. We also show that the presence of z spins enables deterministic, nearly-full, external-magnetic-field-free switching of a uniform perpendicularly magnetized FeCoB layer, the core structure of magnetic tunnel junctions, with high coercivity at a low current density. These results establish the first universal, energy-efficient, integration-friendly approach to generate z-spin current by electric asymmetry design for dense and low-power spin-torque memory and computing technologies and will stimulate investigation of z-spin currents in various polycrystalline materials.

cond-mat.mtrl-sci↗

On algebraic degrees of inverted Kloosterman sums

The study of $n$-dimensional inverted Kloosterman sums was suggested by Katz (1995) who handled the case when $n=1$ from complex point of view. For general $n\geq 1$, the $n$-dimensional inverted Kloosterman sums were studied from both complex and $p$-adic point of view in our previous paper. In this note, we study the algebraic degree of the inverted $n$-dimensional Kloosterman sum as an algebraic integer.

math.NT↗

The Ninth NTIRE 2024 Efficient Super-Resolution Challenge Report

This paper provides a comprehensive review of the NTIRE 2024 challenge, focusing on efficient single-image super-resolution (ESR) solutions and their outcomes. The task of this challenge is to super-resolve an input image with a magnification factor of x4 based on pairs of low and corresponding high-resolution images. The primary objective is to develop networks that optimize various aspects such as runtime, parameters, and FLOPs, while still maintaining a peak signal-to-noise ratio (PSNR) of approximately 26.90 dB on the DIV2K_LSDIR_valid dataset and 26.99 dB on the DIV2K_LSDIR_test dataset. In addition, this challenge has 4 tracks including the main track (overall performance), sub-track 1 (runtime), sub-track 2 (FLOPs), and sub-track 3 (parameters). In the main track, all three metrics (ie runtime, FLOPs, and parameter count) were considered. The ranking of the main track is calculated based on a weighted sum-up of the scores of all other sub-tracks. In sub-track 1, the practical runtime performance of the submissions was evaluated, and the corresponding score was used to determine the ranking. In sub-track 2, the number of FLOPs was considered. The score calculated based on the corresponding FLOPs was used to determine the ranking. In sub-track 3, the number of parameters was considered. The score calculated based on the corresponding parameters was used to determine the ranking. RLFN is set as the baseline for efficiency measurement. The challenge had 262 registered participants, and 34 teams made valid submissions. They gauge the state-of-the-art in efficient single-image super-resolution. To facilitate the reproducibility of the challenge and enable other researchers to build upon these findings, the code and the pre-trained model of validated solutions are made publicly available at https://github.com/Amazingren/NTIRE2024_ESR/.

cs.CV↗

Impact of Charge Density Waves on Superconductivity and Topological Properties in AV$_3$Sb$_5$ Kagome Superconductors

We investigates the electronic structure and superconducting gaps in the charge density wave (CDW) states of vanadium-based Kagome superconductors AV$_3$Sb$_5$, focusing on the concurrent presence of CDW and superconducting orders. Two predominant CDW configurations are explored: the trihexagonal (TrH) and star-of-David (SoD) patterns, involving charge bond order (CBO) and chiral flux phase (CFP), corresponding to real and imaginary bond orders. In the isotropic $s$-wave superconducting state, the presence of CBO alone maintains an isotropic superconducting gap, whereas the introduction of CFP induces anisotropy in the gap, manifesting time-reversal symmetry breaking due to the CFP. Our analysis extends to the topological properties of these states, revealing a marked topological phase transition in the TrH configuration from a trivial to a non-trivial state with increasing CFP intensity. This transition suggests that the introduction of CFP could catalyze the emergence of topological superconductivity, potentially leading to the presence of Majorana excitations. The results contribute significantly to understanding the complex interplay between various CDW patterns and superconductivity in Kagome superconductors. They provide a theoretical framework for the diverse experimental observations of energy gaps and open new avenues for research into topological superconductivity and its potential applications. This study underscores the necessity for further experimental and theoretical exploration to unveil novel interwinded quantum states and functionalities in these intriguing materials.

cond-mat.supr-con↗

Dual Degradation Representation for Joint Deraining and Low-Light Enhancement in the Dark

Rain in the dark poses a significant challenge to deploying real-world applications such as autonomous driving, surveillance systems, and night photography. Existing low-light enhancement or deraining methods struggle to brighten low-light conditions and remove rain simultaneously. Additionally, cascade approaches like ``deraining followed by low-light enhancement'' or the reverse often result in problematic rain patterns or overly blurred and overexposed images. To address these challenges, we introduce an end-to-end model called L$^{2}$RIRNet, designed to manage both low-light enhancement and deraining in real-world settings. Our model features two main components: a Dual Degradation Representation Network (DDR-Net) and a Restoration Network. The DDR-Net independently learns degradation representations for luminance effects in dark areas and rain patterns in light areas, employing dual degradation loss to guide the training process. The Restoration Network restores the degraded image using a Fourier Detail Guidance (FDG) module, which leverages near-rainless detailed images, focusing on texture details in frequency and spatial domains to inform the restoration process. Furthermore, we contribute a dataset containing both synthetic and real-world low-light-rainy images. Extensive experiments demonstrate that our L$^{2}$RIRNet performs favorably against existing methods in both synthetic and complex real-world scenarios. All the code and dataset can be found in \url{https://github.com/linxin0/Low_light_rainy}.

eess.IV↗

EduNLP: Towards a Unified and Modularized Library for Educational Resources

Educational resource understanding is vital to online learning platforms, which have demonstrated growing applications recently. However, researchers and developers always struggle with using existing general natural language toolkits or domain-specific models. The issue raises a need to develop an effective and easy-to-use one that benefits AI education-related research and applications. To bridge this gap, we present a unified, modularized, and extensive library, EduNLP, focusing on educational resource understanding. In the library, we decouple the whole workflow to four key modules with consistent interfaces including data configuration, processing, model implementation, and model evaluation. We also provide a configurable pipeline to unify the data usage and model usage in standard ways, where users can customize their own needs. For the current version, we primarily provide 10 typical models from four categories, and 5 common downstream-evaluation tasks in the education domain on 8 subjects for users' usage. The project is released at: https://github.com/bigdata-ustc/EduNLP.

cs.CL↗

DOP: Diagnostic-Oriented Prompting for Large Language Models in Mathematical Correction

Math world problems correction(MWPC) is a novel task dedicated to rectifying reasoning errors in the process of solving mathematical problems. In this paper, leveraging the advancements in large language models (LLMs), we address two key objectives:(1) Distinguishing between mathematical reasoning and error correction; (2) Exploring strategies to enhance the error correction capabilities of LLMs in mathematics to solve MWPC task. We noticed that, in real-time education,assisting students in recognizing their mistakes is more crucial than simply providing correct answers. However, current research tends to prioritize obtaining accurate solutions to math problems rather than correcting potentially incorrect ones. Therefore, we modify the research paradigm, demonstrating that improving mathematical reasoning abilities does not equate to mastery in error correction. Meanwhile, we propose a novel method called diagnostic-oriented promping(DOP) aimed at facilitating LLMs to excel in error correction. In experiments, DOP has shown outstanding performance, highlighting its significant impact. We argue that in mathematical education, the demand for outstanding correctors surpasses that for proficient reasoners. Codes and data are available on https://github.com/ChenhaoEcnuCS/Reason-Correct.

cs.CL↗

Enhancing Confidence Expression in Large Language Models Through Learning from Past Experience

Large Language Models (LLMs) have exhibited remarkable performance across various downstream tasks, but they may generate inaccurate or false information with a confident tone. One of the possible solutions is to empower the LLM confidence expression capability, in which the confidence expressed can be well-aligned with the true probability of the generated answer being correct. However, leveraging the intrinsic ability of LLMs or the signals from the output logits of answers proves challenging in accurately capturing the response uncertainty in LLMs. Therefore, drawing inspiration from cognitive diagnostics, we propose a method of Learning from Past experience (LePe) to enhance the capability for confidence expression. Specifically, we first identify three key problems: (1) How to capture the inherent confidence of the LLM? (2) How to teach the LLM to express confidence? (3) How to evaluate the confidence expression of the LLM? Then we devise three stages in LePe to deal with these problems. Besides, to accurately capture the confidence of an LLM when constructing the training data, we design a complete pipeline including question preparation and answer sampling. We also conduct experiments using the Llama family of LLMs to verify the effectiveness of our proposed method on four datasets.

cs.CL↗