SearcharxivSearch

arXiv subjects

Jiaming Li

Publications and source records attributed to Jiaming Li.

At least 55 records · Page 3Linked to original sources

First-principles Investigation of Exceptional Coarsening-resistant V-Sc(Al2Cu)4 Nanoprecipitates in Al-Cu-Mg-Ag-Sc Alloys

Aluminum-copper-magnesium-sliver (Al-Cu-Mg-Ag) alloys are extensively utilized in aerospace industries due to the formation of Omega nano-plates.However, the rapid coarsening of these nano-plates above 475 K restricts their application at elevated temperatures.When introducing scandium (Sc) to these alloys, the service temperature of the resultant alloys can reach an unprecedented 675 K, attributed to the in situ formation of a coarsening-resistant V-Sc(Al2Cu)4 phase within the Omega nano-plates. However, the fundamental thermodynamic properties and mechanisms behind the remarkable coarsening resistance of V nano-plates remain unexplored.Here, we employ first-principles calculations to investigate the phase stability of V-Sc(Al2Cu)4 phase, the basic kinetic features of V phase formation within Omega nano-plates, and the origins of the extremely high thermal stability of V nano-plates. Our results indicate that V-Sc(Al2Cu)4 is meta-stable and thermodynamically tends to evolve into a stable ScAl7Cu5 phase. We also demonstrate that kinetic factors are mainly responsible for the temperature dependence of V phase formation. Notably, the formation of V-Sc(Al2Cu)4 within Omega nano-plates modifies the Kagome lattice in the shell layer of the Omega nano-plates, inhibiting further thickening of V nano-plates through the thickening pathway of Omega nano-plates. This interface transition leads to the exceptional coarsening resistance of the V nano-plates. Moreover, we also screened 14 promising element substitutions for Sc. These findings are anticipated to accelerate the development of high-performance Al alloys with superior heat resistance.

cond-mat.mtrl-sci

PersonaMath: Boosting Mathematical Reasoning via Persona-Driven Data Augmentation

While closed-source Large Language Models (LLMs) demonstrate strong mathematical problem-solving abilities, open-source models still face challenges with such tasks. To bridge this gap, we propose a data augmentation approach and introduce PersonaMathQA, a dataset derived from MATH and GSM8K, on which we train the PersonaMath models. Our approach consists of two stages: the first stage focuses on learning from Persona Diversification, and the second stage emphasizes learning from Reflection. In the first stage, we regenerate detailed chain-of-thought (CoT) solutions as instructions using a closed-source LLM and introduce a persona-driven data augmentation technique. This technique innovatively classifies personas based on occupations, significantly enhancing the dataset's diversity and quality. In the second stage, we incorporate reflection to fully leverage more challenging and valuable questions. Evaluation of our PersonaMath models on MATH and GSM8K reveals that the PersonaMath-7B model (based on Qwen2.5-7B) achieves an accuracy of 61.2% on MATH and 87.8% on GSM8K, surpassing all baseline methods and achieving state-of-the-art performance. Notably, our dataset contains only 128.9K data points-merely 32.6% of MetaMathQA and 49.5% of MathInstruct-yet our model outperforms these baselines, demonstrating the high quality and diversity of our dataset, which enables more efficient model training. We open-source the PersonaMathQA dataset, PersonaMath models, and our code for public usage.

cs.CL

II-Bench: An Image Implication Understanding Benchmark for Multimodal Large Language Models

The rapid advancements in the development of multimodal large language models (MLLMs) have consistently led to new breakthroughs on various benchmarks. In response, numerous challenging and comprehensive benchmarks have been proposed to more accurately assess the capabilities of MLLMs. However, there is a dearth of exploration of the higher-order perceptual capabilities of MLLMs. To fill this gap, we propose the Image Implication understanding Benchmark, II-Bench, which aims to evaluate the model's higher-order perception of images. Through extensive experiments on II-Bench across multiple MLLMs, we have made significant findings. Initially, a substantial gap is observed between the performance of MLLMs and humans on II-Bench. The pinnacle accuracy of MLLMs attains 74.8%, whereas human accuracy averages 90%, peaking at an impressive 98%. Subsequently, MLLMs perform worse on abstract and complex images, suggesting limitations in their ability to understand high-level semantics and capture image details. Finally, it is observed that most models exhibit enhanced accuracy when image sentiment polarity hints are incorporated into the prompts. This observation underscores a notable deficiency in their inherent understanding of image sentiment. We believe that II-Bench will inspire the community to develop the next generation of MLLMs, advancing the journey towards expert artificial general intelligence (AGI). II-Bench is publicly available at https://huggingface.co/datasets/m-a-p/II-Bench.

cs.CL

Decoupled Pseudo-labeling for Semi-Supervised Monocular 3D Object Detection

We delve into pseudo-labeling for semi-supervised monocular 3D object detection (SSM3OD) and discover two primary issues: a misalignment between the prediction quality of 3D and 2D attributes and the tendency of depth supervision derived from pseudo-labels to be noisy, leading to significant optimization conflicts with other reliable forms of supervision. We introduce a novel decoupled pseudo-labeling (DPL) approach for SSM3OD. Our approach features a Decoupled Pseudo-label Generation (DPG) module, designed to efficiently generate pseudo-labels by separately processing 2D and 3D attributes. This module incorporates a unique homography-based method for identifying dependable pseudo-labels in BEV space, specifically for 3D attributes. Additionally, we present a DepthGradient Projection (DGP) module to mitigate optimization conflicts caused by noisy depth supervision of pseudo-labels, effectively decoupling the depth gradient and removing conflicting gradients. This dual decoupling strategy-at both the pseudo-label generation and gradient levels-significantly improves the utilization of pseudo-labels in SSM3OD. Our comprehensive experiments on the KITTI benchmark demonstrate the superiority of our method over existing approaches.

cs.CV

Observation of quantum-classical transition behavior of LGI in a dissipative quantum gas

The Leggett-Garg inequality (LGI) is a powerful tool for distinguishing between quantum and classical properties in studies of macroscopic systems. Applying the LGI to non-Hermitian systems with dissipation presents a fascinating opportunity, as competing mechanisms can either strengthen or weaken LGI violations. On one hand, dissipation-induced nonlinear interactions amplify LGI violations compared to Hermitian systems; on the other hand, dissipation leads to decoherence, which could weaken the LGI violation. In this paper, we investigate a non-Hermitian system of ultracold Fermi gas with dissipation. Our experiments reveal that as dissipation increases, the upper bound of the third-order LGI parameter $K_3$ initially rises, reaching its maximum at the exceptional point (EP), where $K_3 = C_{21} + C_{32} - C_{31}$, encompassing three two-time correlation functions. Beyond a certain dissipation threshold, the LGI violation weakens, approaching the classical limit, indicating a quantum-to-classical transition (QCT). Furthermore, we observe that the LGI violation decreases with increasing evolution time, reinforcing the QCT in the time domain. This study provides a crucial stepping stone for using the LGI to explore the QCT in many-body open quantum systems.

cond-mat.quant-gas

Benchmarking Large Language Models for Image Classification of Marine Mammals

As Artificial Intelligence (AI) has developed rapidly over the past few decades, the new generation of AI, Large Language Models (LLMs) trained on massive datasets, has achieved ground-breaking performance in many applications. Further progress has been made in multimodal LLMs, with many datasets created to evaluate LLMs with vision abilities. However, none of those datasets focuses solely on marine mammals, which are indispensable for ecological equilibrium. In this work, we build a benchmark dataset with 1,423 images of 65 kinds of marine mammals, where each animal is uniquely classified into different levels of class, ranging from species-level to medium-level to group-level. Moreover, we evaluate several approaches for classifying these marine mammals: (1) machine learning (ML) algorithms using embeddings provided by neural networks, (2) influential pre-trained neural networks, (3) zero-shot models: CLIP and LLMs, and (4) a novel LLM-based multi-agent system (MAS). The results demonstrate the strengths of traditional models and LLMs in different aspects, and the MAS can further improve the classification performance. The dataset is available on GitHub: https://github.com/yeyimilk/LLM-Vision-Marine-Animals.git.

cs.CV

Instruction-Guided Visual Masking

Instruction following is crucial in contemporary LLM. However, when extended to multimodal setting, it often suffers from misalignment between specific textual instruction and targeted local region of an image. To achieve more accurate and nuanced multimodal instruction following, we introduce Instruction-guided Visual Masking (IVM), a new versatile visual grounding model that is compatible with diverse multimodal models, such as LMM and robot model. By constructing visual masks for instruction-irrelevant regions, IVM-enhanced multimodal models can effectively focus on task-relevant image regions to better align with complex instructions. Specifically, we design a visual masking data generation pipeline and create an IVM-Mix-1M dataset with 1 million image-instruction pairs. We further introduce a new learning technique, Discriminator Weighted Supervised Learning (DWSL) for preferential IVM training that prioritizes high-quality data samples. Experimental results on generic multimodal tasks such as VQA and embodied robotic control demonstrate the versatility of IVM, which as a plug-and-play tool, significantly boosts the performance of diverse multimodal models, yielding new state-of-the-art results across challenging multimodal benchmarks. Code, model and data are available at https://github.com/2toinf/IVM.

cs.CV

Ruler: A Model-Agnostic Method to Control Generated Length for Large Language Models

The instruction-following ability of large language models enables humans to interact with AI agents in a natural way. However, when required to generate responses of a specific length, large language models often struggle to meet users' needs due to their inherent difficulty in accurately perceiving numerical constraints. To explore the ability of large language models to control the length of generated responses, we propose the Target Length Generation Task (TLG) and design two metrics, Precise Match (PM) and Flexible Match (FM) to evaluate the model's performance in adhering to specified response lengths. Furthermore, we introduce a novel, model-agnostic approach called Ruler, which employs Meta Length Tokens (MLTs) to enhance the instruction-following ability of large language models under length-constrained instructions. Specifically, Ruler equips LLMs with the ability to generate responses of a specified length based on length constraints within the instructions. Moreover, Ruler can automatically generate appropriate MLT when length constraints are not explicitly provided, demonstrating excellent versatility and generalization. Comprehensive experiments show the effectiveness of Ruler across different LLMs on Target Length Generation Task, e.g., at All Level 27.97 average gain on PM, 29.57 average gain on FM. In addition, we conduct extensive ablation experiments to further substantiate the efficacy and generalization of Ruler. Our code and data is available at https://github.com/Geaming2002/Ruler.

cs.CL

Small gaps of GSE

In this paper, we study the smallest gaps for the Gaussian symplectic ensemble (GSE). We prove that the rescaled smallest gaps and their locations converge to a Poisson point process with an explicit rate. The approach provides an alternative proof for the GOE case and complements the results in \cite{FTW}. By combining the main results from \cite{BB, FTW, FW2}, the study of the smallest gaps for the classical random matrix ensembles C$β$E and G$β$E for $β= 1, 2,$ and $4$ is now complete.

math.PR

Hierarchical Context Pruning: Optimizing Real-World Code Completion with Repository-Level Pretrained Code LLMs

Some recently developed code large language models (Code LLMs) have been pre-trained on repository-level code data (Repo-Code LLMs), enabling these models to recognize repository structures and utilize cross-file information for code completion. However, in real-world development scenarios, simply concatenating the entire code repository often exceeds the context window limits of these Repo-Code LLMs, leading to significant performance degradation. In this study, we conducted extensive preliminary experiments and analyses on six Repo-Code LLMs. The results indicate that maintaining the topological dependencies of files and increasing the code file content in the completion prompts can improve completion accuracy; pruning the specific implementations of functions in all dependent files does not significantly reduce the accuracy of completions. Based on these findings, we proposed a strategy named Hierarchical Context Pruning (HCP) to construct completion prompts with high informational code content. The HCP models the code repository at the function level, maintaining the topological dependencies between code files while removing a large amount of irrelevant code content, significantly reduces the input length for repository-level code completion. We applied the HCP strategy in experiments with six Repo-Code LLMs, and the results demonstrate that our proposed method can significantly enhance completion accuracy while substantially reducing the length of input. Our code and data are available at https://github.com/Hambaobao/HCP-Coder.

cs.CL

Observation of a broad state-to-state spin-exchange collision near a p-wave Feshbach resonances of $^6$Li atoms

The study of state-to-state spin-exchange collisions in the vicinity of $p$-wave Feshbach resonances offer great opportunities to explore many-body interactions and novel quantum phases. Here, we report the observation of a spin-exchange collision near a $p$-wave Feshbach resonance within a mixture of the lowest and third-lowest hyperfine states of $^6$Li atoms. The spin-exchange interaction is observed over a range of ten gausses and produces a pair of atoms in the second-lowest hyperfine states that are captured by a deep optical dipole trap. We apply a coupled-channel method to calculate the scattering properties of this system. We find that the $p$-wave resonance exhibits a low inelastic collision rate and a broad resonance profile, which is due to the modification by the accompanying spin-exchange collisions. These findings open up new possibilities for the creation of long-lived, strongly interacting $p$-wave Fermi gases.

cond-mat.quant-gas

Learning Background Prompts to Discover Implicit Knowledge for Open Vocabulary Object Detection

Open vocabulary object detection (OVD) aims at seeking an optimal object detector capable of recognizing objects from both base and novel categories. Recent advances leverage knowledge distillation to transfer insightful knowledge from pre-trained large-scale vision-language models to the task of object detection, significantly generalizing the powerful capabilities of the detector to identify more unknown object categories. However, these methods face significant challenges in background interpretation and model overfitting and thus often result in the loss of crucial background knowledge, giving rise to sub-optimal inference performance of the detector. To mitigate these issues, we present a novel OVD framework termed LBP to propose learning background prompts to harness explored implicit background knowledge, thus enhancing the detection performance w.r.t. base and novel categories. Specifically, we devise three modules: Background Category-specific Prompt, Background Object Discovery, and Inference Probability Rectification, to empower the detector to discover, represent, and leverage implicit object knowledge explored from background proposals. Evaluation on two benchmark datasets, OV-COCO and OV-LVIS, demonstrates the superiority of our proposed method over existing state-of-the-art approaches in handling the OVD tasks.

cs.CV

DiffAM: Diffusion-based Adversarial Makeup Transfer for Facial Privacy Protection

With the rapid development of face recognition (FR) systems, the privacy of face images on social media is facing severe challenges due to the abuse of unauthorized FR systems. Some studies utilize adversarial attack techniques to defend against malicious FR systems by generating adversarial examples. However, the generated adversarial examples, i.e., the protected face images, tend to suffer from subpar visual quality and low transferability. In this paper, we propose a novel face protection approach, dubbed DiffAM, which leverages the powerful generative ability of diffusion models to generate high-quality protected face images with adversarial makeup transferred from reference images. To be specific, we first introduce a makeup removal module to generate non-makeup images utilizing a fine-tuned diffusion model with guidance of textual prompts in CLIP space. As the inverse process of makeup transfer, makeup removal can make it easier to establish the deterministic relationship between makeup domain and non-makeup domain regardless of elaborate text prompts. Then, with this relationship, a CLIP-based makeup loss along with an ensemble attack strategy is introduced to jointly guide the direction of adversarial makeup domain, achieving the generation of protected face images with natural-looking makeup and high black-box transferability. Extensive experiments demonstrate that DiffAM achieves higher visual quality and attack success rates with a gain of 12.98% under black-box setting compared with the state of the arts. The code will be available at https://github.com/HansSunY/DiffAM.

cs.CV

Gradient-based Sampling for Class Imbalanced Semi-supervised Object Detection

Current semi-supervised object detection (SSOD) algorithms typically assume class balanced datasets (PASCAL VOC etc.) or slightly class imbalanced datasets (MS-COCO, etc). This assumption can be easily violated since real world datasets can be extremely class imbalanced in nature, thus making the performance of semi-supervised object detectors far from satisfactory. Besides, the research for this problem in SSOD is severely under-explored. To bridge this research gap, we comprehensively study the class imbalance problem for SSOD under more challenging scenarios, thus forming the first experimental setting for class imbalanced SSOD (CI-SSOD). Moreover, we propose a simple yet effective gradient-based sampling framework that tackles the class imbalance problem from the perspective of two types of confirmation biases. To tackle confirmation bias towards majority classes, the gradient-based reweighting and gradient-based thresholding modules leverage the gradients from each class to fully balance the influence of the majority and minority classes. To tackle the confirmation bias from incorrect pseudo labels of minority classes, the class-rebalancing sampling module resamples unlabeled data following the guidance of the gradient-based reweighting module. Experiments on three proposed sub-tasks, namely MS-COCO, MS-COCO to Object365 and LVIS, suggest that our method outperforms current class imbalanced object detectors by clear margins, serving as a baseline for future research in CI-SSOD. Code will be available at https://github.com/nightkeepers/CI-SSOD.

cs.CV

Single-pixel imaging based on deep learning

Single-pixel imaging can collect images at the wavelengths outside the reach of conventional focal plane array detectors. However, the limited image quality and lengthy computational times for iterative reconstruction still impede the practical application of single-pixel imaging. Recently, deep learning has been introduced into single-pixel imaging, which has attracted a lot of attention due to its exceptional reconstruction quality, fast reconstruction speed, and the potential to complete advanced sensing tasks without reconstructing images. Here, this advance is discussed and some opinions are offered. Firstly, based on the fundamental principles of single-pixel imaging and deep learning, the principles and algorithms of single-pixel imaging based on deep learning are described and analyzed. Subsequently, the implementation technologies of single-pixel imaging based on deep learning are reviewed. They are divided into super-resolution single-pixel imaging, single-pixel imaging through scattering media, photon-level single-pixel imaging, optical encryption based on single-pixel imaging, color single-pixel imaging, and image-free sensing according to diverse application fields. Finally, major challenges and corresponding feasible approaches are discussed, as well as more possible applications in the future.

eess.IV

Scaling law for three-body collisions near a narrow s-wave Feshbach resonance

Ultracold atomic gases provide a controllable system to study the inelastic processes for three-body systems, where the three-body recombination rate depends on the scattering length scaling. Such scalings have been confirmed in bosonic systems with various interaction strengths, but their existence with fermionic atoms remains elusive. In this work, we report on an experimental investigation of the scaling law for the three-body atomic loss rate $L_3$ in a two-component $^6$Li Fermi gas with the scattering length $a<0$. The scaling law is validated within a certain range of $a$ near the narrow $s$-wave Feshbach resonance, where $L_3\propto T|a|^{2.60(5)}$, and $T$ is the gas temperature. The scaling law is observed to have an upper and a lower bound in terms of the scattering length. For the upper bound, when $a\rightarrow \infty$, the power-law scaling is suppressed by the unitary behavior of the resonance caused by the strong three-body collisions. For the lower bound, $a\rightarrow 0$, the finite range effect modifies the scaling law by the effective scattering length $L_e$. These results indicate that the three-body recombination rate in a fermionic system could be characterized by the scaling law associated with the generalized Efimov physics.

cond-mat.quant-gas

Controllable Production of Degenerate Fermi Gases of $^6$Li Atoms in the 2D-3D Crossover

The many-body physics in the dimensional crossover regime attracts much attention in cold atom experiments, but yet to explore systematically. One of the technical difficulties existed in the experiments is the lack of the experimental technique to quantitatively tune the atom occupation ratio of the different lattice bands. In this letter, we report such techniques in a process of transferring a 3D Fermi gas into a 1D optical lattice, where the capability of tuning the occupation of the energy band is realized by varying the trapping potentials of the optical dipole trap (ODT) and the lattice, respectively. We could tune a Fermi gas with the occupation in the lowest band from unity to 50$\%$ quantitatively. This provides a route to experimentally study the dependence of many-body interaction on the dimensionality in a Fermi gas.

cond-mat.quant-gas