Searcharxiv⌕ Search

arXiv subjects

Ping Li

Publications and source records attributed to Ping Li.

At least 109 records · Page 6Linked to original sources

Pseudo-labeling with Keyword Refining for Few-Supervised Video Captioning

Video captioning generate a sentence that describes the video content. Existing methods always require a number of captions (\eg, 10 or 20) per video to train the model, which is quite costly. In this work, we explore the possibility of using only one or very few ground-truth sentences, and introduce a new task named few-supervised video captioning. Specifically, we propose a few-supervised video captioning framework that consists of lexically constrained pseudo-labeling module and keyword-refined captioning module. Unlike the random sampling in natural language processing that may cause invalid modifications (\ie, edit words), the former module guides the model to edit words using some actions (\eg, copy, replace, insert, and delete) by a pretrained token-level classifier, and then fine-tunes candidate sentences by a pretrained language model. Meanwhile, the former employs the repetition penalized sampling to encourage the model to yield concise pseudo-labeled sentences with less repetition, and selects the most relevant sentences upon a pretrained video-text model. Moreover, to keep semantic consistency between pseudo-labeled sentences and video content, we develop the transformer-based keyword refiner with the video-keyword gated fusion strategy to emphasize more on relevant words. Extensive experiments on several benchmarks demonstrate the advantages of the proposed approach in both few-supervised and fully-supervised scenarios. The code implementation is available at https://github.com/mlvccn/PKG_VidCap

cs.CV↗

Toward Generalizing Visual Brain Decoding to Unseen Subjects

Visual brain decoding aims to decode visual information from human brain activities. Despite the great progress, one critical limitation of current brain decoding research lies in the lack of generalization capability to unseen subjects. Prior works typically focus on decoding brain activity of individuals based on the observation that different subjects exhibit different brain activities, while it remains unclear whether brain decoding can be generalized to unseen subjects. This study aims to answer this question. We first consolidate an image-fMRI dataset consisting of stimulus-image and fMRI-response pairs, involving 177 subjects in the movie-viewing task of the Human Connectome Project (HCP). This dataset allows us to investigate the brain decoding performance with the increase of participants. We then present a learning paradigm that applies uniform processing across all subjects, instead of employing different network heads or tokenizers for individuals as in previous methods, which can accommodate a large number of subjects to explore the generalization capability across different subjects. A series of experiments are conducted and we have the following findings. First, the network exhibits clear generalization capabilities with the increase of training subjects. Second, the generalization capability is common to popular network architectures (MLP, CNN and Transformer). Third, the generalization performance is affected by the similarity between subjects. Our findings reveal the inherent similarities in brain activities across individuals. With the emerging of larger and more comprehensive datasets, it is possible to train a brain decoding foundation model in the future. Codes and models can be found at https://github.com/Xiangtaokong/TGBD.

cs.CV↗

Ferroelectricity-Driven Metallicity and Magnetic Skyrmions in van der Waals Cr2Ge2Te6/Hf2Ge2Te6 Multiferroic Heterostructure

Two-dimensional (2D) multiferroic heterostructures present a promising platform for advanced spin devices by leveraging the coexisting ferromagnetic (FM) and ferroelectric (FE) orders. Through first-principles calculations and micromagnetic simulations, we reveal non-volatile control of metallicity and topological spin textures in the Cr2Ge2Te6/Hf2Ge2Te6(CGT/HGT) heterostructure. Notably, manipulating ferroelectric polarization in HGT significantly modulates the magnetic anisotropy energy (MAE) and Dzyaloshinskii-Moriya interaction (DMI) of CGT/HGT, reversing the easy magnetization axis from in-plane to out-of-plane. By analyzing the atomic-resolved SOC energy (ΔEsoc), it is found that the cause of the change comes from the Fert-Levy mechanism. Additionally, this polarization control enables the creation and annihilation of bimerons and skyrmions, with interlayer sliding further altering magnetic ordering. Our findings offer valuable insights into magnetoelectric coupling and spin texture manipulation in 2D magnets, highlighting their potential for next-generation spintronic and memory devices.

cond-mat.mtrl-sci↗

KM UMa: An active short-period detached eclipsing binary in a hierarchical quadruple system

The first detailed photometric and spectroscopic analysis of the G-type eclipsing binary KM UMa is presented, which indicates that the system is a short-period detached eclipsing binary. The radial velocity curves were calculated using the cross-correlation function method based on Large Sky Area Multi-Object Fiber Spectroscopic Telescope, Sloan Digital Sky Survey, and our observations, which determined the mass ratio as $q=0.45\ (\pm0.04)$. Based on the light curves from the Transiting Exoplanet Survey Satellite, other survey data, and our multiband observations, the positive and negative O'Connell effects have been detected evolving gradually and alternately over the last 20 yr, which can be explained by the presence of spots on the primary component. A superflare event was detected in the SuperWASP data on 2007 February 28, further indicating that KM UMa is a very active system. We calculated its energy to be $5\times10^{34}$ erg by assuming it occurred on the primary star. Utilizing hundreds of medium-resolution spectra and one low-resolution spectrum, the equivalent width variations of the $H_α$ line were calculated, indicating the presence of a 5.21 ($\pm0.67$) yr magnetic activity cycle. The orbital period variations were analyzed using the O-C method, detecting a long-term decrease superimposed with a periodic variation. The amplitude of the cyclic variation is $0.01124\ (\pm0.00004)$ day, with a period of $33.66\ (\pm 0.0012)$ yr, which exceeds the 5.21 yr activity cycle, suggesting that this is more likely attributable to the light travel time effect of a third body. Simultaneously, a visual companion has been detected based on the Gaia astrometric data, indicating that KM UMa is actually in a 2+1+1 hierarchical quadruple system.

astro-ph.SR↗

Computer-aided Colorization State-of-the-science: A Survey

This paper reviews published research in the field of computer-aided colorization technology. We argue that the colorization task originates from computer graphics, prospers by introducing computer vision, and tends to the fusion of vision and graphics, so we put forward our taxonomy and organize the whole paper chronologically. We extend the existing reconstruction-based colorization evaluation techniques, considering that aesthetic assessment of colored images should be introduced to ensure that colorization satisfies human visual-related requirements and emotions more closely. We perform the colorization aesthetic assessment on seven representative unconditional colorization models and discuss the difference between our assessment and the existing reconstruction-based metrics. Finally, this paper identifies unresolved issues and proposes fruitful areas for future research and development. Access to the project associated with this survey can be obtained at https://github.com/DanielCho-HK/Colorization.

cs.CV↗

Projective Proximal Gradient Descent for A Class of Nonconvex Nonsmooth Optimization Problems: Fast Convergence Without Kurdyka-Lojasiewicz (KL) Property

Nonconvex and nonsmooth optimization problems are important and challenging for statistics and machine learning. In this paper, we propose Projected Proximal Gradient Descent (PPGD) which solves a class of nonconvex and nonsmooth optimization problems, where the nonconvexity and nonsmoothness come from a nonsmooth regularization term which is nonconvex but piecewise convex. In contrast with existing convergence analysis of accelerated PGD methods for nonconvex and nonsmooth problems based on the Kurdyka-Łojasiewicz (KŁ) property, we provide a new theoretical analysis showing local fast convergence of PPGD. It is proved that PPGD achieves a fast convergence rate of $\cO(1/k^2)$ when the iteration number $k \ge k_0$ for a finite $k_0$ on a class of nonconvex and nonsmooth problems under mild assumptions, which is locally Nesterov's optimal convergence rate of first-order methods on smooth and convex objective function with Lipschitz continuous gradient. Experimental results demonstrate the effectiveness of PPGD.

math.OC↗

Conflict-free chromatic index of trees

A graph $G$ is conflict-free $k$-edge-colorable if there exists an assignment of $k$ colors to $E(G)$ such that for every edge $e\in E(G)$, there is a color that is assigned to exactly one edge among the closed neighborhood of $e$. The smallest $k$ such that $G$ is conflict-free $k$-edge-colorable is called the conflict-free chromatic index of $G$, denoted $χ'_{CF}(G)$. Dȩbski and Przyby\a{l}o showed that $2\leχ'_{CF}(T)\le 3$ for every tree $T$ of size at least two. In this paper, we present an algorithm to determine the conflict-free chromatic index of a tree without 2-degree vertices, in time $O(|V(T)|)$. This partially answer a question raised by Kamyczura, Meszka and Przyby\a{l}o.

cs.DM↗

Laboratorial radiative shocks with multiple parameters and first quantifying verifications to core-collapse supernovae

We present experiments to reproduce the characteristics of core-collapse supernovae with different stellar masses and initial explosion energies in the laboratory. In the experiments, shocks are driven in 1.2 atm and 1.9 atm xenon gas by laser with energy from 1600J to 2800J on the SGIII prototype laser facility. The average shock velocities and shocked densities are obtained from experiments. Experimental results reveal that higher laser energy and lower Xe gas density led to higher shock velocity, and lower Xe gas initial density has a higher compression. Modeling of the experiments using the 2D radiation hydrodynamic codes Icefire shows excellent agreement with the experimental results and gives the temperature. These results will contribute to time-domain astrophysical systems, such as gravitational supernovae, where a strong radiative shock propagates outward from the center of the star after the core collapses.

astro-ph.HE↗

Decay estimates for Beam equations with potentials in dimension three

This paper is devoted to studying time decay estimates of the solution for Beam equation (higher order type wave equation) with a potential $$u_{t t}+\big(Δ^2+V\big)u=0, \,\ u(0, x)=f(x),\ u_{t}(0, x)=g(x)$$ in dimension three, where $V$ is a real-valued and decaying potential on $\R^3$. Assume that zero is a regular point of $H:= Δ^2+V $, we first prove the following optimal time decay estimates of the solution operators \begin{equation*} \big\|\cos (t\sqrt{H})P_{ac}(H)\big\|_{L^{1} \rightarrow L^{\infty}} \lesssim|t|^{-\frac{3}{2}}\ \ \hbox{and} \ \ \Big\|\frac{\sin(t\sqrt{H})}{\sqrt{H}} P_{a c}(H)\Big\|_{L^{1} \rightarrow L^{\infty}} \lesssim|t|^{-\frac{1}{2}}. \end{equation*} Moreover, if zero is a resonance of $H$, then time decay of the solution operators above also are considered. It is noticed that the first kind resonance does not effect the decay rates for the propagator operators $\cos(t\sqrt{H})$ and $\frac{\sin(t\sqrt{H})}{\sqrt{H}}$, but their decay will be dramatically changed for the second and third resonance types.

math.AP↗

Ferroelectric tuning of the valley polarized metal-semiconductor transition in Mn2P2S3Se3/Sc2CO2 van der Waals heterostructures and application to nonlinear Hall effect devices

In order to promote the development of the next generation of nano-spintronic devices, it is of great significance to tune the freedom of valley in two-dimensional (2D) materials. Here, we propose a mechanism for manipulating the valley and nonlinear Hall effect by the 2D ferroelectric substrate. The monolayer Mn2P2S3Se3 is a robust antiferromagnetic valley polarized semiconductor. Importantly, the valley polarized metal-semiconductor phase transition of Mn2P2S3Se3 can be effectively tuned by switching the ferroelectric polarization of Sc2CO2. We reveal the microscopic mechanism of phase transition, which origins from the charge transfer and band alignment. Additionally, we find that transformed polarization direction of Sc2CO2 flexibly manipulate the Berry curvature dipole. Based on this discovery, we present the detection valley polarized metal-semiconductor transition by the nonlinear Hall effect devices. These findings not only offer a scheme to tune the valley degree of freedom, but also provide promising platform to design the nonlinear Hall effect devices.

cond-mat.mtrl-sci↗

Large Margin Prototypical Network for Few-shot Relation Classification with Fine-grained Features

Relation classification (RC) plays a pivotal role in both natural language understanding and knowledge graph completion. It is generally formulated as a task to recognize the relationship between two entities of interest appearing in a free-text sentence. Conventional approaches on RC, regardless of feature engineering or deep learning based, can obtain promising performance on categorizing common types of relation leaving a large proportion of unrecognizable long-tail relations due to insufficient labeled instances for training. In this paper, we consider few-shot learning is of great practical significance to RC and thus improve a modern framework of metric learning for few-shot RC. Specifically, we adopt the large-margin ProtoNet with fine-grained features, expecting they can generalize well on long-tail relations. Extensive experiments were conducted by FewRel, a large-scale supervised few-shot RC dataset, to evaluate our framework: LM-ProtoNet (FGF). The results demonstrate that it can achieve substantial improvements over many baseline approaches.

cs.CL↗

MOBIUS: Towards the Next Generation of Query-Ad Matching in Baidu's Sponsored Search

Baidu runs the largest commercial web search engine in China, serving hundreds of millions of online users every day in response to a great variety of queries. In order to build a high-efficiency sponsored search engine, we used to adopt a three-layer funnel-shaped structure to screen and sort hundreds of ads from billions of ad candidates subject to the requirement of low response latency and the restraints of computing resources. Given a user query, the top matching layer is responsible for providing semantically relevant ad candidates to the next layer, while the ranking layer at the bottom concerns more about business indicators (e.g., CPM, ROI, etc.) of those ads. The clear separation between the matching and ranking objectives results in a lower commercial return. The Mobius project has been established to address this serious issue. It is our first attempt to train the matching layer to consider CPM as an additional optimization objective besides the query-ad relevance, via directly predicting CTR (click-through rate) from billions of query-ad pairs. Specifically, this paper will elaborate on how we adopt active learning to overcome the insufficiency of click history at the matching layer when training our neural click networks offline, and how we use the SOTA ANN search technique for retrieving ads more efficiently (Here ``ANN'' stands for approximate nearest neighbor search). We contribute the solutions to Mobius-V1 as the first version of our next generation query-ad matching system.

cs.IR↗

Highly Efficient and Stable Perovskite Solar Cells via MultiFunctional Curcumin Modified Buried Interface

The buried interface between the electron transport layer and the perovskite layer suffers from severe interface defects and imperfect energy level alignment. To address this issue, this study employs a multifunctional organic molecule, curcumin, to modify the interface between SnO2 and the perovskite layer. The functional groups on curcumin effectively passivate the defects on both sides of the interface, reducing -OH and oxygen vacancy defects on the SnO2 surface and passivating uncoordinated Pb2+ in the perovskite layer. This results in a more compatible energy level alignment and lower defect density at the interface, enhancing carrier transport across it. Consequently, the devices based on curcumin achieve an impressive champion power conversion efficiency (PCE) of 24.46%, compared to 22.03% for control devices. This work demonstrates a simple, green, hydrophobic, and efficient molecular modification method for the buried interface, laying the foundation for the development of high-performance and stable perovskite solar cells.

cond-mat.mtrl-sci↗

Tilted Disk Precession and Negative Superhumps in HS 2325+8205: A Multi-Window Analysis

Tilted disk precession exists in different objects. Negative superhumps (NSHs) in cataclysmic variable stars (CVs) are believed to arise from the interaction between the reverse precession of a tilted disk and the streams from the secondary star.Utilizing TESS photometry, we present a comprehensive investigation into the tilted disk precession and NSHs in the dwarf nova (DN) HS 2325+8205, employing eclipse minima, eclipse depths, NSH frequencies, and NSH amplitudes and the correlation between them as the windows. We identified NSHs with a period of 0.185671(17) days in HS 2325+8205. The NSH frequency exhibits variability with a period of 3.943(9) days, akin to the tilted disk precession period validated in novae-like stars (NLs, SDSS J0812) and intermediate polars (IPs, TV Col).The O-C of eclipse minima were similarly found to vary cyclically in period 4.135(5) days, characterized by a faster rise than fall. Furthermore, the NSH amplitude exhibits complex and diverse variations, which may be linked to changes in the disk radius, mass transfer rate, and the apparent area of the hot spot. For the first time in DNe, we observe bi-periodic variations in eclipse depth (P1= 4.131(4) d and P2= 2.065(2) d ~ Pprec/2), resembling those seen in IPs, suggesting that variations with P2 are not attributable to an accretion curtain, as previously suspected. Moreover, NSH amplitude and eclipse depth decrease with increasing NSH frequency, while NSH amplitude correlates positively with eclipse depth.These complex variations observed across multiple observational windows provide substantial evidence for understanding of tilted disk precession and NSHs.

astro-ph.SR↗

Timeline and Boundary Guided Diffusion Network for Video Shadow Detection

Video Shadow Detection (VSD) aims to detect the shadow masks with frame sequence. Existing works suffer from inefficient temporal learning. Moreover, few works address the VSD problem by considering the characteristic (i.e., boundary) of shadow. Motivated by this, we propose a Timeline and Boundary Guided Diffusion (TBGDiff) network for VSD where we take account of the past-future temporal guidance and boundary information jointly. In detail, we design a Dual Scale Aggregation (DSA) module for better temporal understanding by rethinking the affinity of the long-term and short-term frames for the clipped video. Next, we introduce Shadow Boundary Aware Attention (SBAA) to utilize the edge contexts for capturing the characteristics of shadows. Moreover, we are the first to introduce the Diffusion model for VSD in which we explore a Space-Time Encoded Embedding (STEE) to inject the temporal guidance for Diffusion to conduct shadow detection. Benefiting from these designs, our model can not only capture the temporal information but also the shadow property. Extensive experiments show that the performance of our approach overtakes the state-of-the-art methods, verifying the effectiveness of our components. We release the codes, weights, and results at \url{https://github.com/haipengzhou856/TBGDiff}.

cs.CV↗

A Convex-optimization-based Layer-wise Post-training Pruner for Large Language Models

Pruning is a critical strategy for compressing trained large language models (LLMs), aiming at substantial memory conservation and computational acceleration without compromising performance. However, existing pruning methods often necessitate inefficient retraining for billion-scale LLMs or rely on heuristic methods such as the optimal brain surgeon framework, which degrade performance. In this paper, we introduce FISTAPruner, the first post-training pruner based on convex optimization models and algorithms. Specifically, we propose a convex optimization model incorporating $\ell_1$ norm to induce sparsity and utilize the FISTA solver for optimization. FISTAPruner incorporates an intra-layer cumulative error correction mechanism and supports parallel pruning. We comprehensively evaluate FISTAPruner on models such as OPT, LLaMA, LLaMA-2, and LLaMA-3 with 125M to 70B parameters under unstructured and 2:4 semi-structured sparsity, demonstrating superior performance over existing state-of-the-art methods across various language benchmarks.

cs.LG↗

Under-confidence Backdoors Are Resilient and Stealthy Backdoors

By injecting a small number of poisoned samples into the training set, backdoor attacks aim to make the victim model produce designed outputs on any input injected with pre-designed backdoors. In order to achieve a high attack success rate using as few poisoned training samples as possible, most existing attack methods change the labels of the poisoned samples to the target class. This practice often results in severe over-fitting of the victim model over the backdoors, making the attack quite effective in output control but easier to be identified by human inspection or automatic defense algorithms. In this work, we proposed a label-smoothing strategy to overcome the over-fitting problem of these attack methods, obtaining a \textit{Label-Smoothed Backdoor Attack} (LSBA). In the LSBA, the label of the poisoned sample $\bm{x}$ will be changed to the target class with a probability of $p_n(\bm{x})$ instead of 100\%, and the value of $p_n(\bm{x})$ is specifically designed to make the prediction probability the target class be only slightly greater than those of the other classes. Empirical studies on several existing backdoor attacks show that our strategy can considerably improve the stealthiness of these attacks and, at the same time, achieve a high attack success rate. In addition, our strategy makes it able to manually control the prediction probability of the design output through manipulating the applied and activated number of LSBAs\footnote{Source code will be published at \url{https://github.com/v-mipeng/LabelSmoothedAttack.git}}.

cs.CR↗

Less is More: Sparse Watermarking in LLMs with Enhanced Text Quality

With the widespread adoption of Large Language Models (LLMs), concerns about potential misuse have emerged. To this end, watermarking has been adapted to LLM, enabling a simple and effective way to detect and monitor generated text. However, while the existing methods can differentiate between watermarked and unwatermarked text with high accuracy, they often face a trade-off between the quality of the generated text and the effectiveness of the watermarking process. In this work, we present a novel type of LLM watermark, Sparse Watermark, which aims to mitigate this trade-off by applying watermarks to a small subset of generated tokens distributed across the text. The key strategy involves anchoring watermarked tokens to words that have specific Part-of-Speech (POS) tags. Our experimental results demonstrate that the proposed watermarking scheme achieves high detectability while generating text that outperforms previous LLM watermarking methods in quality across various tasks

cs.CR↗