SearcharxivSearch

arXiv subjects

Chunyu Sui

Publications and source records attributed to Chunyu Sui.

5 recordsLinked to original sources

SignX: Continuous Sign Recognition in Compact Pose-Rich Latent Space

The complexity of Sign Language (SL) data processing brings many challenges. The current approach to recognition of SL signs aims to translate RGB sign language videos through pose information into Word-based ID Glosses, which serve to uniquely identify signs. This paper proposes SignX, a novel framework for continuous sign language recognition (SLR) in compact pose-rich latent space. First, we construct a unified latent representation that encodes heterogeneous pose formats (SMPLer-X, DWPose, Mediapipe, PrimeDepth, and Sapiens Segmentation) into a compact, information-dense space. Second, we train a ViT-based Video-to-Pose module to extract this latent representation directly from raw videos. Finally, we develop a temporal modeling and sequence refinement method that operates entirely in this latent space. This multi-stage design achieves end-to-end SLR while significantly reducing computational consumption. Experimental results demonstrate that SignX achieves SOTA accuracy on continuous SLR and Translation task, delivering nearly a 50-fold acceleration over pixel-space baselines.

cs.CV

SignLLM: Sign Language Production Large Language Models

In this paper, we propose SignLLM, a multilingual Sign Language Production (SLP) large language model, which includes two novel multilingual SLP modes MLSF and Prompt2LangGloss that allow sign language gestures generation from query texts input and question-style prompts input respectively. Both modes can use a new RL loss based on reinforcement learning and a new RL module named Priority Learning Channel. These RL components can accelerate the training by enhancing the model's capability to sample high-quality data. To train SignLLM, we introduce Prompt2Sign, a comprehensive multilingual sign language dataset, which builds from public data, including American Sign Language (ASL) and seven others. This dataset standardizes information by extracting pose information from sign language videos into a unified compressed format. We extensively evaluate SignLLM, demonstrating that our model achieves state-of-the-art performance on SLP tasks across eight sign languages.

cs.CV

SignDiff: Diffusion Model for American Sign Language Production

In this paper, we propose a dual-condition diffusion pre-training model named SignDiff that can generate human sign language speakers from a skeleton pose. SignDiff has a novel Frame Reinforcement Network called FR-Net, similar to dense human pose estimation work, which enhances the correspondence between text lexical symbols and sign language dense pose frames, reduces the occurrence of multiple fingers in the diffusion model. In addition, we propose a new method for American Sign Language Production (ASLP), which can generate ASL skeletal pose videos from text input, integrating two new improved modules and a new loss function to improve the accuracy and quality of sign language skeletal posture and enhance the ability of the model to train on large-scale data. We propose the first baseline for ASL production and report the scores of 17.19 and 12.85 on BLEU-4 on the How2Sign dev/test sets. We evaluated our model on the previous mainstream dataset PHOENIX14T, and the experiments achieved the SOTA results. In addition, our image quality far exceeds all previous results by 10 percentage points in terms of SSIM.

cs.CV

Applying Back Propagation Algorithm and Analytic Hierarchy Process to Environment Assessment

This paper designs a new and scientific environmental quality assessment method, and takes Saihan dam as an example to explore the environmental improvement degree to the local and Beijing areas. AHP method is used to assign values to each weight 7 primary indicators and 21 secondary indicators were used to establish an environmental quality assessment model. The conclusion shows that after the establishment of Saihan dam, the local environmental quality has been improved by 7 times, and the environmental quality in Beijing has been improved by 13%. Then the future environmental index is predicted. Finally the Spearson correlation coefficient is analyzed, and it is proved that correlation is 99% when the back-propagation algorithm is used to test and prove that the error is little.

cs.CY

GES Model :Combining Pearson Correlation Coefficient Analysis with Multilayer Perceptron

With the development of technological progress, mining on asteroids is becoming a reality. This paper focuses on how to distribute asteroid mineral resources in a reasonable way to ensure global equity. To distribute asteroid resources fairly, 7 primary indicators and 20 secondary indicators are introduced to build a mathematical model to evaluate global equity and the weights are given by Analytic Hierarchy Process (AHP). Then Global Equity Score(GES) Model based on 12 primary indicators and 40 secondary indicators is built and TOPSIS method is applied to rank all countries. A t-distribution probability density function is applied to simulate the rate of asteroid mining. The Backward Algorithm is applied to quantitatively measure the impact of changing indicators on global equity. Then Pearson correlation coefficient analysis is conducted for each indicator, and t-test is performed lastly. The results demonstrate that asteroid mining promotes global equity that poor countries can be allocated slightly more mineral resources, and a schedule of the implementation of each measure is given. To gain more insight, sensitivity analysis is conducted and the results demonstrate that scores vary less than 7%. It can be concluded that our GES model have great potential as its robustness, accuracy and strengths.

cs.CE