Searcharxiv⌕ Search

arXiv subjects

Songsong Xiong

Publications and source records attributed to Songsong Xiong.

6 recordsLinked to original sources

LM-MCVT: A Lightweight Multi-modal Multi-view Convolutional-Vision Transformer Approach for 3D Object Recognition

In human-centered environments such as restaurants, homes, and warehouses, robots often face challenges in accurately recognizing 3D objects. These challenges stem from the complexity and variability of these environments, including diverse object shapes. In this paper, we propose a novel Lightweight Multi-modal Multi-view Convolutional-Vision Transformer network (LM-MCVT) to enhance 3D object recognition in robotic applications. Our approach leverages the Globally Entropy-based Embeddings Fusion (GEEF) method to integrate multi-views efficiently. The LM-MCVT architecture incorporates pre- and mid-level convolutional encoders and local and global transformers to enhance feature extraction and recognition accuracy. We evaluate our method on the synthetic ModelNet40 dataset and achieve a recognition accuracy of 95.6% using a four-view setup, surpassing existing state-of-the-art methods. To further validate its effectiveness, we conduct 5-fold cross-validation on the real-world OmniObject3D dataset using the same configuration. Results consistently show superior performance, demonstrating the method's robustness in 3D object recognition across synthetic and real-world 3D data.

cs.CV↗

HMT-Grasp: A Hybrid Mamba-Transformer Approach for Robot Grasping in Cluttered Environments

Robot grasping, whether handling isolated objects, cluttered items, or stacked objects, plays a critical role in industrial and service applications. However, current visual grasp detection methods based on Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) often struggle to adapt to diverse scenarios, as they tend to emphasize either local or global features exclusively, neglecting complementary cues. In this paper, we propose a novel hybrid Mamba-Transformer approach to address these challenges. Our method improves robotic visual grasping by effectively capturing both global and local information through the integration of Vision Mamba and parallel convolutional-transformer blocks. This hybrid architecture significantly improves adaptability, precision, and flexibility across various robotic tasks. To ensure a fair evaluation, we conducted extensive experiments on the Cornell, Jacquard, and OCID-Grasp datasets, ranging from simple to complex scenarios. Additionally, we performed both simulated and real-world robotic experiments. The results demonstrate that our method not only surpasses state-of-the-art techniques on standard grasping datasets but also delivers strong performance in both simulation and real-world robot applications.

cs.RO↗

Lifelong Ensemble Learning based on Multiple Representations for Few-Shot Object Recognition

Service robots are integrating more and more into our daily lives to help us with various tasks. In such environments, robots frequently face new objects while working in the environment and need to learn them in an open-ended fashion. Furthermore, such robots must be able to recognize a wide range of object categories. In this paper, we present a lifelong ensemble learning approach based on multiple representations to address the few-shot object recognition problem. In particular, we form ensemble methods based on deep representations and handcrafted 3D shape descriptors. To facilitate lifelong learning, each approach is equipped with a memory unit for storing and retrieving object information instantly. The proposed model is suitable for open-ended learning scenarios where the number of 3D object categories is not fixed and can grow over time. We have performed extensive sets of experiments to assess the performance of the proposed approach in offline, and open-ended scenarios. For the evaluation purpose, in addition to real object datasets, we generate a large synthetic household objects dataset consisting of 27000 views of 90 objects. Experimental results demonstrate the effectiveness of the proposed method on online few-shot 3D object recognition tasks, as well as its superior performance over the state-of-the-art open-ended learning approaches. Furthermore, our results show that while ensemble learning is modestly beneficial in offline settings, it is significantly beneficial in lifelong few-shot learning situations. Additionally, we demonstrated the effectiveness of our approach in both simulated and real-robot settings, where the robot rapidly learned new categories from limited examples.

cs.RO↗

Enhancing Fine-Grained 3D Object Recognition using Hybrid Multi-Modal Vision Transformer-CNN Models

Robots operating in human-centered environments, such as retail stores, restaurants, and households, are often required to distinguish between similar objects in different contexts with a high degree of accuracy. However, fine-grained object recognition remains a challenge in robotics due to the high intra-category and low inter-category dissimilarities. In addition, the limited number of fine-grained 3D datasets poses a significant problem in addressing this issue effectively. In this paper, we propose a hybrid multi-modal Vision Transformer (ViT) and Convolutional Neural Networks (CNN) approach to improve the performance of fine-grained visual classification (FGVC). To address the shortage of FGVC 3D datasets, we generated two synthetic datasets. The first dataset consists of 20 categories related to restaurants with a total of 100 instances, while the second dataset contains 120 shoe instances. Our approach was evaluated on both datasets, and the results indicate that it outperforms both CNN-only and ViT-only baselines, achieving a recognition accuracy of 94.50 % and 93.51 % on the restaurant and shoe datasets, respectively. Additionally, we have made our FGVC RGB-D datasets available to the research community to enable further experimentation and advancement. Furthermore, we successfully integrated our proposed method with a robot framework and demonstrated its potential as a fine-grained perception tool in both simulated and real-world robotic scenarios.

cs.CV↗

Chlorine and Bromine Isotope Fractionation of Halogenated Organic Compounds in Electron Ionization Mass Spectrometry

Revelation of chlorine and bromine isotope fractionation of halogenated organic compounds (HOCs) in electron ionization mass spectrometry (EI-MS) is crucial for compound-specific chlorine/bromine isotope analysis (CSIA-Cl/Br) using gas chromatography EI-MS (GC-EI-MS). This study systematically investigated chlorine/bromine isotope fractionation in EI-MS of HOCs including 12 organochlorines and 5 organobromines using GC-double focus magnetic-sector high resolution MS (GC-DFS-HRMS). Chlorine/bromine isotope fractionation behaviors of the HOCs in EI-MS showed varied isotope fractionation patterns and extents depending on compounds. Besides, isotope fractionation patterns and extents varied at different EI energies, demonstrating potential impacts of EI energy on the chlorine/bromine isotope fractionation. Hypotheses of inter-ion and intra-ion isotope fractionations were applied to interpreting the isotope fractionation behaviors. The inter-ion and intra-ion isotope fractionations counteractively contributed to the apparent isotope ratio for a certain dehalogenated product ion. The isotope fractionation mechanisms were tentatively elucidated on basis of the quasi-equilibrium theory. In the light of the findings of this study, isotope ratio evaluation scheme using complete molecular ions and the EI source with sufficient stable EI energies may be helpful to achieve optimal precision and accuracy of CSIA-Cl/Br data. The method and results of this study can help to predict isotope fractionation of HOCs during dehalogenation processes and further to reveal the dehalogenation pathways.

physics.chem-ph↗

Chlorine and Bromine Isotope Fractionation of Halogenated Organic Pollutants on Gas Chromatography Columns

Compound-specific chlorine/bromine isotope analysis (CSIA-Cl/Br) has become a useful approach for degradation pathway investigation and source appointment of halogenated organic pollutants (HOPs). CSIA-Cl/Br is usually conducted by gas chromatography-mass spectrometry (GC-MS), which could be negatively impacted by chlorine and bromine isotope fractionation of HOPs on GC columns. In this study, 31 organochlorines and 4 organobromines were systematically investigated in terms of Cl/Br isotope fractionation on GC columns using GC-double focus magnetic-sector high resolution MS (GC-DFS-HRMS). On-column chlorine/bromine isotope fractionation behaviors of the HOPs were explored, presenting various isotope fractionation modes and extents. Twenty-nine HOPs exhibited inverse isotope fractionation, and only polychlorinated biphenyl-138 (PCB-138) and PCB-153 presented normal isotope fractionation. And no observable isotope fractionation was found for the rest four HOPs, i.e., PCB-101, 1,2,3,7,8-pentachlorodibenzofuran, PCB-180 and 2,3,7,8-tetrachlorodibenzofuran. The isotope fractionation extents of different HOPs varied from below the observable threshold (0.50%) to 7.31% (PCB-18). The mechanisms of the on-column chlorine/bromine isotope fractionation were tentatively interpreted with the Craig-Gordon model and a modified two-film model. Inverse isotope effects and normal isotope effects might contribute to the total isotope effects together and thus determine the isotope fractionation directions and extents. Proposals derived from the main results of this study for CSIA-Cl/Br research were provided for improving the precision and accuracy of CSIA-Cl/Br results. The findings of this study will shed light on the development of CSIA-Cl/Br methods using GC-MS techniques, and help to implement the research using CSIA-Cl/Br to investigate the environmental behaviors and pollution sources of HOPs.

physics.chem-ph↗