SearcharxivSearch

arXiv subjects

Wanyi Zhang

Publications and source records attributed to Wanyi Zhang.

10 recordsLinked to original sources

GSAM: A Generalizable and Safe Robotic Framework for Articulated Object Manipulation

Articulated object manipulation is a unique challenge for service robots. Existing methods employ end-to-end policy learning, visionmotion planning, and large-language/visual-language model (LLM/VLM), but often overlook the diversity of articulated objects and the complexity of interactions between end-effector and handle, leading to limited generalization and destructive collisions. To address this, we propose GSAM, a generalizable and safe robotic framework for articulated object manipulation. Specifically, a vision-based perceiver generates the kinematic parameters. Considering that pre-trained markers in perceiver yield raw estimations that may deviate from commonsense, we present a f ine-tuned VLM-based refiner, using chain-of-thought (COT) commonsense reasoning to refine perception. To prevent destructive collisions, we design an interaction constraint function generator, integrating articulated object, interaction pose, and obstacle avoidance knowledge into a base. LLM then functionalize these constraints and apply them to trajectory and posture planning. A kinematic-aware manipulation planner verifies reachability for trajectory and posture. Experiments on 50 hinge tasks across 5 object categories and 50 randomly initialized end-effectorhandle configurations show that GSAM reduces standard deviation by 3.1% and improves manipulation success rate by 36.0% compared to the best baseline, respectively demonstrating the superior object generalization and interaction safety of GSAM in practical scenarios.

cs.RO

Oxygen permeability and stability in the entropy-stabilized Co-based Perovskite oxygen permeable membranes

Oxygen transport membranes (OTMs), enabling catalytic reaction and gas separation, support crucial chemical engineering processes and decarbonization technologies, but their applications are hindered by limited oxygen permeation fluxes and inadequate long-term stability during operation. Here, a series of high-entropy perovskite OTMs based on La0.5Sr0.5CoO3 were designed and synthesized by the simple sol-gel method. The impact of varying doping ratios on the structure, surface morphology, oxygen permeability, and stability of these high-entropy OTMs was thoroughly examined. At 950 °C, the optimal composition, La0.25Sr0.25Gd0.2Nd0.2Pr0.1CoO3, achieved oxygen permeation fluxes of 1.62 mL min-1 cm-2 under air/He gradient and 1.46 mL min-1 cm-2 under air/CO2, respectively. Remarkably, all high-entropy OTMs demonstrated stable operation for over 100 h in a pure CO2 environment without a significant decline in performance. This finding paves a new way to enhance the structural and oxygen permeation stability of OTMs, and further promotes the application of OTMs in oxy-fuel combustion technologies aimed at improving CO2 capture and storage efficiency.

cond-mat.mtrl-sci

Efficient photocatalytic CO2 Reduction to C2+ Products with Pt1-xPdxSn4 Dirac Nodal Arc Semimetal

The photochemical CO2 reduction reaction represents a zero-carbon pathway for converting CO2 into value-added chemicals, yet its industrial implementation has been constrained by low selectivity and product diversity. Dirac nodal arc semimetals characterized by ultrahigh carrier mobility with over 25000 cm2 V-1 s-1 offer a promising platform to search for efficient catalysts for CO2 conversion. Herein, we demonstrate that strategic Pt incorporation into PdSn4 optimizes the electronic structure and carrier dynamics of this Dirac semimetal. Experimental and theoretical analyses reveal that the resulting Pd-Sn-Pt local electronic structure redistributes charge density around Pd and Pt atoms, which facilitates C-C coupling via *OC-COH and *OC-CHOH intermediates and enhances carrier mobility by 40% versus the pristine PdSn4 single crystal. The optimized Pd0.4Pt0.6Sn4 single crystal achieves C2H4 with formation rate of 0.000328 mol g-1 h-1, product selectivity of 73.1% and electron-based selectivity of 89%. This work establishes electronic-structure-tunable Dirac semimetals as a new paradigm for multi-carbon photochemical CO2 reduction, providing a design strategy for next-generation photocatalysts.

cond-mat.mtrl-sci

Photonic restricted Boltzmann machine for content generation tasks

The restricted Boltzmann machine (RBM) is a neural network based on the Ising model, well known for its ability to learn probability distributions and stochastically generate new content. However, the high computational cost of Gibbs sampling in content generation tasks imposes significant bottlenecks on electronic implementations. Here, we propose a photonic restricted Boltzmann machine (PRBM) that leverages photonic computing to accelerate Gibbs sampling, enabling efficient content generation. By introducing an efficient encoding method, the PRBM eliminates the need for computationally intensive matrix decomposition and reduces the computational complexity of Gibbs sampling from $O(N)$ to $O(1)$. Moreover, its non-Von Neumann photonic computing architecture circumvents the memory storage of interaction matrices, providing substantial advantages for large-scale RBMs. We experimentally validate the photonic-accelerated Gibbs sampling by simulating a two-dimensional Ising model, where the observed phase transition temperature closely matches the theoretical predictions. Beyond physics-inspired tasks, the PRBM demonstrates robust capabilities in generating and restoring diverse content, including images and temporal sequences, even in the presence of noise and aberrations. The scalability and reduced training cost of the PRBM framework underscore its potential as a promising pathway for advancing photonic computing in generative artificial intelligence.

physics.optics

SLTNet: Efficient Event-based Semantic Segmentation with Spike-driven Lightweight Transformer-based Networks

Event-based semantic segmentation has great potential in autonomous driving and robotics due to the advantages of event cameras, such as high dynamic range, low latency, and low power cost. Unfortunately, current artificial neural network (ANN)-based segmentation methods suffer from high computational demands, the requirements for image frames, and massive energy consumption, limiting their efficiency and application on resource-constrained edge/mobile platforms. To address these problems, we introduce SLTNet, a spike-driven lightweight transformer-based network designed for event-based semantic segmentation. Specifically, SLTNet is built on efficient spike-driven convolution blocks (SCBs) to extract rich semantic features while reducing the model's parameters. Then, to enhance the long-range contextural feature interaction, we propose novel spike-driven transformer blocks (STBs) with binary mask operations. Based on these basic blocks, SLTNet employs a high-efficiency single-branch architecture while maintaining the low energy consumption of the Spiking Neural Network (SNN). Finally, extensive experiments on DDD17 and DSEC-Semantic datasets demonstrate that SLTNet outperforms state-of-the-art (SOTA) SNN-based methods by at most 9.06% and 9.39% mIoU, respectively, with extremely 4.58x lower energy consumption and 114 FPS inference speed. Our code is open-sourced and available at https://github.com/longxianlei/SLTNet-v1.0.

cs.CV

Strongly correlated electronic superconductivity in the noncentrosymmetric Re-Os-based high/medium-entropy alloys

The class of unconventional superconductors, particularly noncentrosymmetric superconductors, has been highly considered as potential materials for understanding the complex properties of quantum materials. Here, five previously unreported Re3.5Os3.5Ta0.5Hf0.5Nb3, Re3Os3Ta0.5Hf0.5Nb3, Re3.5Os3.5Mo0.5Hf0.5Nb3, Re3.5Os3.5Mo0.5W0.5Nb3, and Re3Os3Mo0.5Hf0.5Nb3 Re-Os-based high/medium-entropy alloys (MEAs-HEAs) with valence electron count ranging from 6.45 to 6.81 were synthesized and investigated using x-ray diffraction, transport, magnetization, and specific heat measurements. Our analyses confirm that all five compounds crystallize in a noncentrosymmetric α-Mn-type structure and exhibit type-II superconductivity with Tc values from 4.20 K to 5.11 K, respectively. Unexpectedly, despite being immersed in an acidic environment for one month, the structures and superconducting properties of HEAs remain stable. Our findings indicate that the Tc increases with an increasing valence electron count in MEAs-HEAs. Furthermore, these noncentrosymmetric α-Mn-type HEA superconductors have large Kadowaki-Woods ratios (KWR), implying the presence of strong electronic correlations.

cond-mat.supr-con

Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice

Simultaneous Interpretation (SI) represents one of the most daunting frontiers in the translation industry, with product-level automatic systems long plagued by intractable challenges: subpar transcription and translation quality, lack of real-time speech generation, multi-speaker confusion, and translated speech inflation, especially in long-form discourses. In this study, we introduce Seed-LiveInterpret 2.0, an end-to-end SI model that delivers high-fidelity, ultra-low-latency speech-to-speech generation with voice cloning capabilities. As a fully operational product-level solution, Seed-LiveInterpret 2.0 tackles these challenges head-on through our novel duplex speech-to-speech understanding-generating framework. Experimental results demonstrate that through large-scale pretraining and reinforcement learning, the model achieves a significantly better balance between translation accuracy and latency, validated by human interpreters to exceed 70% correctness in complex scenarios. Notably, Seed-LiveInterpret 2.0 outperforms commercial SI solutions by significant margins in translation quality, while slashing the average latency of cloned speech from nearly 10 seconds to a near-real-time 3 seconds, which is around a near 70% reduction that drastically enhances practical usability.

cs.CL

Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition

Modern automatic speech recognition (ASR) model is required to accurately transcribe diverse speech signals (from different domains, languages, accents, etc) given the specific contextual information in various application scenarios. Classic end-to-end models fused with extra language models perform well, but mainly in data matching scenarios and are gradually approaching a bottleneck. In this work, we introduce Seed-ASR, a large language model (LLM) based speech recognition model. Seed-ASR is developed based on the framework of audio conditioned LLM (AcLLM), leveraging the capabilities of LLMs by inputting continuous speech representations together with contextual information into the LLM. Through stage-wise large-scale training and the elicitation of context-aware capabilities in LLM, Seed-ASR demonstrates significant improvement over end-to-end models on comprehensive evaluation sets, including multiple domains, accents/dialects and languages. Additionally, Seed-ASR can be further deployed to support specific needs in various scenarios without requiring extra language models. Compared to recently released large ASR models, Seed-ASR achieves 10%-40% reduction in word (or character, for Chinese) error rates on Chinese and English public test sets, further demonstrating its powerful performance.

eess.AS

Multi-Modal Subjective Context Modelling and Recognition

Applications like personal assistants need to be aware ofthe user's context, e.g., where they are, what they are doing, and with whom. Context information is usually inferred from sensor data, like GPS sensors and accelerometers on the user's smartphone. This prediction task is known as context recognition. A well-defined context model is fundamental for successful recognition. Existing models, however, have two major limitations. First, they focus on few aspects, like location or activity, meaning that recognition methods based onthem can only compute and leverage few inter-aspect correlations. Second, existing models typically assume that context is objective, whereas in most applications context is best viewed from the user's perspective. Neglecting these factors limits the usefulness of the context model and hinders recognition. We present a novel ontological context model that captures five dimensions, namely time, location, activity, social relations and object. Moreover, our model defines three levels of description(objective context, machine context and subjective context) that naturally support subjective annotations and reasoning.An initial context recognition experiment on real-world data hints at the promise of our model.

cs.AI

Breast Cancer Classification with Ultrasound Images Based on SLIC

Ultrasound image diagnosis of breast tumors has been widely used in recent years. However, there are some problems of it, for instance, poor quality, intense noise and uneven echo distribution, which has created a huge obstacle to diagnosis. To overcome these problems, we propose a novel method, a breast cancer classification with ultrasound images based on SLIC (BCCUI). We first utilize the Region of Interest (ROI) extraction based on Simple Linear Iterative Clustering (SLIC) algorithm and region growing algorithm to extract the ROI at the super-pixel level. Next, the features of ROI are extracted. Furthermore, the Support Vector Machine (SVM) classifier is applied. The calculation states that the accuracy of this segment algorithm is up to 88.00% and the sensitivity of the algorithm is up to 92.05%, which proves that the classifier presents in this paper has certain research meaning and applied worthiness.

cs.CV