SearcharxivSearch

arXiv subjects

Bojun Zhang

Publications and source records attributed to Bojun Zhang.

11 recordsLinked to original sources

ReViCo: Unveiling the Limitations of VLMs in Visual Text Understanding via Error Correction

Vision Language Models (VLMs) have shown great success in general visual tasks, yet they still struggle to deeply understand text within images. In this paper, we introduce ReViCo (Real Visual Correction), a benchmark designed to evaluate VLM text understanding through a novel task of visual text error correction. ReViCo challenges models to identify and fix text errors in real-world images, which requires a profound understanding of the interplay between visual text and its surrounding visual context. We benchmark various VLMs using two distinct paradigms: prompt-based strategy and targeted model training, both aimed at pushing the limits of current models. Our experiments reveal a striking performance gap between even the best VLMs and human, and further analysis also shows that most models struggle to accurately perceive the visual text, resulting in frequent correction errors. By highlighting these gaps, ReViCo provides a new benchmark foundation for developing more robust and text-aware VLMs.

cs.CV

RETO: A Rotary-Enhanced Transformer Operator for High-Fidelity Prediction of Automotive Aerodynamics

Rapid aerodynamic evaluation is crucial for modern vehicle design, yet existing neural operators struggle to capture intricate spatial correlations. We propose the rotary-enhanced transformer operator (RETO), a novel neural solver featuring a dual-stage spatial awareness mechanism: sinusoidal-cosine encodings for global referencing and rotary positional encodings (RoPE) for relative displacements. RoPE encodes spatial relations via unitary rotations, enforcing translation invariance and enhancing local gradient resolution. RETO is validated on ShapeNet and the high-fidelity DrivAerML benchmark. On ShapeNet, RETO achieves a relative $L_2$ error of 0.063, outperforming RegDGCNN at 0.125 and representing a 16\% improvement over the Transolver baseline, which yields an error of 0.075. These performance gains are further amplified on the DrivAerML dataset, where RETO achieves relative $L_2$ errors of 0.089 for surface pressure and 0.097 for velocity. In comparison, Transolver results in errors of 0.116 and 0.121 for the same metrics, indicating that RETO achieves precision enhancements of 23\% and 19\%, respectively. For comprehensive comparison, the surface pressure and velocity errors for AB-UBT are 0.102 and 0.124, while RegDGCNN yields 0.235 and 0.312, respectively. Information-theoretical analysis shows that the entropy peak of RETO at 0.35 is significantly lower than that of Transolver at 0.75 under $10^4$ resolution, indicating a focused attentional mechanism capable of preserving localized gradients against global diffusion.

eess.IV

PinPoint3D: Fine-Grained 3D Part Segmentation from a Few Clicks

Fine-grained 3D part segmentation is crucial for enabling embodied AI systems to perform complex manipulation tasks, such as interacting with specific functional components of an object. However, existing interactive segmentation methods are largely confined to coarse, instance-level targets, while non-interactive approaches struggle with sparse, real-world scans and suffer from a severe lack of annotated data. To address these limitations, we introduce PinPoint3D, a novel interactive framework for fine-grained, multi-granularity 3D segmentation, capable of generating precise part-level masks from only a few user point clicks. A key component of our work is a new 3D data synthesis pipeline that we developed to create a large-scale, scene-level dataset with dense part annotations, overcoming a critical bottleneck that has hindered progress in this field. Through comprehensive experiments and user studies, we demonstrate that our method significantly outperforms existing approaches, achieving an average IoU of around 55.8% on each object part under first-click settings and surpassing 71.3% IoU with only a few additional clicks. Compared to current state-of-the-art baselines, PinPoint3D yields up to a 16% improvement in IoU and precision, highlighting its effectiveness on challenging, sparse point clouds with high efficiency. Our work represents a significant step towards more nuanced and precise machine perception and interaction in complex 3D environments.

cs.CV

Vision Language Models Are Not (Yet) Spelling Correctors

Spelling correction from visual input poses unique challenges for vision language models (VLMs), as it requires not only detecting but also correcting textual errors directly within images. We present ReViCo (Real Visual Correction), the first benchmark that systematically evaluates VLMs on real-world visual spelling correction across Chinese and English. ReViCo contains naturally occurring errors collected from real-world image data and supports fine-grained evaluation at both image and token levels. Through comprehensive experiments on representative cascaded (Qwen) and native (InternVL) open-source models, as well as closed-source systems (GPT-4o, Claude), we show that current VLMs fall significantly short of human performance, particularly in correction. To address these limitations, we explore two solution paradigms: a Joint OCR-Correction pipeline and a Background Information enhanced approach, both of which yield consistent performance gains. Our analysis highlights fundamental limitations of existing architectures and provides actionable insights for advancing multimodal spelling correction.

cs.CL

Electromagnetic Signal Modulation Recognition based on Subgraph Embedding Learning

Automatic Modulation Recognition (AMR) detects modulation schemes of received signals for further processing of signals without any priori information, which is critically important for civil spectrum regulation, information countermea sures, and communication security. Due to the powerful feature extraction and classification capabilities of Deep Learning (DL), DL-based AMR algorithms have achieved excellent performance gains compared with traditional modulation detection algorithms. However, all existing DL-based AMR algorithms, to the best of our knowledge, are designed for specific channels and systems, because data dimension of the used training dataset is fixed. To this end, we takes the first step to propose a Subgraph Embedding Learning (SEL) structure to address the classical AMR problem, and the proposed algorithm is called SEL-AMR. Our algorithm treats the communication system as a subgraph and uses the relationship between samples to smooth the effects brought by noise and different channels to extract robust features. Thus, the proposed SEL-AMR algorithm can adapt to any dynamic channels and systems. We use 5 public real datasets and a small amount of simulation data to evaluate our SEL-AMR algorithm. Experimental results reveal that SEL-AMR can well adapt to different channels and systems, and always outperforms the state of-the-art algorithms by improving up to 20% macro-average recognition precision and 30% recognition accuracy.

cs.NI

A Low-Cost, High-Precision Human-Machine Interaction Solution Based on Multi-Coil Wireless Charging Pads

Wireless charging pads are common, yet their functionality is mainly restricted to charging. Existing gesture recognition techniques, such as those based on machine vision and WiFi, have drawbacks like high costs and poor precision. This paper presents a new human machine interaction solution using multicoil wireless charging pads. The proposed approach leverages the pads existing modules without additional wearable sensors. It determines gestures by monitoring current and power changes in different coils. The data processing includes noise removal, sorting, highpass filtering, and slicing. A Bayesian network and particle filtering are employed for motion tracking. Through experiments, this solution proves to have wide applications, high recognition accuracy, and low cost. It can effectively identify diverse gestures, increasing the value of wireless charging pads. It outperforms traditional methods, with a 0.73 improvement in recognition accuracy and better environmental adaptability.

cs.HC

Multi-Failure Localization in High-Degree ROADM-based Optical Networks using Rules-Informed Neural Networks

To accommodate ever-growing traffic, network operators are actively deploying high-degree reconfigurable optical add/drop multiplexers (ROADMs) to build large-capacity optical networks. High-degree ROADM-based optical networks have multiple parallel fibers between ROADM nodes, requiring the adoption of ROADM nodes with a large number of inter-/intra-node components. However, this large number of inter-/intra-node optical components in high-degree ROADM networks increases the likelihood of multiple failures simultaneously, and calls for novel methods for accurate localization of multiple failed components. To the best of our knowledge, this is the first study investigating the problem of multi-failure localization for high-degree ROADM-based optical networks. To solve this problem, we first provide a description of the failures affecting both inter-/intra-node components, and we consider different deployments of optical power monitors (OPMs) to obtain information (i.e., optical power) to be used for automated multi-failure localization. Then, as our main and original contribution, we propose a novel method based on a rules-informed neural network (RINN) for multi-failure localization, which incorporates the benefits of both rules-based reasoning and artificial neural networks (ANN). Through extensive simulations and experimental demonstrations, we show that our proposed RINN algorithm can achieve up to around 20 higher localization accuracy compared to baseline algorithms, incurring only around 4.14 ms of average inference time.

cs.NI

Privacy-Preserving Gesture Tracking System Utilizing Frequency-Hopping RFID Signals

Gesture tracking technology provides users with a hands free interactive experience without the need to hold or touch devices. However, current gesture tracking research has primarily focused on tracking accuracy while neglecting issues of user privacy protection and security. This study aims to develop a gesture tracking system based on frequency hopping RFID signals that effectively protects user privacy without compromising tracking efficiency and accuracy. By introducing frequency hopping technology, we have designed a mechanism that prevents potential eavesdroppers from obtaining raw RFID signals, thereby enhancing the systems privacy protection capabilities. The system architec ture includes the collection of RFID signals, data processing, signal recovery, and gesture tracking. Experimental results show that our method significantly improves privacy protection levels while maintaining real time and accuracy. This research not only provides a new perspective for the field of gesture tracking but also offers valuable insights for the use of RFID technology in privacy-sensitive applications.

cs.CR

PERL: Pinyin Enhanced Rephrasing Language Model for Chinese ASR N-best Error Correction

Chinese ASR correction is challenging because errors are often \emph{phonetic} (many characters share similar Pinyin) while the correction model must also obey a \emph{length constraint} under noisy N-best hypotheses. Existing approaches either exploit Pinyin only at the prompt/feature level without integrating it into model representations or rely on generative decoding that can drift in length. We propose \textbf{PERL}, a \textbf{constrained rephrasing pipeline} for Chinese N-best ASR correction that (i) predicts the target length and enforces it via mask budgeting, and (ii) fuses \emph{semantic} and \emph{phonetic} (Pinyin) representations through token-wise gates conditioned on sentence semantics. Experiments on Aishell-1 and our new domain N-best benchmark \textbf{DoAD} show that PERL consistently reduces CER (29.11\% on Aishell-1 and up to $\sim$70\% on DoAD) while maintaining low latency. We also provide analyzes of length generalization and phonetic--semantic interactions, showing when PERL relies on phonetic cues versus semantic constraints.

cs.CL

Investigating the Star-Formation Characteristics of Radio Active Galactic Nuclei

The coevolution of supermassive black holes and their host galaxies represents a fundamental question in astrophysics. One approach to investigating this question involves comparing the star-formation rates (SFRs) of active galactic nuclei (AGNs) with those of typical star-forming galaxies. At relatively low redshifts ($z\lesssim 1$), radio AGNs manifest diminished SFRs, indicating suppressed star formation, but their behavior at higher redshifts is unclear. To examine this, we leveraged galaxy and radio AGN data from the well-characterized W-CDF-S, ELAIS-S1, and XMM-LSS fields. We established two mass-complete reference star-forming galaxy samples and two radio AGN samples, consisting of 1,763 and 6,766 radio AGNs, the former being higher in purity and the latter more complete. We subsequently computed star-forming fractions ($f_{\text{SF}}$; the fraction of star-forming galaxies to all galaxies) for galaxies and radio-AGN-host galaxies and conducted a robust comparison between them up to $z\approx3$. We found that the tendency for radio AGNs to reside in massive galaxies primarily accounts for their low $f_{\text{SF}}$, which also shows a strong negative dependence upon $M_{\star}$ and a strong positive evolution with $z$. To investigate further the star-formation characteristics of those star-forming radio AGNs, we constructed the star-forming main sequence (MS) and investigated the behavior of the position of AGNs relative to the MS at $z\approx0-3$. Our results reveal that radio AGNs display lower SFRs than star-forming galaxies in the low-$z$ and high-$M_{\star}$ regime and, conversely, exhibit comparable or higher SFRs than MS star-forming galaxies at higher redshifts or lower $M_{\star}$.

astro-ph.GA

Poster: Flexible Scheduling of Network and Computing Resources for Distributed AI Tasks

Many emerging Artificial Intelligence (AI) applications require on-demand provisioning of large-scale computing, which can only be enabled by leveraging distributed computing services interconnected through networking. To address such increasing demand for networking to serve AI tasks, we investigate new scheduling strategies to improve communication efficiency and test them on a programmable testbed. We also show relevant challenges and research directions.

cs.NI