SearcharxivSearch

arXiv subjects

Zibo Zhang

Publications and source records attributed to Zibo Zhang.

16 recordsLinked to original sources

Negative correlation for the random-cluster model below one on high-degree regular graphs

The random-cluster model with cluster parameter \(0<Q<1\) is conjectured to exhibit negative dependence, but even pairwise negative correlation between distinct edges remains open on general finite graphs. We first show that restricting the problem to regular graphs of diverging degree does not essentially weaken the pairwise conjecture: validity for all such graph sequences, even when restricted to edge pairs at distance \(o(d/\log d)\), is equivalent to validity on arbitrary finite graphs. We then prove strict pairwise negative correlation for a broad class of high-degree regular graph sequences satisfying a two-scale edge-isoperimetric condition. More precisely, for every fixed \(p\in(0,1)\), all sufficiently large graphs in the sequence have strictly negative covariance between any two distinct edge indicators at distance \(o(d/\log d)\), uniformly in \(Q\in[0,1)\) and in the choice of the two edges. In particular, the result includes the \(Q=0\) endpoint, corresponding to Bernoulli bond percolation conditioned to be connected. The isoperimetric hypothesis is satisfied by high-degree expander families and, with probability tending to one, by uniformly random regular graphs of diverging degree; its two-scale form also allows product geometries such as hypercubes and fixed-side high-dimensional discrete tori. The proof is based on a polymer representation and cluster expansion, together with a geometric identification of the leading contribution to the two-edge correlation.

math.PR

ReferTrack: Referring Then Tracking for Embodied Visual Tracking

Embodied visual tracking (EVT) requires a mobile agent to continuously follow a specific target described in natural language using only onboard vision. While recent vision-language-action (VLA) policies unify target identification and trajectory planning, their chain-of-thought (CoT) reasoning often operates in abstract spatial latents that are difficult to supervise and weakly aligned with explicit image-space detections. To address this, we introduce ReferTrack, a referring-then-tracking paradigm that grounds EVT using a single forward-facing camera. Our model first selects the target from an indexed set of bounding boxes, then decodes tracking waypoints conditioned on this image-grounded decision. To preserve target motion cues over time, ReferTrack maintains a sliding-window queue of previously selected bounding boxes, injecting their geometric features into the visual history via temporal-viewpoint-bbox indicator (TVBI) tokens. We further enhance target identification by co-training on a custom Refer-QA dataset. On EVT-Bench, ReferTrack achieves state-of-the-art single-view performance with success rates of 89.4%, 73.3%, and 74.1% on the single-target, distracted, and ambiguity tracking splits, respectively -- matching or even surpassing several multi-camera baselines on identification-heavy tasks. Finally, real-world deployments on legged and humanoid robots validate its robust sim-to-real transfer capabilities. Code is available at https://github.com/MedlarTea/referTrack.

cs.RO

From second moments to pairwise negative correlation: applications to minimal and uniform spanning trees

We uncover a close connection between the second moment of the degree of a typical vertex in a random subgraph and the pairwise negative correlation (p-NC) property. On one hand, we exploit this connection to prove the p-NC property for non-adjacent edges in minimal spanning trees on complete graphs. On the other hand, we apply the classical p-NC property of uniform spanning trees to derive a universal upper bound on the second moment of the degree of a uniformly chosen vertex in uniform spanning trees on finite, connected, regular graphs, thereby resolving an open question posed by Nachmias and Peres. Furthermore, we determine that the optimal upper bound is exactly 6, and the method for achieving this optimal bound is interesting in itself -- the proof uses Edmonds' matroid polytope theorem.

math.PR

Pairwise Negative Correlation for Uniform Spanning Subgraphs of the Complete Graph

We investigate the pairwise negative correlation (p-NC) property for uniform probability measures on several families of spanning subgraphs of the complete graph $K_n$. Motivated by conjectured negative dependence properties of the random-cluster model with $q<1$, we focus on three natural families: the set of all connected spanning subgraphs, the set of forests with exactly $k$ components, and the set of connected spanning subgraphs with excess $k$, where $k$ is a fixed integer. We prove that for each of these families, the associated uniform measure satisfies the p-NC property provided $n$ is sufficiently large. Our results extend earlier work on uniform forests and provide the first verification of the p-NC property for uniform connected subgraphs and their truncations on complete graphs.

math.PR

Biomimetic Metamaterial-based Interface for Decoding Heterogeneous Mechanodermal Activity

Human skin acts as a dynamic biomechanical interface that conveys critical physiological and behavioural information through spatiotemporally distributed deformations. Due to the limited capabilities of current sensing technologies, the spatiotemporal diversity of its mechanical cues has remained underutilised to date, preventing these mechanisms from being used to capture and decode the full spectrum of underlying physiological states. In this work, we define this heterogeneous set of mechanical signals as mechanodermal activity (MDA) and introduce the biomimetic metamaterial-based interface (BMMI), an engineered auxetic metamaterial substrate that reproduces the microrelief and mechanoreceptor architecture of natural skin. The BMMI allows selective capture of diverse MDA signals from adjacent skin regions with simultaneous signal amplification and noise suppression, and permits straightforward modulation to accommodate various scenarios. Combined with bespoke algorithms, the wireless BMMI device decodes MDA accurately and robustly for multimodal communication interfaces, unleashing applications in healthcare monitoring and human-machine interaction.

q-bio.NC

Physiology-informed layered sensing for intelligent human-exoskeleton interaction

Wearable exoskeletons hold transformative promise for restoring mobility across diverse users with muscular weakness or other impairments. However, their translation beyond laboratory environments remains limited by sensing systems that capture movement but not underlying physiology. Here, we present a soft, lightweight smart leg sleeve that achieves anatomically aligned, layered multimodal sensing by integrating textile-based surface electromyography (sEMG) electrodes, ultrasensitive textile strain sensors, and inertial measurement units (IMUs). Each sensing modality targets a distinct physiological layer: IMUs track joint kinematics at the skeletal level, sEMG monitors muscle activation at the muscular level, and strain sensors detect skin deformation at the cutaneous level. Together, these sensors provide real-time perception to support three core objectives: controlling personalized assistance, optimizing user effort, and safeguarding against injury risks. The system is skin-conformal, mechanically compliant, and seamlessly integrated with a custom exoskeleton ($<20$~g total sensor and electronics weight). We demonstrate: (1) accurate ankle joint moment estimation (RMSE = 0.13~Nm/kg), (2) real-time classification of metabolic trends (accuracy = 97.1\%), and (3) injury risk detection within 100~ms (recall = 0.96), all validated on unseen users using a leave-one-subject-out protocol. This work establishes a physiology-aligned sensing architecture that reframes exoskeleton perception from motion tracking to real-time physiological decoding, offering a pathway towards intelligent, adaptive, and personalized wearable robotics.

eess.SY

AI-Driven Smart Sportswear for Real-Time Fitness Monitoring Using Textile Strain Sensors

Wearable biosensors have revolutionized human performance monitoring by enabling real-time assessment of physiological and biomechanical parameters. However, existing solutions lack the ability to simultaneously capture breath-force coordination and muscle activation symmetry in a seamless and non-invasive manner, limiting their applicability in strength training and rehabilitation. This work presents a wearable smart sportswear system that integrates screen-printed graphene-based strain sensors with compact electronics for wireless data transfer and a deep learning framework for real-time classification of exercise execution quality. By leveraging 1D ResNet-18 for feature extraction, the system achieves 92.1% classification accuracy across six exercise conditions, distinguishing between breathing irregularities and asymmetric muscle exertion. Additionally, t-SNE analysis and Grad-CAM-based explainability visualization confirm that the network accurately captures biomechanically relevant features, ensuring robust interpretability. The proposed system establishes a foundation for next-generation AI-powered sportswear, with applications in fitness optimization, injury prevention, and adaptive rehabilitation training.

eess.SP

How to Make Your Multi-Image Posts Popular? An Approach to Enhanced Grid for Nine Images on Social Media

The nine-grid layout is commonly used for multi-image posts, arranging nine images in a tic-tac-toe board. This layout effectively presents content within limited space. Moreover, due to the numerous possible arrangements within the nine-image grid, the optimal arrangement that yields the highest level of attractiveness remains unknown. Our study investigates how the arrangement of images within a nine-grid layout affects the overall popularity of the image set, aiming to explore alignment schemes more aligned with user preferences. Based on survey results regarding user preferences in image arrangement, we have identified two ordering sequences that are widely recognized: sequential order and center prioritization, considering both image visual content and aesthetic quality as alignment metrics, resulting in four layout schemes. Finally, we recruited participants to annotate various layout schemes of the same set of images. Our experience-centered evaluation indicates that layout schemes based on aesthetic quality outperformed others. This research yields empirical evidence supporting the optimization of the nine-grid layout for multi-image posts, thereby furnishing content creators with valuable insights to enhance both attractiveness and user experience.

cs.HC

An AI-Driven Multimodal Smart Home Platform for Continuous Monitoring and Assistance in Post-Stroke Motor Impairment

At-home rehabilitation for post-stroke patients presents significant challenges, as continuous, personalized care is often limited outside clinical settings. Moreover, the lack of integrated solutions capable of simultaneously monitoring motor recovery and providing intelligent assistance in home environments hampers rehabilitation outcomes. Here, we present a multimodal smart home platform designed for continuous, at-home rehabilitation of post-stroke patients, integrating wearable sensing, ambient monitoring, and adaptive automation. A plantar pressure insole equipped with a machine learning pipeline classifies users into motor recovery stages with up to 94\% accuracy, enabling quantitative tracking of walking patterns during daily activities. An optional head-mounted eye-tracking module, together with ambient sensors such as cameras and microphones, supports seamless hands-free control of household devices with a 100\% success rate and sub-second response time. These data streams are fused locally via a hierarchical Internet of Things (IoT) architecture, ensuring low latency and data privacy. An embedded large language model (LLM) agent, Auto-Care, continuously interprets multimodal data to provide real-time interventions -- issuing personalized reminders, adjusting environmental conditions, and notifying caregivers. Implemented in a post-stroke context, this integrated smart home platform increased mean user satisfaction from 3.9 $\pm$ 0.8 in conventional home environments to 8.4 $\pm$ 0.6 with the full system ($n=20$). Beyond stroke, the system offers a scalable, patient-centered framework with potential for long-term use in broader neurorehabilitation and aging-in-place applications.

cs.HC

Wearable intelligent throat enables natural speech in stroke patients with dysarthria

Wearable silent speech systems hold significant potential for restoring communication in patients with speech impairments. However, seamless, coherent speech remains elusive, and clinical efficacy is still unproven. Here, we present an AI-driven intelligent throat (IT) system that integrates throat muscle vibrations and carotid pulse signal sensors with large language model (LLM) processing to enable fluent, emotionally expressive communication. The system utilizes ultrasensitive textile strain sensors to capture high-quality signals from the neck area and supports token-level processing for real-time, continuous speech decoding, enabling seamless, delay-free communication. In tests with five stroke patients with dysarthria, IT's LLM agents intelligently corrected token errors and enriched sentence-level emotional and logical coherence, achieving low error rates (4.2% word error rate, 2.9% sentence error rate) and a 55% increase in user satisfaction. This work establishes a portable, intuitive communication platform for patients with dysarthria with the potential to be applied broadly across different neurological conditions and in multi-language support systems.

eess.AS

Fiber-level Woven Fabric Capture from a Single Photo

Accurately rendering the appearance of fabrics is challenging, due to their complex 3D microstructures and specialized optical properties. If we model the geometry and optics of fabrics down to the fiber level, we can achieve unprecedented rendering realism, but this raises the difficulty of authoring or capturing the fiber-level assets. Existing approaches can obtain fiber-level geometry with special devices (e.g., CT) or complex hand-designed procedural pipelines (manually tweaking a set of parameters). In this paper, we propose a unified framework to capture fiber-level geometry and appearance of woven fabrics using a single low-cost microscope image. We first use a simple neural network to predict initial parameters of our geometric and appearance models. From this starting point, we further optimize the parameters of procedural fiber geometry and an approximated shading model via differentiable rasterization to match the microscope photo more accurately. Finally, we refine the fiber appearance parameters via differentiable path tracing, converging to accurate fiber optical parameters, which are suitable for physically-based light simulations to produce high-quality rendered results. We believe that our method is the first to utilize differentiable rendering at the microscopic level, supporting physically-based scattering from explicit fiber assemblies. Our fabric parameter estimation achieves high-quality re-rendering of measured woven fabric samples in both distant and close-up views. These results can further be used for efficient rendering or converted to downstream representations. We also propose a patch-space fiber geometry procedural generation and a two-scale path tracing framework for efficient rendering of fabric scenes.

cs.GR

HABD: a houma alliance book ancient handwritten character recognition database

The Houma Alliance Book, one of history's earliest calligraphic examples, was unearthed in the 1970s. These artifacts were meticulously organized, reproduced, and copied by the Shanxi Provincial Institute of Cultural Relics. However, because of their ancient origins and severe ink erosion, identifying characters in the Houma Alliance Book is challenging, necessitating the use of digital technology. In this paper, we propose a new ancient handwritten character recognition database for the Houma alliance book, along with a novel benchmark based on deep learning architectures. More specifically, a collection of 26,732 characters samples from the Houma Alliance Book were gathered, encompassing 327 different types of ancient characters through iterative annotation. Furthermore, benchmark algorithms were proposed by combining four deep neural network classifiers with two data augmentation methods. This research provides valuable resources and technical support for further studies on the Houma Alliance Book and other ancient characters. This contributes to our understanding of ancient culture and history, as well as the preservation and inheritance of humanity's cultural heritage.

cs.CV

A deep learning-enabled smart garment for accurate and versatile sleep conditions monitoring in daily life

In wearable smart systems, continuous monitoring and accurate classification of different sleep-related conditions are critical for enhancing sleep quality and preventing sleep-related chronic conditions. However, the requirements for device-skin coupling quality in electrophysiological sleep monitoring systems hinder the comfort and reliability of night wearing. Here, we report a washable, skin-compatible smart garment sleep monitoring system that captures local skin strain signals under weak device-skin coupling conditions without positioning or skin preparation requirements. A printed textile-based strain sensor array responds to strain from 0.1% to 10% with a gauge factor as high as 100 and shows independence to extrinsic motion artefacts via strain-isolating printed pattern design. Through reversible starching treatment, ink penetration depth during direct printing on garments is controlled to achieve batch-to-batch performance variation < 10%. Coupled with deep learning, explainable artificial intelligence (XAI), and transfer learning data processing, the smart garment is capable of classifying six sleep states with an accuracy of 98.6%, maintaining excellent explainability (classification with low bias) and generalization (95% accuracy on new users with few-shot learning less than 15 samples per class) in practical applications, paving the way for next-generation daily sleep healthcare management.

eess.SP

Facilitating Mixed-Methods Analysis with Computational Notebooks

Data exploration is an important aspect of the workflow of mixed-methods researchers, who conduct both qualitative and quantitative analysis. However, there currently exists few tools that adequately support both types of analysis simultaneously, forcing researchers to context-switch between different tools and increasing their mental burden when integrating the results. To address this gap, we propose a unified environment that facilitates mixed-methods analysis in a computational notebook-based settings. We conduct a scenario study with three HCI mixed-methods researchers to gather feedback on our design concept and to understand our users' needs and requirements.

cs.HC

EmoWear: Exploring Emotional Teasers for Voice Message Interaction on Smartwatches

Voice messages, by nature, prevent users from gauging the emotional tone without fully diving into the audio content. This hinders the shared emotional experience at the pre-retrieval stage. Research scarcely explored "Emotional Teasers"-pre-retrieval cues offering a glimpse into an awaiting message's emotional tone without disclosing its content. We introduce EmoWear, a smartwatch voice messaging system enabling users to apply 30 animation teasers on message bubbles to reflect emotions. EmoWear eases senders' choice by prioritizing emotions based on semantic and acoustic processing. EmoWear was evaluated in comparison with a mirroring system using color-coded message bubbles as emotional cues (N=24). Results showed EmoWear significantly enhanced emotional communication experience in both receiving and sending messages. The animated teasers were considered intuitive and valued for diverse expressions. Desirable interaction qualities and practical implications are distilled for future design. We thereby contribute both a novel system and empirical knowledge concerning emotional teasers for voice messaging.

cs.HC

Ultrasensitive Textile Strain Sensors Redefine Wearable Silent Speech Interfaces with High Machine Learning Efficiency

Our research presents a wearable Silent Speech Interface (SSI) technology that excels in device comfort, time-energy efficiency, and speech decoding accuracy for real-world use. We developed a biocompatible, durable textile choker with an embedded graphene-based strain sensor, capable of accurately detecting subtle throat movements. This sensor, surpassing other strain sensors in sensitivity by 420%, simplifies signal processing compared to traditional voice recognition methods. Our system uses a computationally efficient neural network, specifically a one-dimensional convolutional neural network with residual structures, to decode speech signals. This network is energy and time-efficient, reducing computational load by 90% while achieving 95.25% accuracy for a 20-word lexicon and swiftly adapting to new users and words with minimal samples. This innovation demonstrates a practical, sensitive, and precise wearable SSI suitable for daily communication applications.

eess.AS