SearcharxivSearch

arXiv subjects

Xiaoming Chen

Publications and source records attributed to Xiaoming Chen.

At least 19 recordsLinked to original sources

CircuTutor: Transforming Static Circuit Problems into Intelligent and Dynamic Tutoring

Learning direct current circuit concepts requires learners to connect invisible physical quantities, such as current, voltage, resistance, and power, with observable outcomes such as bulb brightness. Conventional textbook materials and general-purpose circuit simulators provide opportunities for problem solving and exploration but offer limited support for explaining why circuit behavior changes or diagnosing the reasoning behind incorrect answers. We present CircuTutor, a circuit-state-driven intelligent tutoring system that transforms static textbook circuit problems into an interactive tutoring workflow. CircuTutor first uses multimodal problem parsing to extract the textbook question, circuit topology, component parameters, switch states, and answer options, which are converted into a structured task and validated through circuit simulation. Learners can then interactively explore the circuit (by changing parameters) and submit an answer while a SPICE-compatible solver computes physically consistent circuit states. After the learner submits an answer, CircuTutor presents a before-and-after circuit state animation corresponding to the selected operation, organizes the simulated state changes into a causal reasoning chain that explains the underlying circuit behavior, maps answer discrepancies to likely misconceptions, and generates adaptive follow-up exercises targeted at the diagnosed misconception. Our experimental results demonstrate that CircuTutor effectively improves conceptual learning and the overall learning experience. The proposed framework demonstrates how simulated circuit states can be transformed into intelligent and interactive tutoring for circuit education, with the potential to generalize to other STEM domains.

cs.AI

AccidentSim: Generating Vehicle Collision Videos with Physically Realistic Collision Trajectories from Real-World Accident Reports

Collecting real-world vehicle accident videos for autonomous driving research is challenging due to their rarity and complexity. While existing driving video generation methods may produce visually realistic videos, they often fail to deliver physically realistic simulations because they lack the capability to generate accurate post-collision trajectories. In this paper, we introduce AccidentSim, a novel framework that generates physically realistic vehicle collision videos by extracting and utilizing the physical clues and contextual information available in real-world vehicle accident reports. Specifically, AccidentSim leverages a reliable physical simulator to replicate post-collision vehicle trajectories from the physical and contextual information in the accident reports and to build a vehicle collision trajectory dataset. This dataset is then used to fine-tune a language model, enabling it to respond to user prompts and predict physically consistent post-collision trajectories across various driving scenarios based on user descriptions. Finally, we employ Neural Radiance Fields (NeRF) to render high-quality backgrounds, merging them with the foreground vehicles that exhibit physically realistic trajectories to generate vehicle collision videos. Experimental results demonstrate that the videos produced by AccidentSim excel in both visual and physical authenticity.

cs.CV

VirSqueezer: Generating Realistic Deformations and Squeezing Dynamics in VR from Fine-Grained Squeezing Controls

Squeezing is one of the most natural forms of hand manipulation, inherently involving fine-grained, temporally evolving, per-finger flexion. In VR content creation, squeezing plays a unique role in enabling particular visual effects such as localized deformations and dynamic behaviors, e.g., bursting a Coke can or juicing a fruit, thereby expanding the expressive possibilities of VR content. However, existing techniques, such as 3D Gaussian splatting-based methods and diffusion-based video generation models, are limited in their ability to simulate fine-grained virtual squeezing effects. We introduce VirSqueezer, a framework designed to generate both localized deformations (primary effects) and complex squeezing dynamics, such as rupture and overflow (secondary effects). VirSqueezer captures squeezing control signals using a SenseGlove and provides the user with inferred resistance force feedback during the squeezing process. By estimating object contact areas, inferring physical properties, and simulating physical responses, VirSqueezer computes conditions that guide generation models for visual effect generation, ensuring both visual coherence and temporal synchronization with the simulation. Consequently, VirSqueezer enables the generation of physically realistic visual effects directly from continuous, fine-grained squeezing control signals. Our extensive evaluation demonstrates VirSqueezer's ability to reproduce realistic localized deformations, generate convincing visual dynamics, and maintain consistency in fine-grained squeezing controls.

cs.HC

FracEvent: Event-Camera Simulation via Fractional-Relaxation Pixel Dynamics

Event cameras asynchronously report brightness changes with microsecond-level temporal resolution, but real event data remain difficult to collect at scale because specialized sensors, careful synchronization, and task-specific annotations are required. Event-camera simulation is therefore important to event-based vision tasks. Most practical simulators build on contrast-threshold event generation, some with additional filtering, stochastic noise, or hand-tuned sensor parameters. While effective, such formulations often simplify the temporal structure produced by the lifecycle of each pixel, which can distort event timing and weaken downstream transfer. We introduce FracEvent, an event simulator that models this pixel-level lifecycle with fractional-relaxation voltage dynamics. Given a log-intensity trajectory, FracEvent drives a compact stack of relaxation modes, combines their responses into a voltage state, emits ON/OFF events by localizing threshold crossings on the continuous voltage trajectory, and updates the reference while retaining the underlying memory modes. This retained state links residual voltage response to later event timing. We evaluate FracEvent through event-stream comparison and downstream transfer on image reconstruction and optical flow estimation. Across multiple datasets, FracEvent improves the temporal structure of generated events and achieves stronger downstream-transfer results than competing simulator baselines, showing its practical value for event-camera simulation.

cs.CV

FlashAccel: Leveraging High-Bandwidth Flash (HBF) for High-Throughput LLM Inference

Large language model (LLM) inference is increasingly limited by the capacity of High-Bandwidth Memory (HBM) in GPUs, as model weights and KV cache grow rapidly. High-Bandwidth Flash (HBF) provides higher capacity than HBM while offering comparable bandwidth, making it a promising substrate for capacity-constrained LLM inference. However, its inherently high access latency, low bandwidth utilization, and lack of support for heterogeneous resource management make it difficult to integrate HBF into GPUs for LLM inference. We present FlashAccel, a co-designed system that enables efficient LLM inference using HBF. FlashAccel integrates HBF into HBM-based GPUs, providing architectural support to mitigate access latency. It improves bandwidth utilization through specialized data layouts for both model weights and KV cache, and introduces an HBF-aware storage management layer together with a programming model to organize persistent data in HBF and coordinate heterogeneous memory resources at the system level. Experimental results demonstrate that integrating six HBF stacks into the GPU enables FlashAccel to deliver an average improvement of 2.49$\times$ and 1.93$\times$ in throughput per GPU and energy efficiency over the HBM-only GPU under a 100ms latency constraint, respectively.

cs.AR

Recursive Flow: A Generative Framework for MIMO Channel Estimation

Channel estimation is a fundamental challenge in massive multiple-input multiple-output systems, where estimation accuracy governs the spectral efficiency and link reliability. In this work, we introduce Recursive Flow (RC-Flow), a novel solver that leverages pre-trained flow matching priors to robustly recover channel state information from noisy, under-determined measurements. Different from conventional open-loop generative models, our approach establishes a closed-loop refinement framework via a serial restart mechanism and anchored trajectory rectification. By synergizing flow-consistent prior directions with data-fidelity proximal projections, the proposed RC-Flow achieves robust channel reconstruction and delivers state-of-the-art performance across diverse noise levels, particularly in noise-dominated scenarios. The framework is further augmented by an adaptive dual-scheduling strategy, offering flexible management of the trade-off between convergence speed and reconstruction accuracy. Theoretically, we analyze the Jacobian spectral radius of the recursive operator to prove its global asymptotic stability. Numerical results demonstrate that RC-Flow reduces inference latency by two orders of magnitude while achieving a 2.7 dB performance gain in low signal-to-noise ratio regimes compared to the score-based baseline.

cs.IT

DynFOA: Generating First-Order Ambisonics with Conditional Diffusion for Dynamic and Acoustically Complex 360-Degree Videos

Spatial audio is crucial for immersive 360-degree video experiences, yet most 360-degree videos lack it due to the difficulty of capturing spatial audio during recording. Automatically generating spatial audio such as first-order ambisonics (FOA) from video therefore remains an important but challenging problem. In complex scenes, sound perception depends not only on sound source locations but also on scene geometry, materials, and dynamic interactions with the environment. However, existing approaches only rely on visual cues and fail to model dynamic sources and acoustic effects such as occlusion, reflections, and reverberation. To address these challenges, we propose DynFOA, a generative framework that synthesizes FOA from 360-degree videos by integrating dynamic scene reconstruction with conditional diffusion modeling. DynFOA analyzes the input video to detect and localize dynamic sound sources, estimate depth and semantics, and reconstruct scene geometry and materials using 3D Gaussian Splatting (3DGS). The reconstructed scene representation provides physically grounded features that capture acoustic interactions between sources, environment, and listener viewpoint. Conditioned on these features, a diffusion model generates spatial audio consistent with the scene dynamics and acoustic context. We introduce M2G-360, a dataset of 600 real-world clips divided into MoveSources, Multi-Source, and Geometry subsets for evaluating robustness under diverse conditions. Experiments show that DynFOA consistently outperforms existing methods in spatial accuracy, acoustic fidelity, distribution matching, and perceived immersive experience.

cs.SD

LivePhys: Transforming Static Physics Problems into Interactive Simulations via a Scan-to-Play Framework

Physics problems in textbooks are typically presented as static diagrams accompanied by brief textual descriptions, requiring learners to infer dynamic physical behaviors through mental visualization. This process often imposes high cognitive demands and limits learners' ability to form accurate mental models. In this paper, we present \textbf{LivePhys}, a framework that enables a \emph{Scan-to-Play} paradigm for mechanics learning by transforming static textbook physics problems into executable, interactive simulations. LivePhys decouples multimodal perception from physics-aware reasoning and deterministic simulation. Given a problem diagram and its accompanying text, LivePhys performs text extraction, geometric segmentation, and cross-modal grounding to construct a structured, physics-aware intermediate representation. A multimodal large language model is then used as a reasoning controller to infer entities, parameters, and constraints, which are executed by a physics engine to generate spatially consistent and interactive simulations that allow learners to explore and manipulate problem conditions dynamically. Our evaluation results show that LivePhys significantly outperforms general-purpose multimodal models in simulation executability, spatial accuracy, and interaction fidelity. In addition, a user study demonstrates that interacting with LivePhys-generated simulations reduces learners' perceived cognitive load compared to static textbook materials.

cs.ET

LEO Satellite Internet of Things: Architecture, Technology, and On-Orbit Verification

Low Earth orbit (LEO) satellite constellations are poised to become a cornerstone of the sixth-generation (6G) Internet of Things (IoT), providing truly global coverage and ubiquitous connectivity. This article presents a holistic two-dimensional system architecture for 6G LEO satellite IoT that incorporates composition and functional perspectives to facilitate the seamless integration of LEO satellites and terrestrial networks. Building upon this architecture, we evaluate three pivotal enabling technologies targeting the uplink, downlink, and inter-satellite links (ISLs). Specifically, we analyze massive grant-free random access for efficient uplink connectivity, investigate deep learning-based multibeam precoding for robust downlink transmission, and examine distributed cooperative routing for resilient ISL data delivery. Furthermore, we present an on-orbit verification platform that validates the real-world feasibility and performance of the proposed solutions. Finally, we outline key open challenges and future research directions to guide the realization of future LEO satellite IoT.

cs.IT

Task-Oriented Wave Processing with Stacked Intelligent Metasurfaces: Framework, Fusion, and Challenges

The deep integration of diverse services in sixth-generation (6G) networks poses significant challenges to conventional task-agnostic channels, often resulting in performance conflicts. To resolve these bottlenecks, this article introduces a physical-layer computing paradigm enabled by stacked intelligent metasurfaces (SIMs), transforming the wireless environment from a passive medium into a programmable signal processor. Specifically, we establish a unified framework to map high-level service requirements directly to wave-domain synthesis. We then investigate the fusion of diverse services, demonstrating how the deep computational architecture of SIMs resolves resource conflicts in integrated sensing and communication (ISAC) and integrated communication and computation (ICC) scenarios. Furthermore, we critically analyze fundamental challenges, including diffractive channel modeling and inverse task-to-phase mapping, while validating through numerical results that this approach elevates the system from simple coexistence to true service symbiosis. Finally, we discuss key research directions to pave the way for service-native 6G architectures.

cs.IT

Metasurface Antenna-Enabled LEO Satellite Constellation Communications: Design and Optimization

Next-generation low Earth orbit (LEO) satellite constellations face critical bottlenecks in spectral efficiency and onboard hardware complexity. To overcome these limitations, this paper introduces a novel architecture enabled by metasurface antennas (MAs) at the LEO satellites. In particular, MAs are metasurface-integrated feed antennas that perform high-precision beamforming directly in the wave domain, thereby effectively mitigating multi-user interference. Based on such an antenna architecture, a weighted sum rate (WSR) maximization problem is formulated by jointly optimizing the scheduling of feed antennas to terrestrial users (TUs) and the passive beamforming of the metasurface for system performance enhancement. To address this mixed-integer nonlinear programming (MINLP) challenge, an alternating optimization (AO)-based joint scheduling and beamforming algorithm is proposed. On the one hand, the proposed algorithm incorporates a polynomial-time minimum-cost maximum-flow (MCMF) method, which is dedicated to the optimal scheduling of feed antennas and TUs. On the other hand, it adopts a weighted minimum mean square error (WMMSE) method integrated with semidefinite relaxation (SDR) technique, which is tailored for metasurface beamforming design. Simulation results confirm the effectiveness of the proposed algorithm for MA-enabled LEO satellite constellation communications.

cs.IT

Robust Design of Integrated Sensing and Communication in LEO Satellite Systems

With the growing demand for satellite sensing and communication, the limited wireless resources are difficult to support multiple satellite systems. Therefore, it is desired to investigate integrated sensing and communication (ISAC) in low Earth orbit (LEO) satellite systems to enable multi-functionality within a single satellite, thereby saving both spectrum and orbital resources. In this paper, a framework for ISAC in LEO satellite systems is established, where a satellite can simultaneously sense multiple targets and serve multiple communication users (CUs) over the same spectrum. Considering the limited onboard energy of satellite, a novel robust beamforming design algorithm is developed with the goal of minimizing total transmit power while satisfying the mean squared error (MSE) requirements for sensing and signal-to-interference-plus-noise ratio (SINR) requirements for communication in presence of channel phase uncertainty which exacerbates the cross-functional interference. According to theoretical analysis, the proposed algorithm for ISAC in LEO satellite systems is effective. Moreover, extensive simulations confirm the superiority of the proposed algorithm over baselines.

cs.IT

On the Performance of Integrated Satellite-Terrestrial Maritime Communications

In this paper, we present an integrated terrestrial and satellite maritime communication system, where a shorebased terrestrial base station (TBS) and a low Earth orbit (LEO) satellite cooperatively provide wide-area communication services to maritime users. We conduct performance analysis for the integrated satellite-terrestrial maritime communication system. Specifically, we analyze the transmission rate and coverage probability of near-shore and off-shore users respectively according to the maritime communication environment. Besides, in order to better understand the impact of some key parameters, we also make asymptotic analysis in some special cases. Further, we design an optimization algorithm to maximize the coverage probability of near-shore users by adjusting the transmission power of TBS and the LEO satellite, while ensuring the both off-shore and near-shore users can meet the minimum communication rate requirements. Finally, extensive numerical analysis results verify the accuracy of the theoretical results and the effectiveness of the proposed optimization algorithm in the maritime communication system.

cs.IT

Joint Communication and Sensing Design for Integrated Satellite-Terrestrial Maritime Systems

Joint communication and sensing has been a key technology in 6G. By integrating sensing into maritime communications, ships can communicate with the base station while sensing the surrounding environment to ensure safe navigation. In this paper, we introduce an integrated satellite-terrestrial maritime system (ISTMS) with joint communication and sensing based on the same radio-frequency signals. Specifically, the terrestrial base station (TBS) and low Earth orbit (LEO) satellite provide communication services for near-shore users (NSUs) and off-shore users (OSUs), respectively, while simultaneously performing target sensing. Based on a differential evolution method (DE), we propose a sensing algorithm, which can enhance the location accuracy and reduce resource consumption. Furthermore, we derive the key performance metrics for both communication and sensing. Through joint beamforming optimization of the TBS and LEO satellite, we maximize the sum rate of maritime users while satisfying target localization accuracy requirements and transmit power constraints. Finally, extensive simulation results demonstrate the effectiveness of the proposed algorithms in terms of location accuracy and transmission rate compared with the baseline algorithms.

cs.IT

Parameter Efficient Machine Unlearning on Hybrid Resistive Memory based Compute-in-Memory Accelerators

Resistive memory compute-in-memory accelerators provide energy efficient analogue matrix vector multiplication for neural network inference, but frequent reprogramming of analogue weights remains costly because of device variability and iterative write and verify operations. This limitation hinders their use in edge model adaptation, including approximate machine unlearning and continual learning, where model parameters may need to be updated repeatedly in response to data deletion requests or newly arriving tasks. Here we present a co-design approach across hardware and software that maps frozen pretrained weights to analogue resistive memory arrays while placing trainable low rank adaptation branches in SRAM connected digital compute. By using LoRA style parameter efficient updates, the proposed scheme confines adaptation to a small set of digital parameters and avoids repeated reprogramming of the analogue backbone. To our knowledge, this work provides the first experimental demonstration of approximate machine unlearning on a fabricated resistive memory CIM accelerator. We validate the framework on a 180 nm 128x128 1T1R resistive-memory macro for face recognition, and through circuit-accurate simulations for speaker authentication and stylized image generation tasks, owing to the substantial model sizes involved. Compared with a baseline that directly updates analog weights, our hybrid mapping reduces analog training/update cost by up to 148x, on-chip deployment overhead by up to 388x, and inference energy by up to 59x, while preserving competitive task performance. These results show that hybrid analogue-digital LoRA mapping can enable efficient post-deployment adaptation on RM-CIM hardware, although formal machine-unlearning guarantees and large-scale system integration remain open challenges.

cs.ET

Continuous Aperture Array-Assisted Integrated Communication and Navigation in LEO Satellite Constellations

This paper proposes a novel continuous aperture array (CAPA)-assisted integrated communication and navigation (ICAN) framework for low Earth orbit (LEO) satellite constellations. Within this framework, an electromagnetic-based collaborative transmission model is developed, in which multiple satellites equipped with CAPAs simultaneously radiate downlink data streams and navigation reference signals over shared spectrum. Building upon this, the achievable communication rate and the navigation Cramer-Rao bound (CRB) are derived, which explicitly characterize the intrinsic coupling between the dual-function beamformers and system performance. To improve the positioning accuracy with communication quality of service guarantee, a joint beamforming optimization problem is formulated to minimize the average CRB subject to transmit power budgets and minimum rate constraints. To tackle the inherent infinite-dimensionality of the CAPA beamformer design, an ICAN channel subspace is introduced to equivalently transform the formulation into a tractable finite-dimensional problem, which is then efficiently solved via an iterative convex optimization algorithm. Finally, numerical results demonstrate that the proposed CAPA-assisted beamforming design algorithm significantly outperforms conventional discrete phased array architectures and other benchmark schemes, yielding notable improvements in ICAN performance.

cs.IT

Modeling and Analysis for Multiple-Layer LEO Satellite Internet of Things Constellations

To provide multiple-satellite coverage for global Internet of Things (IoT), a low Earth orbit (LEO) satellite IoT constellation usually contains multiple-layer orbits with different altitudes. However, the performance of multiple-layer LEO satellite IoT constellations under practical Rician fading satellite channels remains unknown due to complex theoretical modeling and intractable mathematical analysis. To address these challenges, this paper proposes a stochastic geometry-based modeling and analysis framework for multiple-layer LEO satellite IoT constellations, integrating Rician channel modeling and Cox point processes. Specifically, we introduce a novel channel approximation method to overcome the intractable expressions caused by the Rician fading. Building on this method, we derive exact closed-form expressions for key performance metrics, including connectivity probability, coverage probability, and transmission rate, especially in the case of IoT short-packet transmission. Extensive simulation results validate the accuracy and effectiveness of the proposed model and reveal significant design insights. The results not only provide new theoretical perspectives for modeling and analysis of LEO satellite IoT constellations but also offer practical guidance for system deployment and optimization.

cs.IT

EscFOA: Enhancing Spatial Learning for Visually Impaired Learners via Generative Spatial Audio in 360-Degree Educational Environments

Immersive 360-degree educational environments often lack accessible spatial structure, limiting visually impaired learners' ability to orient, explore, and construct mental representations. This paper proposes EscFOA, a geometry-aware spatial audio generation framework designed as an \emph{acoustic scaffolding} to support spatial cognition. By integrating 3D Gaussian Splatting (3DGS) with conditional diffusion models, EscFOA reconstructs scene geometry from 360-degree videos to synthesize high-fidelity spatial audio consistent with the environmental structure. Explicitly targeting learning outcomes like independent spatial orientation and reduced cognitive load, EscFOA significantly outperforms conventional monaural and stereo audio in supporting spatial learning behaviors among blindfolded sighted participants (simulating visually impaired learners). These findings demonstrate that geometry-consistent generative audio can effectively enable inclusive access to complex spatial learning materials.

cs.SD