SearcharxivSearch

arXiv subjects

Yuxiang Zhang

Publications and source records attributed to Yuxiang Zhang.

At least 19 recordsLinked to original sources

Reassessing 3GPP NR CSI Codebook Structures in Near-Field Channels: Finite-Feedback Multilayer Precoding and Design Insights

3GPP TR 38.901 Rel-19 introduces antenna-element-level spherical-wave modeling, while NR Type-I and enhanced Type-II (eType-II) CSI codebooks continue to use plane-wave DFT beams. Whether this mismatch materially degrades finite-feedback multilayer precoding in standardized multipath channels remains unclear. To isolate its impact, we evaluate both codebooks over strictly paired far-field (FF) and near-field (NF) 3GPP channels that share user locations, multipath parameters, polarization, and link budgets and differ only in their wavefront models. Simulations cover Rank 1-4 transmission in 7-GHz UMi and 24-GHz InH-linear scenarios. We find no systematic FF/NF shift in singular-mode gains or equal-power SVD (SVD-EP) rates. Instead, spherical-wave phases reorder multipath projections onto plane-wave candidates and thereby alter codeword selection. For Rank-4 InH-linear users at 0.1 normalized Rayleigh distance, given FF-selected Type-I and eType-II codewords incur median direct mismatch losses of 2.28% and 5.34% on the NF channel, respectively; codebook reselection identifies better-matched codewords and reduces these losses to 0.770% and 3.28%. Nested candidate-set comparisons further show that relaxing Type-I interlayer constraints improves the SVD-EP-normalized rate by 20.7 percentage points, whereas finite-range sampling adds only 0.579 points. These results support prioritizing multilayer multibeam representation in large-aperture NR CSI codebooks, with range states providing complementary refinement.

eess.SP

PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models

We present PhysBrain 1.5, a unified model for understanding physical environments, generating actions, and predicting future states. Motivated by the physical loop of observation, interaction, and environmental change, we bring these capabilities into a common learning framework. Starting from a general vision--language model, we encode language responses, end-effector motion, and dense visual targets as discrete sequences and jointly optimize them with autoregressive next-token prediction. Pre-training draws its embodied supervision entirely from human interaction videos, using task-centered episodes to pair semantic and spatial context with recovered motion and subsequent observations. We then adapt the model through supervised fine-tuning on a mixture of human demonstrations, robot trajectories, and simulated experience. Across 28 embodied understanding benchmarks, our 8B model achieves an average score of 72.5, setting a new open-source state of the art and performing on par with leading proprietary models such as GPT-6-Astra and Gemini 3.6 Flash. It achieves the best open-source results on 14 benchmarks while retaining general multimodal capabilities. Beyond these understanding evaluations, qualitative examples show the model's ability to produce end-effector trajectories and predict future scenes through spatially aligned RGB, depth, and robot-mask outputs.

cs.CV

iOSPointMapper: RealTime Pedestrian and Accessibility Mapping with Mobile AI

Accurate, up-to-date sidewalk data is essential for building accessible and inclusive pedestrian infrastructure, yet current approaches to data collection are often costly, fragmented, and difficult to scale. We introduce iOSPointMapper, a mobile application that enables real-time, privacy-conscious sidewalk mapping on the ground, using recent-generation iPhones and iPads. The system leverages on-device semantic segmentation, LiDAR-based depth estimation, and fused GPS/IMU data to detect and localize sidewalk-relevant features such as traffic signs, traffic lights and poles. To ensure transparency and improve data quality, iOSPointMapper incorporates a user-guided annotation interface for validating system outputs before submission. Collected data is anonymized and transmitted to the Transportation Data Exchange Initiative (TDEI), where it integrates seamlessly with broader multimodal transportation datasets. Detailed evaluations of the system's feature detection and spatial mapping performance reveal the application's potential for enhanced pedestrian mapping. Together, these capabilities offer a scalable and user-centered approach to closing critical data gaps in pedestrian

cs.CV

Band-Selective LDOS Engineering of Yb/Er Upconversion: an Electromagnetic-Kinetic Diagnostic Framework

A persistent challenge in plasmonic upconversion is decoupling pump-field enhancement from emission-side local-density-of-optical-states (LDOS) engineering to achieve selective band manipulation. Here, we show that a corrugated SU8/Au/Al2O3 grating coated with a NaYF4:Yb/Er upconversion nanoparticle (UCNP) monolayer realizes a truly band-selective platform. A broad plasmonic resonance near 670 nm modulates the red Er 3 + 3+ decay rate by $\pm$15% as a function of the Al2O3 spacer thickness, while leaving the green transition experimentally invariant (< 1% change). Simultaneously, the 980 nm pump field is monotonically suppressed below free-space levels, ensuring that steady-state and time-resolved observables cleanly probe the emission-side LDOS without pump interference. We analyze this system using a coupled electromagnetic-kinetic framework that integrates finite-difference time-domain (FDTD) calculations of Purcell factors and pump fields with a six-level Yb/Er rate-equation model. The framework quantitatively reproduces the 670 nm plasmonic resonance, the red-band decay-rate modulation, and the monotonic decrease of the green/red intensity ratio. Crucially, the model serves as a powerful diagnostic tool: it overpredicts a green-band rate reduction, but systematic parametric testing rules out geometric (apex smoothing) and material (grain-boundary damping, interband loss) imperfections as the cause. Instead, it isolates the residual discrepancy to measurement-versus-model factors (finite-aperture angular averaging) and missing non-radiative kinetic channels, establishing a clear roadmap for the rational design and validation of future plasmonic-UCNP architectures.

physics.optics

ChainPrune: Evaluating and Reducing Redundancy in Long Chain-of-Thought Reasoning

Chain-of-Thought (CoT) reasoning has significantly enhanced the multi-step problem-solving capabilities of large language models (LLMs) by introducing explicit intermediate reasoning. However, advanced Large Reasoning Models (LRMs) often exhibit overthinking behaviors, including excessively long reasoning steps, redundant steps, and high computational overhead. Existing token-length reward strategies aim to promote concise outputs, but often result in pseudo-conciseness, where token count is reduced, yet redundant reasoning persists, leading to longer and less structurally efficient chains. To address these limitations, we propose ChainPrune, a novel reasoning path semantic structural optimization method to efficiently and controllably synthesize self-generated high-quality training data. We initially consolidate self-generated reasoning paths into a tree-based structure, followed by a multi-criteria dominant path selection process for preference data construction that formulates shallow reasoning trajectories while preserving essential reasoning steps. To further enhance the quality of reasoning, we incorporate a DPO-based preference learning method combined with supervised loss, effectively mitigating false reward suppression. This innovative integration significantly enhances both the efficiency and effectiveness of our reasoning framework. Comprehensive experimental results demonstrate significant reductions in step length and computational overhead, while maintaining or even enhancing accuracy.

cs.LG

Electromagnetic World Model for 6G: A Unified Framework for Joint Environment Reconstruction and Channel Prediction

The integration of sensing, communication, and intelligence is becoming a key enabler for sixth generation (6G) wireless systems, where intelligent terminals are expected to simultaneously support efficient link establishment and reliable environmental sensing. However, existing studies mainly exploit sensing information or communication information to address a single task, such as channel prediction or environment reconstruction. Motivated by the shared dependence of optical and radio-frequency signals on the surrounding environment, we propose the electromagnetic world model (EMWM), the first unified framework for joint environment reconstruction and channel prediction. EMWM learns a common electromagnetic representation with the potential to provide a modeling foundation for 6G tasks. Specifically, partial channel state information (CSI) and multi-view red-green-blue (RGB) images are encoded into CSI and visual tokens and jointly processed by a hierarchical world-model backbone with local and global aggregation. Based on the learned representation, a mixture-of-experts (MoE)-based CSI prediction head reconstructs the complete CSI, while a depth prediction head estimates multi-view depth maps that are further converted into three-dimensional (3D) point clouds. Moreover, a large-scale multi-modal dataset is constructed based on a campus digital twin. Experimental results show that EMWM outperforms conventional neural network and large language model (LLM) baselines in both CSI prediction and environment reconstruction, achieving a squared generalized cosine similarity (SGCS) of 0.9699 for CSI prediction while demonstrating robustness across different signal-to-noise ratio (SNR) conditions and zero-shot generalization at 28 GHz.

eess.SP

WiWorld-RealData: A Real-World Multi-Modal Dataset for 6G Wireless World Models

As sixth-generation wireless systems evolve from reliable connectivity toward environment intelligence, wireless world models aim to learn how physical environments and user states affect wireless propagation, requiring real-world data with explicit correspondences between channel responses and environment observations. However, existing channel-environment datasets are predominantly simulation-based or designed for specific communication tasks, limiting their support for general environment-channel relationship learning. To address this gap, we construct WiWorld-RealData, a real-world multi-band channel and multi-modal environment sensing dataset for 6G wireless world model research. It provides synchronized channel impulse responses measured at 3.7 and 6.775 GHz together with multi-view and panoramic images, light detection and ranging point clouds, millimeter-wave radar observations, and global navigation satellite system trajectories. Unified timestamps, sample identifiers, and metadata establish sample-level correspondences across these heterogeneous modalities. The overall measurement campaign produced approximately 10 TB of data, while the current public release provides aligned channel-environment samples from a representative continuous outdoor route. A path-loss prediction case study further validates the dataset using a continuous test route segment, achieving a mean absolute error of 2.02 dB and a root mean square error of 2.69 dB under few-shot adaptation. WiWorld-RealData supports cross-band propagation analysis, environment-aware channel modeling, wireless digital twins, and channel foundation model research. The dataset is available at https://scc.bupt.edu.cn/dataset-manage/datasets/44 and https://doi.org/10.57760/sciencedb.40663.

eess.SP

SSP: An Event-Matched Syn2Sim2Phy Cross-Domain Evaluation Framework for Autonomous Driving VLA Models

Vision-language-action (VLA) models for autonomous driving jointly produce scene interpretation, language-based reasoning, and driving trajectories. Existing evaluations often use independently selected synthetic, simulated, and physical data, so measured performance gaps can be confounded by changes in scenario content rather than genuine domain sensitivity. We propose SSP (Synthetic-Simulation-Physical), an event-matched Syn2Sim2Phy evaluation framework that anchors cross-domain comparison to the same safety-critical interaction. Starting from a synthetic long-tail video, SSP builds a validated event specification that preserves road topology, participant roles, relative motion, conflict evolution, passing order, response constraints, and event phases. Platform-specific realizations are then constructed in CARLA and on a closed proving ground and are evaluated only after transfer audits confirm preservation of mandatory event properties. SSP maps heterogeneous outputs from OpenEMMA, LLaViDA, and Alpamayo-R1 into common semantic slots and a 1 s trajectory window to assess output validity, semantic accuracy, critical-interaction recognition, trajectory quality, and risk response. Across Cut-in and vulnerable-road-user crossing cases, the macro-averaged Integrated VLA Capability Scores are 0.259, 0.291, and 0.325 in the Synthetic, Simulation, and Physical domains, respectively, while the best domain varies by scenario. Alpamayo-R1, OpenEMMA, and LLaViDA obtain scores of 0.405, 0.338, and 0.131. SSP provides a reproducible scene-transfer chain and an evidence-qualified evaluation of VLA behavior without assuming that the Physical domain is universally superior.

cs.CV

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs

Tool-use LLMs are increasingly asked to act on users' behalf, but existing benchmarks usually focus on profile recall, style imitation, generic tool use, or response-level personalization. We introduce UserToolBench , a benchmark for personalized decision making in tool-use LLMs. UserToolBench tests whether a model can infer latent user preferences from interaction history, recognize when clarification is needed, and produce user-aligned tool-call trajectories under incomplete information. The benchmark is built from privacy-sanitized real interaction traces and combines structured persona profiles, public API-style tool ecosystems, and long-horizon multi-turn trajectories. It includes 10 user profiles, 36 tool sets, 1,065 turns, 170 unique tools, and evaluation-focused task types covering lack-of-information, single-tool, and multi-tool settings. Experiments with strong tool-use LLMs show that current models still have difficulty with personalized delegation. Multi-tool coordination, missing-constraint inference, and long-horizon behavioral consistency remain major bottlenecks. These results suggest that personalization evaluation should move beyond asking whether outputs sound user-specific and instead ask whether LLMs make correct decisions for the users they represent.

cs.LG

ZeroLock: Concurrent Memory-Efficient LLM Training via Modular Update Decoupling

Large language model (LLM) fine-tuning at the edge adapts the model to scenario-specific data while preserving privacy. Although existing studies proposed pipeline parallelism to address the limited memory and computing resources of edge devices, they commonly rely on backpropagation (BP) training, which has a fundamental limitation of update locking and could experience severe throughput and memory bottlenecks. In this work, we propose a BP-free algorithm, called ZeroLock, that decouples the model updates into independent chunk updates by local objective construction. It breaks the update locking of BP and hence can improve throughput at the algorithm level and lower memory usage by reducing activation storage. To the best of our knowledge, we provide the first theoretical framework for such local objective construction-based approaches under general model chunk division by mapping local objectives to the global objective. We prove that ZeroLock has a convergence rate of $\tilde{\mathcal{O}}(1/\sqrt{T})$, which differs from BP only by polylogarithmic factors. We design a system for ZeroLock and build real-world prototypes, incorporating techniques such as early forwarding and failure recovery for efficient and robust implementation. Experiments on the prototype show that compared to BP-based baselines, ZeroLock reduces the memory by 26.5% and improves throughput by 4.9%.

cs.LG

MagicSelector: Joint Optimization for Agent Tool Selection via Counterfactual Decomposition and Progressive Reranking

We present MagicSelector, a joint optimization framework integrating Counterfactual task decomposition, Progressive reranking, and Dynamic Top-K, designed to address the fundamental challenges of tool retrieval in agents. MagicSelector is a specialized framework capable of translating ambiguous user instructions into executable atomic subtasks and guiding high-precision tool retrieval, effectively mitigating redundant noise and severe context distraction in out-of-domain (OOD) scenarios. We empower MagicSelector with these capabilities through three key contributions: (1) a preference-guided counterfactual task decomposition mechanism that utilizes a counterfactual reward to quantify the marginal causal gain of decomposition on retrieval ranking, effectively imposing fine-grained structural supervision on logical coherence; (2) a progressive tool reranking method driven by self-distillation hard negative mining, which optimizes both point-wise and list-wise relevance to enhance fine-grained discrimination among highly similar tools; and (3) a dual semantic boundary-aware dynamic Top-K strategy that adaptively monitors reranking score cliffs and inter-tool semantic shifts to dynamically truncate the candidate list, maximizing relevant tool recall while filtering long-tail noise. Evaluated on MTDTool, the first task decomposition benchmark we constructed tailored for mobile multi-turn interactions with process-level annotations, MagicSelector yields promising performance. Extensive experiments demonstrate that MagicSelector significantly outperforms state-of-the-art methods in terms of tool retrieval accuracy, OOD generalization capability, and overall token efficiency, thereby demonstrating the effectiveness of our proposed framework.

cs.IR

MVLA-GR: A Phase-Free Multipath-Based Geometry Reconstruction Method via Multi-View Likelihood Accumulation for ISAC

Integrated sensing and communication (ISAC) enables wireless systems to reuse communication signals for environmental sensing, where reconstructing the geometry of surrounding objects is a representative sensing task. However, many conventional methods rely on coherent processing and require accurate phase information, which is often hard to guarantee in practical communication systems, particularly at high carrier frequencies. To address this problem, this paper proposes a Multi-View Likelihood Accumulation Geometry Reconstruction (MVLA-GR) method based on channel impulse response (CIR) measurements, which uses only delay and power observations without requiring phase information. The method extracts dominant multipath components from each observation, and for each candidate spatial location, accumulates components across views whose propagation distances match the location as supporting evidence. A soft distance-matching kernel is introduced to tolerate range estimation errors and viewpoint-dependent scattering migration, and the received power of each component is used as a reliability weight. A joint thresholding strategy combining response magnitude and angular support continuity then converts the continuous support map into a binary geometry estimate. Ray-tracing simulations on canonical and complex targets, as well as real-world vehicle measurements at 36 GHz, demonstrate that MVLA-GR can effectively recover target geometry, providing a low-complexity phase-free solution for ISAC.

eess.SP

DeepRT Engine: A Unified GPU-Parallel Ray-Tracing Framework with Hybrid SBR-IM Path Search for 6G Digital Twin Channel

Digital twin channel (DTC) aims to establish a real-time digital counterpart of physical wireless channels for reproducing and predicting site-specific propagation characteristics. As a high-precision channel computation method for realistic propagation scenarios, ray tracing (RT) serves as a key enabler for DTC construction. However, conventional RT suffers from high complexity under serial path-searching workflows. This letter proposes DeepRT Engine (DeepRT-E), a parallel RT acceleration architecture with a three-stage physically-inspired pipeline for real-time DTC construction. Firstly, DeepRT-E constructs a bounding volume hierarchy (BVH) to partition the scene and reduce redundant ray-surface intersections. Secondly, the shooting and bouncing rays (SBR) algorithm is executed through a ray-level parallel tracing framework to identify candidate surface sequences and prune the search space of the image method (IM). Finally, a parallel batched IM solver refines the retained candidates for accurate propagation-path recovery. Simulation results show that DeepRT-E reduces runtime by 96.3% and achieves a converged error of only 0.001 dB, outperforming Wireless InSite and Sionna in efficiency and accuracy.

eess.SP

Human-as-Humanoid: Enabling Zero-Shot Humanoid Learning from Ego-Exo Human Videos with Human-Aligned Embodiments

Vision-language-action (VLA) models across robot embodiments require high-quality observation--action supervision to learn deployable action distributions, yet scaling such robot data remains difficult, especially for high-DoF humanoids. Teleoperation provides controller-aligned supervision, while human egocentric videos capture diverse bimanual manipulation but do not directly provide executable robot actions. We introduce Human-as-Humanoid, a human-to-humanoid supervision framework that enables near-real-time human-centric action generation, making human demonstrations usable for high-DoF humanoid VLA training by jointly aligning the robot embodiment, the sensing setup, and the action-label interface. Built on PrimeU, a human-aligned 60-DoF upper-body humanoid, Human-as-Humanoid uses synchronized ego-exo videos to pair deployment-aligned egocentric observations with exocentric motion recovery, retargets the recovered human motion through staged Inverse Kinematics (IK) into controller-aligned 60-DoF action chunks, and trains the VLA model with Forward Kinematics (FK)-aware supervision to preserve wrist and fingertip task-space geometry. This converts large-scale human demonstrations from visual observations into executable observation--action supervision for the target humanoid. Experiments validate the conversion chain at the motion-recovery, robot-action-space, and real-robot deployment levels. Human-as-Humanoid yields a 4.8--7.2x raw demonstration-throughput gain over humanoid teleoperation in our data-collection analysis, and on several downstream tasks, policies post-trained only with the converted human labels generalize to real-robot deployment without target-task robot demonstrations. The official project website is available at https://zgc-embodyai.github.io/Human-as-Humanoid.

cs.RO

BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents

Stepwise group-based RL is an attractive way to train long-horizon LLM agents without a learned critic: it reuses multiple sampled rollouts to estimate local advantages. Its weakness is less visible but more fundamental: every group-relative estimator assumes that the steps it compares are equivalent for credit assignment. We show that current agentic variants violate this assumption through a state-action credit mismatch. The observation-hash partition is overly fine on the state side, creating singleton groups with zero step-level signal, while a single within-group mean is too coarse on the action side, mixing state-value estimation with action-specific credit. We introduce BiPACE (Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation), a drop-in advantage estimator that fixes both sides without adding a critic, auxiliary loss, or extra rollouts. BiGPO clusters steps by cosine distance in the actor's own hidden-state geometry, an empirical policy-induced proxy for bisimulation that substantially lowers the singleton rate left by observation hashing. PACE then recenters returns within each behavioral cluster using action-conditioned peer baselines; its Q-style instance estimates a local Q(s,a)-V(s) nonparametrically. On ALFWorld/Qwen2.5-7B, BiPACE_Q raises overall validation success from GiGPO's 90.8 to $97.1\pm0.9$ over three seeds, and crosses the 95% threshold on every seed, which GiGPO never does within the same budget. On Qwen2.5-1.5B it reaches $93.5\pm1.2$ versus GiGPO's 86.7, and on WebShop and TextCraft it improves over GRPO and GiGPO at both model scales. The measured BiPACE-specific overhead is 11.3% of a single training-step wall time. Yet it changes the estimator's comparison unit from surface identity to approximate behavioral equivalence plus action-side counterfactuals. The code is available at https://github.com/TianxiangZhao/BiPACE.

cs.CL

Point-Voxel Absorbing Graph Representation Learning for Event Stream based Recognition

Sampled point and voxel methods are usually employed to downsample the dense events into sparse ones. After that, one popular way is to leverage a graph model which treats the sparse points/voxels as nodes and adopts graph neural networks (GNNs) to learn the representation of event data. Although good performance can be obtained, however, their results are still limited mainly due to two issues. (1) Existing event GNNs generally adopt the additional max (or mean) pooling layer to summarize all node embeddings into a single graph-level representation for the whole event data representation. However, this approach fails to capture the importance of graph nodes and also fails to be fully aware of the node representations. (2) Existing methods generally employ either a sparse point or voxel graph representation model which thus lacks consideration of the complementary between these two types of representation models. To address these issues, we propose a novel dual point-voxel absorbing graph representation learning for event stream data representation. To be specific, given the input event stream, we first transform it into the sparse event cloud and voxel grids and build dual absorbing graph models for them respectively. Then, we design a novel absorbing graph convolutional network (AGCN) for our dual absorbing graph representation and learning. The key aspect of the proposed AGCN is its ability to effectively capture the importance of nodes and thus be fully aware of node representations in summarizing all node representations through the introduced absorbing nodes. Extensive experiments on multiple event-based classification benchmark datasets fully validated the effectiveness of our framework.

cs.CV

Ill-Posedness Analysis of CSI-Based Electromagnetic Inverse Scattering for Material Reconstruction in ISAC Systems

Channel state information (CSI)-based electromagnetic inverse scattering for material reconstruction in ISAC systems enables physics-grounded, material-aware DT. Yet the resulting CSI-induced scattering operator is often severely ill-conditioned. To understand the origin of the ill-posedness, this paper analyzes the mathematical properties of the electromagnetic inverse problem and investigates the operator structure of the ISAC scattering matrix jointly shaped by in-domain scattering responses and Tx/Rx propagation channels. We show that background-related matrix columns are highly coherent and dominate the near rank deficiency, whereas scatterer-related columns are comparatively weakly correlated; their coherence decreases with the number of probing frequencies and thus contributes to the effective rank. Motivated by this analysis, we prove that restricting the ROI around the true scatterer yields a provable condition-number reduction and a tightened CRLB, and we quantify the impact of ROI mismatch numerically. To operationalize these insights, an ROI-constrained QP framework is adopted, where a linear sampling method delineates a coarse ROI and the QP update is performed in the reduced subspace. Full-wave FDTD simulations over multiple geometries and SNR validate pronounced conditioning improvement, substantial complexity savings, and improved robustness, consistent with the proposed analysis, compared with the full-domain formulation.

eess.SP

WorldCoder-Bench: Benchmarking Physically Grounded 3D World Synthesis

Large language models (LLMs) are increasingly asked not only to write static interfaces, but to construct executable interactive worlds from natural language. Browser-native 3D, commonly built with Three.js, is a natural next frontier: generated programs must integrate assets, obey spatial and physical constraints, and keep user-facing controls synchronized with hidden runtime state. Existing web-generation benchmarks and evaluators, however, largely observe only pixels or DOM nodes, while the mechanics of a Three.js world unfold inside an opaque . We introduce WorldCoder-Bench, a benchmark for autonomous, physically grounded 3D world synthesis. WorldCoder-Bench contains 2,026 expert-curated tasks across Simulation, Rendering, and Application scenarios, with optional .glb assets and hidden behavioral contracts. We further propose StateProbe, an execution-based protocol that probes generated programs in a sandboxed browser and verifies hidden, mutation-hardened contracts over runtime states and transitions. Beyond verification coverage, we report Return on Automation and Time Efficiency Multiplier to measure correctness-adjusted cost and time savings. Across nine frontier models, the best system reaches only 27.8% verification coverage on WorldCoder-Core and 19.9% on WorldCoder-Robust, with failures dominated by state-schema drift and broken interaction chains rather than missing scene elements. Utility metrics further show that cheap or fast models can still provide substantial value on easier domains. WorldCoder-Bench is available at https://anonymous.4open.science/r/WorldCoder-Bench/.

cs.AI