SearcharxivSearch

arXiv subjects

Ziheng Wang

Publications and source records attributed to Ziheng Wang.

At least 19 recordsLinked to original sources

VLA-ULAP: Interleaving Cloud VLA Calls with Ultra-Lightweight Local Action Prediction at the Edge

Billion-parameter vision--language--action (VLA) policies demand substantial onboard power, while communication delays in remote inference hinder timely responses. We propose VLA-ULAP, which interleaves remote VLA calls with an Ultra-Lightweight Local Action Predictor (ULAP). With approximately 7.4M parameters including the frozen vision encoder, ULAP combines current views, proprioception, and executed action history to predict chunks in one pass. Trained independently, it requires no VLA hidden states, online verification, or server round trips. On Jetson Orin Nano, ULAP takes 19.9 ms and 0.183 J per inference, compared with 284.3 ms and 50.55 J for GR00T on RTX A6000. Across three simulated base-policy/benchmark pairs, selected operating points remove 48.8--76.7\% of VLA calls while retaining 95.0--97.5\% of the baseline success rate. Against local VLA-acceleration alternatives on VLA-JEPA, ULAP uses an estimated 49.2\% less inference time and 51.0\% less GPU energy per successful episode than ACT at comparable success rates, and 77.1\% less time and 79.9\% less energy than SP-VLA at equal success rates. Physical SO-101 experiments retain 95.2--100\% of the baseline success rate across seen and held-out placements while reducing inference time by an estimated 47.9--58.0\% and inference-device energy by 52.1--62.5\%, based on successful-episode call counts and measured device costs. Faster responses also improve dynamic-task success rates: in latency-aware LIBERO-Safety simulation, VLA-ULAP exceeds $π_{0.5}$ by 11.0 and 15.5 percentage points on two tasks while approximately halving VLA calls.

cs.RO

Analytical Channel Modeling and Stability Aware Optimization of Optical Inter Satellite Links

Optical inter-satellite links (OISLs) are key enablers for high-capacity space networks and next-generation satellite constellations. However, their extreme directionality makes link reliability highly sensitive to platform-induced pointing jitter, which causes random misalignment between the transmitter and receiver beams. In this paper, we develop a tractable closed-form statistical channel model for point-to-point OISLs subject to independent pointing errors at both terminals. Accurate Gaussian main-lobe approximations are applied to the transmitter far-field pattern and receiver coupling efficiency. This transforms the diffraction-based channel response into closed-form expressions for the channel-gain distribution, outage probability, and ergodic capacity. The analytical results are validated through Monte Carlo simulations and used to study the impact of terminal stability, beam divergence, and link margin on OISL performance. The results show that outage probability is governed by the weaker terminal in terms of pointing stability, while improving only the stronger terminal provides minimal additional benefit. In contrast, the ergodic-capacity penalty depends on the combined stability of both terminals, revealing a fundamental distinction between reliability and throughput metrics. The proposed framework provides practical design guidelines for selecting beam parameters and specifying pointing and tracking requirements under varying levels of platform instability.

cs.IT

Demonstrate of High-Performance Top-Gate ALD Crystalline In2O3 Transistor Enabled by Lattice-Matched HfO2 and In2O3 Heterostructure

In this work, we demonstrate high-mobility top-gate (TG) atomic-layer-deposited (ALD) crystalline In2O3 transistors through simultaneous interface and crystallinity engineering. First, a HfO2/In2O3/HfO2 stack is employed, enabling epitaxial-like crystallization of the ultrathin In2O3 channel, because of the lattice matching between monoclinic phase HfO2 and cubic phase In2O3. Second, an oxygen-rich gate insulator process is applied using high-dose O3 precursor and elevated deposition temperature, effectively suppressing oxygen scavenging during gate dielectric deposition, significantly reducing interfacial defect formation. Third, the homogeneous In-O bonding network in crystalline In2O3 exhibits substantially enhanced resistance to oxygen scavenging by source/drain contacts, which significantly improves the immunity to threshold voltage (VTH) roll-off at short channel length compared to amorphous In2O3. As a result, high-performance TG long-channel In2O3 transistors are achieved with a high mobility of 163 cm2/V s and a steep subthreshold slope of 64 mV/dec. High-performance TG short-channel In2O3 transistors with high ION of 1650 μA/μm at VD of 1 V, large on/off ratio over 1010 and VTH of -0.27 V are demonstrated. These results establish lattice-engineered crystalline In2O3 as an effective strategy for high-mobility, aggressively scaled TG oxide transistors suitable for BEOL-compatible applications.

cond-mat.mtrl-sci

High-Performance Scaled P-Type SnOx Transistor by Atomic Layer Deposition with CFET Integration

The development of high-performance p-type oxide semiconductors is essential for realizing complementary logic for monolithic 3D integration, yet p-type oxide semiconductors still exhibit substantially inferior performance compared with their n-type counterparts. In this work, we demonstrate high-performance p-type SnOx transistors by atomic layer deposition (ALD), as back-end-of-line compatible devices for monolithic 3D integration. The SnOx transistors exhibit high field-effect mobility of 6.9 cm2/Vs, low subthreshold swing (SS) of 185 mV/dec, decent on/off ratio (ION/IOFF) of 1.8*104 and high bias stability. By scaling the channel length down to 80 nm, a high on-current of 38.7 mA/mm at VDS of -1 V is achieved. It is understood that precursor and reaction engineering to suppress Sn4+ component in SnOx film are the key for performance enhancement. Furthermore, a complementary field-effect transistor with ALD SnOx p-FET vertically stacking on ALD In2O3 n-FET is also demonstrated, achieving maximum voltage gain of 21 V/V at VDD of 4 V. These findings suggest ALD SnOx as a promising candidate for scaled high-performance BEOL p-type transistors.

cond-mat.mtrl-sci

Giant Surface-driven Nonlinear Hall Effect in BiTeCl at Room Temperature

The nonlinear Hall effect (NLHE) provides a pathway to generate a Hall response in time-reversal-symmetric yet inversion-symmetry-broken systems. NLHE can rectify an alternating current into a transverse direct voltage, making it attractive for radio-frequency rectification, energy harvesting, and terahertz detection, applications for which device miniaturization remains a central pursuit. In this context, the inherent inversion symmetry breaking at surfaces is particularly appealing: because symmetry is necessarily broken at the surface of any crystal, irrespective of whether its bulk is centrosymmetric, surface-driven nonlinear responses lift the stringent constraint on bulk symmetry and open a route toward compact device architectures. Here we report the observation of a giant, surface-driven second-order nonlinear Hall effect in the Rashba-type polar semiconductor BiTeCl at room temperature. The determined second-order nonlinear Hall susceptibility at 300 K reaches 1.68 $μ$mV$^{-1}$, which is 80 times larger than that of the best previously reported surface-dominated systems. We attribute this giant response to the synergistic interplay between BiTeCl's polar crystal structure and its rich surface states: the polar stacking renders the top and bottom surfaces inequivalent, so that the nonlinear response originates from a single surface without compensation from the other. Symmetry and scaling analyses suggest that both skew-scattering and side-jump mechanisms contribute to the observed effect. Our findings not only identify BiTeCl as a promising platform for future applications utilizing the NLHE, but also establish the asymmetry between the opposite surfaces of a polar crystal as a general design principle for discovering surface-driven materials with larger nonlinear Hall responses.

cond-mat.mtrl-sci

Diffractive optical element for super-Gaussian beam shaping on intersatellite optical communications

Pointing jitter can significantly degrade the performance of intersatellite optical communication links. This work investigates diffractive optical beam shaping as a means of generating super-Gaussian profiles with reduced sensitivity to transmitter misalignment. Phase screens are designed using a Gerchberg--Saxton phase-retrieval algorithm and evaluated for different super-Gaussian orders. Lower-order profiles are reproduced accurately, whereas higher orders are increasingly limited by numerical discretization and finite-aperture effects. The practical implementation of the phase screens using fused-silica diffractive optical elements is assessed through sensitivity analyses of radial manufacturing resolution, phase quantization, phase-depth errors, and incident-beam wavefront aberrations. The results provide manufacturing tolerances and wavefront-quality requirements for preserving the desired beam shape.

physics.optics

MatrAIx: Simulating the World with 8.3 Billion Persona Agents

Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. MatrAIx has three core components: First, Persona 8B contains 8.3 billion persona records represented by 1,290 categorical dimensions. Records are either sampled from a dependency graph that preserves correlated attributes or derived from human-authored profiles. We release a quality-filtered coreset of approximately 1 million personas, comprising 599,847 human-grounded and 400,000 synthetic records. Second, the MatrAIx Playground provides four environments in which diverse users evaluate and interact with digital products: Survey, AI Chatbot, Web, and App. Third, MatrAIx provides 1,010 application tasks spanning more than 25 domains, including Commerce, Software, Finance, and Healthcare. We conducted 18,189 evaluation trials across eight representative tasks. Persona agents were powered by three LLMs: Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5. The resulting feedback captures how decisions and preferences vary across persona backgrounds, including hesitation after a price increase, willingness to continue after an AI assistant fails, and latency tolerance. We conducted two main validation studies: First, a 400-trial controlled study evaluated persona adherence across ten behavioral attributes and all four environments. The declared behavior was expressed or correctly suppressed in 366 trials (91.5%). Second, human and LLM judges evaluated the extraction quality of human-grounded personas. Overall, MatrAIx provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.

cs.AI

Beyond Symmetric Fusion: Exploiting Task-Dependent Modality Strengths for RGB-Event Small Object Detection

State-of-the-art RGB-Event detectors improve the detection of small, fast-moving objects by combining complementary features from RGB and Event data, yet they typically fuse the two modalities into a unified representation for both localization and classification. Such a task-symmetric design is inconsistent with the intuition that the two modalities should play different roles according to their task-specific strengths. To examine this issue, we conduct a modality-specific evaluation and find that the relative advantage of the two modalities reverses across tasks: Event data are substantially more effective for class-agnostic localization, whereas RGB data provide stronger category evidence within localized target regions. Motivated by this task-dependent asymmetry, we propose an Asymmetric Event-RGB Object Detection Transformer (AERODet). During class-agnostic localization, Scale-wise Uncertainty-aware Reliability Estimation (SURE) calculates the relative reliability of the two modalities from their objectness response heatmaps and accordingly calibrates their contributions when the decoder aggregates multimodal features. Once the candidate boxes are obtained, Task-Decoupled Semantic Refinement (TDSR) decouples classification from localization and uses RGB RoI features for fine-grained classification. Extensive experiments on FRED and NeRDD demonstrate that AERODet achieves state-of-the-art performance. In particular, it surpasses the strongest RGB-Event baseline by 10.7 mAP points on the FRED challenging split.

cs.CV

Minimizing Worst-Case Weighted Latency for Multi-Robot Persistent Monitoring: Theory and RL-Based Solutions

We study multi-robot persistent monitoring on weighted graphs, where node weights encode monitoring priorities and edge weights encode travel distances. The goal is to design joint robot trajectories that minimize the worst-case weighted latency across all nodes over an infinite time horizon. The widely adopted worst-case latency objective evaluates team performance over the entire time horizon and therefore may fail to distinguish strategies with poor transient behavior but strong asymptotic performance. To address this limitation, we propose a family of tail-performance objectives that generalize the standard objective and study the resulting functional optimization problems. We establish several key theoretical properties, including the existence of optimal strategies, relationships among the proposed objectives and their corresponding optimization problems, approximation by periodic solutions to arbitrary accuracy, and reductions to event-driven decision models with discretized waiting times. Building on these results, we construct an equivalent event-driven Markov decision process (MDP), called the Tail Worst-case Latency-Optimizing Markov Decision Process (TWLO-MDP), which reformulates the tail-performance objective as a standard average-reward criterion. We then develop reinforcement-learning-based solution methods for the TWLO-MDP and introduce the multi-robot monitoring benchmark (M2Bench), a unified platform that supports the evaluation and comparison of heuristic and learning-based monitoring algorithms. Experiments on synthetic and realistic monitoring scenarios show that our methods effectively reduce the worst-case weighted latency and outperform representative baselines.

cs.RO

Knowledge-Guided Synthetic Bug Feedback for LLM-Based Unit Test Generation

Large language models (LLMs) have opened new opportunities for unit test generation, but executable tests do not necessarily reveal real defects. This paper studies how historical real-bug mechanisms can be transformed into executable feedback targets for LLM-based unit test generation. The proposed framework constructs structural and semantic representations of real-bug records, retrieves mechanisms applicable to a focal method, and instantiates them as synthetic bugs that guide iterative test enhancement. We evaluate the approach on method-level real-bug detection tasks from Defects4J and show that mechanism-guided synthetic-bug feedback improves real-bug detection over execution-, coverage-, mutation-, knowledge-, and search-based baselines. The results suggest that organizing real-bug mechanisms as retrievable and executable feedback targets is an effective way to guide generated tests toward bug-triggering inputs and behavioral oracles.

cs.SE

RED-Sphere: Hyperspherical Residual Edge Debiasing for Cross-Population Fundus Disease Domain Generalization

Medical image classifiers are often trained within one source population, yet clinical deployment requires robustness to patients whose appearance, acquisition style, and disease prevalence differ from the source cohort. Existing fairness and robustness methods often require group supervision or treat appearance variation as an undifferentiated nuisance, which is insufficient when population-correlated low-level cues and lesion evidence share edge and texture structure. We study a strict source-only cross-population setting, where external populations are unseen during optimization, validation, scheduling, hyperparameter and model selection. We propose RED-Sphere, a plug-and-play robustness framework for image classification under unseen population shifts. It estimates shortcut-sensitive nuisance responses with an edge and feature energy prior, attenuates dominant responses through residual soft gating, regularizes masked nuisance views with counterfactual-inspired consistency and separation losses, and predicts labels with normalized spherical prototypes. It favours angular semantic evidence over source-correlated activation magnitude while preserving lesion structure. Although demonstrated on 2D Scanning Laser Ophthalmoscopy (SLO) fundus classification for Age-Related Macular Degeneration (AMD) and Diabetic Retinopathy (DR), RED-Sphere is not tied to retinal anatomy: the same principle can be adapted with modality-specific nuisance priors wherever appearance shortcuts and semantic evidence are entangled. Under a strict White-only Harvard-FairVision protocol, RED-Sphere improves held-out macro-F1 across all 20 task and backbone comparisons, with average gains of 1.28 and 2.98 F1 points on AMD and DR. Gains in AUC and PR-AUC, visual diagnostics, ablations, and sensitivity analyses further support stronger external semantic alignment and more stable angular disease geometry.

cs.CV

High-Mobility and High-Reliability Top-Gate Oxide Semiconductor Transistors by Oxygen Engineering

In this work, we investigate the role of oxygen (O) on the performance of top-gate (TG) atomic-layer-deposited (ALD) oxide semiconductor transistors. The results reveal distinct defect characteristics and positive bias temperature instability (PBTI) degradation mechanisms between oxygen-rich (O-rich) and oxygen-deficient (O-poor) devices. It is found that an O-rich device fabrication process followed by O-free annealing can effectively achieve TG indium-rich (In-rich) oxide semiconductor transistors with high mobility, high reliability and high stability in hydrogen environment because O-rich process can suppress oxygen vacancies and their interaction with hydrogen, while O-free annealing plays a critical role in minimizing the formation of O-rich defects such as oxygen dimers (O-O bonds). Consequently, TG In-rich transistors with high mobility, steep subthreshold slope, and high PBTI reliability at high temperature are demonstrated. The understanding of O-rich defects provides a new insight to overcome the mobility-stability trade-off.

cond-mat.mtrl-sci

Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety

General-purpose models often struggle to reliably identify and understand real-world multimodal risks, largely due to the inherent multimodal adversarial nature of content and AI safety. We present Yuvion VL, a family of multimodal large language models purpose-built for content and AI safety, with both instruction-tuned and reasoning-oriented variants. Yuvion VL addresses this gap by treating safety as an inherently adversarial and multimodal problem and designing the entire pipeline around adversarial robustness. For data construction, we develop an automated pipeline integrating adversarial-aware data synthesis with multi-stage quality control, producing large-scale, high-quality multimodal samples augmented with domain knowledge and reasoning annotations. For training, we adopt a three-stage pipeline that includes continued pretraining for risk-concept cross-modal alignment, instruct post-training for production-grade safety tasks, and reasoning post-training for enhanced interpretability and performance in complex tasks. We further introduce Confuse-then-Contrast Fine-Tuning, a contrastive framework that mines model-specific confusions and constructs multi-image contrastive groups to enforce explicit discrimination of fine-grained visual-semantic elements, enabling the model to distinguish between visually similar cases with different safety implications in adversarial safety tasks. To support rigorous evaluation, we further introduce Yuvion VL RiskEval (YVRE), a collection of benchmarks covering diverse open and internal evaluations, with a focus on content and AI safety, adversarial robustness, and real-world capability requirements. Experiments show that Yuvion VL-32B achieves industry-leading safety performance, surpassing comparably sized open-source models and best closed-source commercial models, while maintaining comparable general capabilities.

cs.CV

Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety

As large language models are increasingly deployed in real-world systems, safety failures can still lead to harmful outputs and dangerous misuse. We argue that the essence of safety is adversarial: many failures arise not from natural inputs alone, but from strategic attempts to evade model policies and safeguards. However, existing general-purpose model development largely overlook this adversarial nature, and often remain insufficient for realistic safety scenarios involving planning, tool use, and multi-step reasoning, causing measured safety performance to overestimate real deployment robustness. To address this gap, we present Yuvion LLM, a large language model built for adversarially robust content safety and broader AI safety. Yuvion LLM treats adversarial robustness and agentic capability as first-class objectives. Its pipeline combines adversarially aware data construction, knowledge-enhanced continued pretraining, and policy-grounded multi-task safety post-training, including risk-aware supervised fine-tuning and reinforcement learning-based policy optimization, together with safety-aware agentic reinforcement learning for tool use and multi-step reasoning in complex safety scenarios. We further introduce the Yuvion LLM RiskEval (YLRE), a collection of 93 benchmarks across four evaluation categories, covering diverse open and internal evaluations with a focus on safety, adversarial robustness, and real-world capability requirements. Across these evaluations, Yuvion LLM demonstrates clear advantages on safety-focused benchmarks and particularly strong robustness under adversarial conditions, while maintaining solid overall capability. Notably, Yuvion-8B outperforms most state-of-the-art baselines, including substantially larger models such as GPT-5.4 and Qwen3-MAX, on several safety tasks.

cs.CL

MoVerse: Real-Time Video World Modeling with Panoramic Gaussian Scaffold

We present MoVerse, a real-time video world model that creates an interactively navigable scene from a single narrow-field-of-view image. This setting is challenging because the input observes only a small fraction of the environment, while interactive roaming requires a complete surrounding world, persistent geometry, controllable camera motion, and temporally coherent high-fidelity observations. MoVerse addresses this problem by separating world construction from observation rendering. It first expands the input into a gravity-aligned 360$^\circ$ panorama with topology-aware diffusion, closing the missing field of view before 3D reasoning. It then lifts the panorama into a persistent 3D Gaussian scaffold using panoramic geometry-aware residual prediction, yielding a dense and directly renderable spatial memory. Finally, a Gaussian-conditioned video renderer translates scaffold renderings along user-specified camera trajectories into photorealistic video. To make this renderer practical for interaction, we train a bidirectional diffusion teacher for high-quality conditional rendering and distill it into a causal autoregressive student for bounded-latency streaming. This design combines the controllability and long-range consistency of explicit 3D representations with the perceptual quality of generative video models. MoVerse supports real-time scene roaming at 8~FPS on a single NVIDIA RTX~4090 GPU, demonstrating a practical path toward single-image world creation with interactive video output.

cs.CV

Weak Convergence Analysis of Online Neural Actor-Critic Algorithms

We prove that a single-layer neural network trained with the online actor critic algorithm converges in distribution to a random ordinary differential equation (ODE) as the number of hidden units and the number of training steps $\rightarrow \infty$. In the online actor-critic algorithm, the distribution of the data samples dynamically changes as the model is updated, which is a key challenge for any convergence analysis. We establish the geometric ergodicity of the data samples under a fixed actor policy. Then, using a Poisson equation, we prove that the fluctuations of the model updates around the limit distribution due to the randomly-arriving data samples vanish as the number of parameter updates $\rightarrow \infty$. Using the Poisson equation and weak convergence techniques, we prove that the actor neural network and critic neural network converge to the solutions of a system of ODEs with random initial conditions. Analysis of the limit ODE shows that the limit critic network will converge to the true value function, which will provide the actor an asymptotically unbiased estimate of the policy gradient. We then prove that the limit actor network will converge to a stationary point.

cs.LG

Intuitive Surgical SurgToolLoc and SurgVU Challenges Results: 2022-2025

Robotic assisted (RA) surgery promises to transform surgical intervention. Intuitive Surgical is committed to fostering these changes and the machine learning models and algorithms that will enable them. With these goals in mind we have invited the surgical data science community to participate in a yearly competition hosted through the Medical Imaging Computing and Computer Assisted Interventions (MICCAI) conference. With varying changes from year to year, we have challenged the community to solve difficult machine learning problems in the context of advanced RA applications. Here we document the results of these challenges, focusing on surgical tool localization (SurgToolLoc) and surgical visual understanding (SurgVU). The publicly released dataset that accompanies these challenges is detailed in a separate paper arXiv:2501.09209 [1].

cs.CV

MoCam: Unified Novel View Synthesis via Structured Denoising Dynamics

Generative novel view synthesis faces a fundamental dilemma: geometric priors provide spatial alignment but become sparse and inaccurate under view changes, while appearance priors offer visual fidelity but lack geometric correspondence. Existing methods either propagate geometric errors throughout generation or suffer from signal conflicts when fusing both statically. We introduce MoCam, which employs structured denoising dynamics to orchestrate a coordinated progression from geometry to appearance within the diffusion process. MoCam first leverages geometric priors in early stages to anchor coarse structures and tolerate their incompleteness, then switches to appearance priors in later stages to actively correct geometric errors and refine details. This design naturally unifies static and dynamic view synthesis by temporally decoupling geometric alignment and appearance refinement within the diffusion process. Experiments demonstrate that MoCam significantly outperforms prior methods, particularly when point clouds contain severe holes or distortions, achieving robust geometry-appearance disentanglement.

cs.CV